A prospecting data anomaly detection and early warning system

Through the multi-dimensional data fusion and spatial anomaly analysis of prospecting data abnormality detection and early warning system, the problems of incomplete and inaccurate abnormality detection in the existing technology are solved, and more efficient abnormality detection and early warning are achieved.

CN119830224BActive Publication Date: 2025-06-13THE THIRD INST OF GEOLOGY & MINERALS EXPLORATION GANSU PROVINCIAL BUREAU OF GEOLOGY & MINERALS EXPLORATION & DEV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510303575.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing anomaly detection methods for prospecting data cannot fully capture the correlation between multi-source data, resulting in inaccurate abnormal detection and inability to capture the abnormal characteristics of prospecting data in space, resulting in incomplete detection and low accuracy.

Method used

Provide a prospecting data abnormality detection and early warning system, including a data acquisition module, a data fusion module, a data processing module, a detection module and an early warning module. Through multi-dimensional data fusion, multi-level abnormality detection and spatial abnormality analysis, abnormality detection indicators are calculated and the real abnormality data and spatial abnormality characteristics are determined, and early warning information is finally generated.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of the abnormal detection of prospecting data, can more accurately capture the correlation and spatial abnormal characteristics between multi-source data, and improves the effect of abnormal detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830224B_ABST
    Figure CN119830224B_ABST
Patent Text Reader

Abstract

The present invention provides a prospecting data anomaly detection and early warning system, which relates to the technical field of data processing. It includes: a data acquisition module for collecting various prospecting data within a preset time period to obtain original data; a data fusion module for performing data fusion on the original data to obtain a multi-dimensional data set corresponding to each moment; a data processing module for calculating an anomaly detection index according to the differences between the multi-dimensional data sets at adjacent moments; a first detection module for determining real anomaly data and fault data according to the distribution of the anomaly detection index; a second detection module for determining the spatial anomaly characteristics of the real anomaly data according to a spatial analysis algorithm to obtain spatial anomaly data; and a warning module for generating corresponding warning information according to the spatial anomaly data. The present invention solves the problems of incomplete anomaly detection and low accuracy of prospecting data in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a prospecting data anomaly detection and early warning system. Background Art

[0002] Prospecting data collection and anomaly detection are important steps in exploring mineral resources (such as metal ores, non-metal ores, etc.) in exploration geology. Various prospecting data play a key role in the positioning, development, and management of mineral resources. In order to accurately and efficiently explore mineral resources in geology, a prior patent discloses a method for prospecting data collection and logging data anomaly detection. Based on time series analysis, by calculating the differences in target logging data (such as acoustic data, resistivity data, natural gamma radiation data) at adjacent times, the initial anomaly degree, overall deviation degree, and actual anomaly degree are calculated, and then the anomaly data is screened out. And the target logging data is matched by the Spearman rank correlation coefficient to distinguish real anomaly data and fault data. However, the above patent only analyzes a single type of logging data, while prospecting data usually includes multi-source data such as geophysical data, geochemical data, and geological data. The analysis of a single data source cannot comprehensively capture the correlation between multi-source data, resulting in inaccurate anomaly detection. And the above patent is mainly based on time series analysis and cannot capture the spatial anomaly characteristics of prospecting data, resulting in incomplete and inaccurate anomaly detection. Summary of the Invention

[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a prospecting data anomaly detection and early warning system, which solves the problems of incomplete and inaccurate prospecting data anomaly detection in the prior art.

[0004] To achieve the above purpose, the present invention provides the following solutions:

[0005] A prospecting data anomaly detection and early warning system, comprising:

[0006] A data collection module, configured to collect a variety of prospecting data within a preset time period to obtain original data;

[0007] A data fusion module, which fuses the original data to obtain a multi-dimensional data set corresponding to each moment;

[0008] A data processing module, configured to calculate anomaly detection indicators according to the differences in the multi-dimensional data sets at adjacent times, and the anomaly detection indicators include: an initial anomaly degree, an overall deviation degree, and an actual anomaly degree;

[0009] A first detection module, configured to determine real anomaly data and fault data according to the distribution of the anomaly detection indicators;

[0010] A second detection module, configured to determine the spatial anomaly characteristics of the true anomaly data according to a spatial analysis algorithm, so as to obtain spatial anomaly data;

[0011] An early warning module, configured to generate corresponding early warning information according to the spatial anomaly data.

[0012] Preferably, the data fusion module includes:

[0013] A data preprocessing sub-module, configured to perform data cleaning and normalization processing on the original data to obtain the original data under the same dimension;

[0014] A trend calculation sub-module, configured to decompose the prospecting data of each dimension of the original data by using a time series decomposition algorithm to obtain a data change trend sequence;

[0015] A correlation calculation sub-module, configured to calculate the Pearson correlation coefficient between the change trend sequences of the data of each dimension to obtain a multi-dimensional trend correlation;

[0016] A trend offset sequence calculation sub-module, configured to calculate the absolute value of the difference between the data sequence of each dimension of the original data and the trend sequence to obtain a trend offset sequence;

[0017] An abnormal moment detection sub-module, configured to use the LOF algorithm to detect the abnormal data in the trend offset sequence to obtain a trend abnormal moment;

[0018] A trend jump degree calculation sub-module, configured to calculate the trend jump degree according to the distribution of the trend abnormal moments;

[0019] A fusion participation degree calculation module, configured to calculate a fusion participation degree according to the trend jump degree and the multi-dimensional trend correlation;

[0020] A weight calculation sub-module, configured to calculate a fusion weight according to the fusion participation degree;

[0021] A generation sub-module, configured to perform data fusion on the original data according to the fusion weight to obtain a multi-dimensional data set corresponding to each moment.

[0022] Preferably, the data processing module includes:

[0023] A first index calculation sub-module, configured to calculate the initial anomaly degree of the multi-dimensional data set;

[0024] A second index calculation sub-module, configured to calculate the overall deviation degree according to the initial anomaly degree;

[0025] A third index calculation sub-module, configured to calculate the actual anomaly degree according to the overall deviation degree.

[0026] Preferably, the expression for the initial anomaly degree is:

[0027] ;

[0028] where is the initial anomaly degree, is the observation value of the th dimension at time, is the moving standard deviation of the th dimension at time, is the moving mean of the th dimension at time, and n is the total number of dimensions;

[0029] The calculation formula for the overall deviation degree is:

[0030] ;

[0031] where is the overall deviation degree, is the predicted value of the th dimension at time;

[0032] The calculation expression for the actual anomaly degree is:

[0033] ;

[0034] where is the actual anomaly degree, is the local consistency factor, , and are the first weight coefficient, the second weight coefficient, and the third weight coefficient, respectively;

[0035] The calculation expression for the local consistency factor is:

[0036] ;

[0037] where is the size of the local time window, is the mean of the standard deviations of all dimensions within the time window, is the standard deviation of the th dimension within the time window, is the data sequence of the th dimension within the time window.

[0038] Preferably, the first detection module includes:

[0039] A normalization sub-module for normalizing the anomaly detection metrics to obtain normalized data;

[0040] A modeling sub-module for jointly modeling the anomaly metrics of the normalized data using a multivariate Gaussian distribution to obtain a probability distribution model;

[0041] An anomaly probability calculation sub-module for obtaining the anomaly probability at each moment according to the probability distribution model;

[0042] A clustering analysis sub-module for clustering the anomaly probabilities at each moment to obtain initial true anomaly data and initial fault data;

[0043] A result optimization sub-module for performing secondary classification on the initial true anomaly data and initial fault data to obtain final true anomaly data and final fault data.

[0044] Preferably, the expression of the probability distribution model is:

[0045] ;

[0046] where, is the probability distribution model, is the normalized data, and the normalized data includes: the normalized initial anomaly degree, the normalized overall deviation degree, and the normalized actual anomaly degree, is the covariance matrix, is the mean vector.

[0047] Preferably, the anomaly probability calculation expression is:

[0048] ;

[0049] where, is the anomaly probability.

[0050] Preferably, the second detection module includes:

[0051] A position conversion sub-module for converting the position information of the final true anomaly data into a unified spatial coordinate system to obtain unified position information data;

[0052] A spatial first processing sub-module for analyzing the unified position information data according to a spatial autocorrelation algorithm to obtain spatial distribution characteristics;

[0053] A spatial second processing sub-module for classifying and extracting anomaly characteristics from the spatial distribution characteristics using a spatial clustering algorithm to obtain spatial anomaly data.

[0054] Preferably, the spatial second processing sub-module includes:

[0055] A clustering unit, configured to perform clustering analysis on the spatial distribution features by using a spatial clustering algorithm to obtain spatial aggregation anomalies and spatial isolation anomalies;

[0056] An anomaly extraction unit, configured to extract anomaly features according to the spatial aggregation anomalies and spatial isolation anomalies to obtain spatial aggregation anomaly features and spatial isolation anomaly features;

[0057] A classification unit, configured to perform data classification on the spatial aggregation anomaly features and spatial isolation anomaly features to obtain spatial anomaly data.

[0058] Preferably, the anomaly extraction unit includes:

[0059] An aggregation center position determination subunit, configured to determine the coordinate mean of data points within the spatial aggregation anomaly and spatial isolation anomaly regions to obtain an aggregation center;

[0060] An aggregation radius determination subunit, configured to determine the maximum distance from all data points within the spatial aggregation anomaly and spatial isolation anomaly regions to the coordinate mean to obtain an aggregation radius;

[0061] An aggregation density determination subunit, configured to determine the aggregation density of all data points within the spatial aggregation anomaly and spatial isolation anomaly regions to obtain an aggregation density;

[0062] A feature determination subunit, configured to determine spatial aggregation anomaly features and spatial isolation anomaly features according to the aggregation center, aggregation radius, and aggregation density.

[0063] The present invention discloses the following technical effects:

[0064] The present invention provides a prospecting data anomaly detection and early warning system, including: a data acquisition module, configured to acquire various prospecting data within a preset time period to obtain original data; a data fusion module, configured to perform data fusion on the original data to obtain a multi-dimensional data set corresponding to each moment; a data processing module, configured to calculate an anomaly detection index according to the difference between multi-dimensional data sets at adjacent moments, where the anomaly detection index includes: an initial anomaly degree, an overall deviation degree, and an actual anomaly degree; a first detection module, configured to determine real anomaly data and fault data according to the distribution of the anomaly detection index; a second detection module, configured to determine the spatial anomaly features of the real anomaly data according to a spatial analysis algorithm to obtain spatial anomaly data; and an early warning module, configured to generate a corresponding early warning message according to the spatial anomaly data. The present invention significantly improves the comprehensiveness and accuracy of prospecting data anomaly detection through multi-dimensional data fusion, multi-level anomaly detection, and spatial anomaly analysis. Description of the Drawings

[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0066] Figure 1 Schematic diagram of a prospecting data anomaly detection and early warning system provided by an embodiment of the present invention.

[0067] Explanation of reference numerals:

[0068] 1 - Data acquisition module, 2 - Data fusion module, 3 - Data processing module, 4 - First detection module, 5 - Second detection module, 6 - Early warning module. Specific embodiments

[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0070] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0071] As Figure 1 shown, the present invention provides a prospecting data anomaly detection and early warning system, including:

[0072] A data acquisition module 1, configured to collect various prospecting data within a preset time period to obtain raw data;

[0073] Specifically, the prospecting data includes:

[0074] Geological data: such as rock type, stratigraphic structure, mineral composition, etc.; geophysical data: such as gravity, magnetism, electrical method, seismic, etc.; geochemical data: such as element content in soil, water sample, gas; remote sensing data: such as satellite images, aerial photos, etc.; borehole data: such as core, wellbore, well depth, etc.

[0075] A data fusion module 2, which performs data fusion on the raw data to obtain a multi-dimensional data set corresponding to each moment;

[0076] The data processing module 3 is used to calculate anomaly detection metrics based on the differences in multi-dimensional data sets at adjacent times. The anomaly detection metrics include: initial anomaly degree, overall deviation degree, and actual anomaly degree;

[0077] The first detection module 4 is used to determine real anomaly data and fault data according to the distribution of the anomaly detection metrics;

[0078] The second detection module 5 is used to determine the spatial anomaly characteristics of the real anomaly data according to the spatial analysis algorithm to obtain spatial anomaly data;

[0079] The warning module 6 is used to generate corresponding warning information according to the spatial anomaly data.

[0080] Furthermore, the data fusion module 2 includes:

[0081] The data preprocessing sub-module is used to perform data cleaning and normalization processing on the original data to obtain the original data under the same dimension; scale the data to the range of [0, 1] or [-1, 1].

[0082] The trend calculation sub-module is used to decompose the prospecting data of each dimension of the original data by using the time series decomposition algorithm to obtain the data change trend sequence; use the time series decomposition algorithm to decompose the data into trend, seasonal, and residual parts and extract the change trend sequence.

[0083] The correlation calculation sub-module is used to calculate the Pearson correlation coefficient between the change trend sequences of each dimension data to obtain multi-dimensional trend correlation;

[0084] Specifically, calculate the local change rate (such as difference or derivative) of the trend sequence of each dimension; dynamically adjust the size of the sliding window according to the change rate: shrink the window when the change rate is large, and expand the window when the change rate is small; calculate the Pearson correlation coefficient between the trend sequences within each window; integrate the results of all windows to obtain the multi-dimensional trend correlation matrix.

[0085] The trend offset sequence calculation sub-module is used to calculate the absolute value of the difference between the data sequence and the trend sequence of each dimension of the original data to obtain the trend offset sequence;

[0086] The anomaly moment detection sub-module is used to detect the anomaly data in the trend offset sequence by using the LOF algorithm to obtain the trend anomaly moment;

[0087] The trend jump degree calculation sub-module is used to calculate the trend jump degree according to the distribution of the trend anomaly moments;

[0088] The fusion participation degree calculation module calculates the fusion participation degree according to the degree of trend jump and the multi-dimensional trend correlation;

[0089] The weight calculation sub-module is used to calculate the fusion weight according to the fusion participation degree;

[0090] The generation sub-module is used to perform data fusion on the original data according to the fusion weight to obtain the multi-dimensional data set corresponding to each moment.

[0091] Furthermore, the data processing module 3 includes:

[0092] The first index calculation sub-module is used to calculate the initial anomaly degree of the multi-dimensional data set;

[0093] The second index calculation sub-module is used to calculate the overall deviation degree according to the initial anomaly degree;

[0094] The third index calculation sub-module is used to calculate the actual anomaly degree according to the overall deviation degree.

[0095] Furthermore, the expression of the initial anomaly degree is:

[0096] ;

[0097] Where is the initial anomaly degree, is the th dimension at time, is the th dimension at time, the moving standard deviation, is the th dimension at time, the moving mean, and n is the total number of dimensions; this formula uses the moving mean and standard deviation to dynamically capture the change trend of the data, avoids misjudgment caused by a fixed threshold, and eliminates the influence of different dimension dimensions through standardization processing, making the results more comparable.

[0098] The calculation formula of the overall deviation degree is:

[0099] ;

[0100] Where is the overall deviation degree, is the th dimension at time;

[0101] Specifically, a time series prediction model is introduced to predict the data values under normal conditions, which are compared with the actual observed values to capture the overall deviation trend and use the Euclidean distance to measure the overall deviation degree, and the joint abnormal features of multi-dimensional data are comprehensively considered.

[0102] The calculation expression of the actual abnormal degree is as follows:

[0103] ;

[0104] Wherein, is the actual abnormal degree, is the local consistency factor, , and are the first weight coefficient, the second weight coefficient and the third weight coefficient respectively;

[0105] The calculation expression of the local consistency factor is as follows:

[0106] ;

[0107] Wherein, is the size of the local time window, is the mean value of the standard deviations of all dimensions in the time window, is the standard deviation of the th dimension in the time window, is the th dimension data sequence in the time window.

[0108] Specifically, by calculating the standard deviation within the local time window, the stability characteristics of the data within the local time window are captured, and is used to standardize the local consistency. The larger the value, the more stable the data.

[0109] Furthermore, the first detection module 4 includes:

[0110] A normalization sub-module for normalizing the anomaly detection index to obtain normalized data;

[0111] A modeling sub-module for jointly modeling the anomaly index distribution of the normalized data using a multivariate Gaussian distribution to obtain a probability distribution model;

[0112] An anomaly probability calculation sub-module for obtaining the anomaly probability at each moment according to the probability distribution model;

[0113] A clustering analysis sub-module for clustering the anomaly probabilities at each moment to obtain initial true anomaly data and initial fault data;

[0114] Specifically, clustering algorithms (such as K-Means, DBSCAN) are used to cluster the anomaly probabilities.

[0115] According to the clustering results, the data is divided into initial true anomaly data and initial fault data.

[0116] The result optimization sub-module is used to perform secondary classification on the initial true anomaly data and initial fault data to obtain final true anomaly data and final fault data.

[0117] Specifically, classification algorithms (such as SVM, random forest) are used to perform secondary classification on the initial classification data. According to the classification results, final true anomaly data and final fault data are obtained.

[0118] Furthermore, the expression of the probability distribution model is:

[0119] ;

[0120] where is the probability distribution model, is the standardized data, and the standardized data includes: the standardized initial anomaly degree, the standardized overall deviation degree, and the standardized actual anomaly degree, is the covariance matrix, is the mean vector.

[0121] Specifically, modeling the standardized anomaly detection metrics through a multivariate Gaussian distribution can effectively capture the correlations in multi-dimensional data and provide a probabilistic anomaly detection method.

[0122] Furthermore, the anomaly probability calculation expression is:

[0123] ;

[0124] where is the anomaly probability.

[0125] Specifically, if the data point at a certain moment falls within the distribution range of normal data (i.e., is large), then the probability that it belongs to anomaly data is small;

[0126] if the data point at a certain moment is far from the distribution range of normal data (i.e., is small), then the probability that it belongs to anomaly data is large.

[0127] The distribution mean of normal data is = [0.1, 0.2, 0.3], and the covariance matrix is = I (identity matrix); at time t, = [0.1, 0.2, 0.3], at this time, is larger (because is close to ), so is smaller, then it means that the data at this moment is more likely to be normal data;

[0128] At time t', = [5, 5, 5], at this time is smaller (because is far from ), so is larger, then it means that the data at this moment is more likely to be abnormal data.

[0129] Further, the second detection module 5 includes:

[0130] A position conversion sub-module, used to convert the position information of the final true abnormal data into a unified spatial coordinate system to obtain unified position information data;

[0131] A spatial first processing sub-module, used to analyze the unified position information data according to the spatial autocorrelation algorithm to obtain spatial distribution characteristics;

[0132] Specifically, check whether the data has missing values or outliers. If there are missing values, interpolation or mean filling methods can be used for processing; if there are outliers, the 3σ principle or IQR method can be used to remove or correct them. Organize the position information data into a format suitable for spatial autocorrelation analysis (such as a two-dimensional array or GeoDataFrame) to obtain the cleaned unified position information data; determine the neighborhood relationship of each point according to the position information; construct a spatial weight matrix according to the neighborhood relationship; Global spatial autocorrelation analysis: Use Moran's I or Geary's C algorithm to calculate the global spatial autocorrelation index; Local spatial autocorrelation analysis: Use the local Moran's I (LISA) algorithm to calculate the local spatial autocorrelation index of each point; Spatial distribution feature extraction:

[0133] Summarize global features: According to the global spatial autocorrelation index, summarize the overall spatial distribution characteristics of the data (such as clustering, dispersion, or randomness).

[0134] Extract local features: According to the local spatial autocorrelation index, extract local spatial aggregation regions (such as HH or LL regions) and abnormal regions (such as HL or LH regions), where HH and LL: represent positive spatial autocorrelation, that is, the values in a region are similar to those in its surrounding regions (high-value aggregation or low-value aggregation).

[0135] HL and LH: indicating negative spatial autocorrelation, that is, there are significant differences in values between a region and its surrounding regions (isolated high values or isolated low values).

[0136] Generate a spatial distribution feature report: Integrate global and local features into a spatial distribution feature report, including charts and text descriptions.

[0137] The second spatial processing sub-module is used to classify the spatial distribution features and extract abnormal features by using a spatial clustering algorithm to obtain spatial abnormal data.

[0138] Furthermore, the second spatial processing sub-module includes:

[0139] The clustering unit is used to perform clustering analysis on the spatial distribution features by using a spatial clustering algorithm to obtain spatial aggregation anomalies and spatial isolation anomalies;

[0140] The abnormal feature extraction unit is used to extract abnormal features according to the spatial aggregation anomalies and spatial isolation anomalies to obtain spatial aggregation abnormal features and spatial isolation abnormal features;

[0141] The classification unit is used to classify the spatial aggregation abnormal features and spatial isolation abnormal features to obtain spatial abnormal data.

[0142] Furthermore, the abnormal feature extraction unit includes:

[0143] The aggregation center position determination sub-unit is used to determine the coordinate mean of the data points within the spatial aggregation anomaly and spatial isolation anomaly regions to obtain the aggregation center;

[0144] The aggregation radius determination sub-unit is used to determine the maximum distance from all the data points within the spatial aggregation anomaly and spatial isolation anomaly regions to the coordinate mean to obtain the aggregation radius;

[0145] The aggregation density determination sub-unit is used to determine the aggregation density of all the data points within the spatial aggregation anomaly and spatial isolation anomaly regions to obtain the aggregation density;

[0146] The feature determination sub-unit is used to determine the spatial aggregation abnormal features and spatial isolation abnormal features according to the aggregation center, aggregation radius and aggregation density.

[0147] Specifically, based on the data point coordinates within the spatially aggregated anomaly and spatially isolated anomaly regions, calculate the mean of the data point coordinates (such as longitude and latitude or UTM coordinates) within each anomaly region to obtain the aggregated center coordinates; according to the data point coordinates and their aggregated centers within the spatially aggregated anomaly and spatially isolated anomaly regions, calculate the Euclidean distance from each data point to the aggregated center to obtain the aggregated radius; calculate the area of the anomaly region and determine the aggregation density based on the data point coordinates and their aggregated radii within the spatially aggregated anomaly and spatially isolated anomaly regions; determine the characteristics of spatially aggregated anomalies:

[0148] Aggregated center: Represents the core position of the anomaly region.

[0149] Aggregated radius: Represents the scope size of the anomaly region.

[0150] Aggregation density: Represents the density degree of data points within the anomaly region.

[0151] Determine the characteristics of spatially isolated anomalies:

[0152] Aggregated center: Represents the position of the isolated anomaly point.

[0153] Aggregated radius: Represents the influence range of the isolated anomaly point (usually small).

[0154] Aggregation density: Represents the density of the isolated anomaly point (usually low).

[0155] Generate a feature report: Integrate the characteristics of spatially aggregated anomalies and spatially isolated anomalies into a feature report, including charts and text descriptions.

[0156] Furthermore, the warning level classification of the warning system: Classify the warnings into three levels: low, medium, and high according to the degree of anomaly (such as anomaly probability, influence range, etc.).

[0157] Formulate response strategies: Develop differentiated response strategies for each warning level, such as only recording logs for low-level warnings, notifying relevant personnel for medium-level warnings, and activating emergency response plans for high-level warnings.

[0158] Real-time hierarchical warning: According to the anomaly detection results, trigger corresponding-level warnings in real time and execute corresponding response strategies.

[0159] More specifically, the implementation method of the warning system is:

[0160] Data visualization: Intuitively display anomaly data and warning information through forms such as charts, heat maps, and maps.

[0161] Interactive analysis: Support users to perform interactive analysis on anomaly data, such as filtering, zooming in, and comparison.

[0162] Real-time monitoring: Update abnormal data and warning information in real time to ensure that users can always keep abreast of the latest situation.

[0163] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0164] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A prospecting data anomaly detection and early warning system, characterized in that: include: The data acquisition module is used to collect various prospecting data within a preset time period to obtain raw data; A data fusion module performs data fusion on the original data to obtain a multidimensional data set corresponding to each moment; A data processing module, used to calculate anomaly detection indicators according to the differences of multidimensional data sets at adjacent moments, wherein the anomaly detection indicators include: initial anomaly degree, overall deviation degree and actual anomaly degree; A first detection module, used to determine real abnormal data and fault data according to the distribution of the abnormal detection index; A second detection module is used to determine the spatial anomaly characteristics of the real abnormal data according to a spatial analysis algorithm to obtain spatial anomaly data; An early warning module, used to generate corresponding early warning information according to the spatial abnormal data; The first detection module comprises: A standardization submodule is used to standardize the anomaly detection index to obtain standardized data; A modeling submodule, for performing abnormal index joint distribution modeling on the standardized data using multivariate Gaussian distribution to obtain a probability distribution model; The abnormal probability calculation submodule is used to obtain the abnormal probability at each moment according to the probability distribution model; A cluster analysis submodule, used to cluster the abnormality probabilities at each moment to obtain initial true abnormality data and initial fault data; A result optimization submodule, used for performing secondary classification on the initial true abnormal data and the initial fault data to obtain final true abnormal data and final fault data; The expression of the probability distribution model is: ; in, is the probability distribution model, The standardized data includes: the initial abnormality degree after standardization, the overall deviation degree after standardization and the actual abnormality degree after standardization. ; The second detection module comprises: A position conversion submodule, used to convert the position information of the final real abnormal data into a unified spatial coordinate system to obtain unified position information data; A spatial first processing submodule, used for analyzing the unified position information data according to a spatial autocorrelation algorithm to obtain spatial distribution characteristics; The second spatial processing submodule is used to classify the spatial distribution features and extract abnormal features using a spatial clustering algorithm to obtain spatial abnormal data.

2. A mining data anomaly detection and early warning system according to claim 1, characterized in that: The data fusion module comprises: A data preprocessing submodule is used to perform data cleaning and normalization on the raw data to obtain raw data of the same dimension; A trend calculation submodule is used to decompose the prospecting data of each dimension of the original data using a time series decomposition algorithm to obtain a data change trend sequence; The correlation calculation submodule is used to calculate the Pearson correlation coefficient between the change trend sequences of each dimension data to obtain the multi-dimensional trend correlation; A trend shift sequence calculation submodule is used to calculate the absolute value of the difference between the data sequence of each dimension of the original data and the trend sequence to obtain a trend shift sequence; An abnormal moment detection submodule, used to detect abnormal data in the trend offset sequence using a LOF algorithm to obtain the abnormal moment of the trend; The trend jump degree calculation submodule is used to calculate the trend jump degree according to the distribution of trend abnormal moments; A fusion participation degree calculation module is used to calculate the fusion participation degree according to the trend jump degree and the multi-dimensional trend correlation; A weight calculation submodule, used for calculating the fusion weight according to the fusion participation degree; The generating submodule is used to perform data fusion on the original data according to the fusion weight to obtain a multidimensional data set corresponding to each moment.

3. A prospecting data anomaly detection and early warning system according to claim 1, characterized in that: The data processing module comprises: A first indicator calculation submodule, used to calculate the initial abnormality degree of the multidimensional data set; A second indicator calculation submodule, used for calculating the overall deviation degree according to the initial abnormality degree; The third indicator calculation submodule is used to calculate the actual abnormality degree according to the overall deviation degree.

4. A prospecting data anomaly detection and early warning system according to claim 1, characterized in that: The expression of the initial abnormality degree is: ; in, is the initial abnormality level, For the Dimensions in The observed value at time, For the Dimensions in The sliding standard deviation of the moment, For the Dimensions in The sliding mean of time, n is the total number of dimensions; The calculation formula of the overall deviation degree is: ; in, is the overall deviation degree, For the Dimensions in The predicted value at the moment; The calculation expression of the actual abnormality degree is: ; in, is the actual degree of abnormality, is the local consistency factor, , and are respectively the first weight coefficient, the second weight coefficient and the third weight coefficient; The calculation expression of the local consistency factor is: ; in, is the size of the local time window, is the mean of the standard deviations of all dimensions in the time window, For the The standard deviation of the dimension in the time window, For the A data sequence in a time window.

5. A mining data anomaly detection and early warning system according to claim 1, characterized in that: The abnormal probability calculation expression is: ; in, is the abnormal probability.

6. A prospecting data anomaly detection and early warning system according to claim 1, characterized in that: The second spatial processing submodule comprises: A clustering unit, used to perform cluster analysis on the spatial distribution characteristics using a spatial clustering algorithm to obtain spatial cluster anomalies and spatial isolated anomalies; An anomaly extraction unit is used to extract anomaly features according to the spatially clustered anomalies and the spatially isolated anomalies to obtain spatially clustered anomaly features and spatially isolated anomaly features; The classification unit is used to classify the spatially aggregated abnormal features and the spatially isolated abnormal features to obtain spatial abnormal data.

7. A mining data anomaly detection and early warning system according to claim 6, characterized in that: The abnormality extraction unit comprises: A clustering center position determination subunit, used to determine the coordinate mean of the data points in the spatial clustering anomaly and the spatial isolated anomaly area to obtain the clustering center; A clustering radius determination subunit, used to determine the maximum distance from all data points in the spatial clustering anomaly and spatial isolated anomaly area to the coordinate mean to obtain a clustering radius; A cluster density determination subunit, used to determine the cluster density of all data points in the spatial cluster anomaly and spatial isolated anomaly region to obtain a cluster density; The feature determination subunit is used to determine the spatial clustering anomaly feature and the spatial isolation anomaly feature according to the clustering center, clustering radius and clustering density.

Citation Information

Patent Citations

  • Multi-source data fusion method based on digital twinning

    CN117407744A

  • Exploration data acquisition and logging data anomaly detection method

    CN117574270A

  • Processing method and system based on power transaction decision data

    CN118014615A