A method and device for predicting gas concentration based on adaptive nearest neighbor model

By using an adaptive nearest neighbor model in coal mines, the problem of gas concentration prediction in dynamic environments is solved, and high-precision and stable prediction effects are achieved.

CN115982587BActive Publication Date: 2025-05-16INNER MONGOLIA GUODING INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310036344.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-05-16
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict coal mine gas concentration in a dynamically changing geological environment, resulting in the prediction model being unable to effectively predict in a dynamic environment.

Method used

The gas concentration prediction method based on the adaptive nearest neighbor model is adopted. The initial gas data is obtained by setting monitoring points down the mine, gas data is collected in real time and preprocessed. The training set of the nearest neighbor model is updated using the sample denoising method and sample reuse strategy to achieve high-precision gas concentration prediction in a dynamic environment.

Benefits of technology

High-precision gas concentration prediction in dynamically changing geological environments is achieved, with better fitting of the predicted output with the actual change, better stability, and improved the reliability of the prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115982587B_ABST
    Figure CN115982587B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of coal mine gas, and specifically relates to a method and device for predicting gas concentration based on an adaptive nearest neighbor model; the method obtains initial gas data of a monitoring point, pre-processes the data to obtain gas training data as a training set of a nearest neighbor model; obtains the latest hour's gas data and normalizes the data to obtain the latest hour's gas sample, and adds the data to the training set of the nearest neighbor model; uses a sample denoising method to process the training set of the nearest neighbor model; searches for deleted data based on the latest hour's processed gas data, uses a sample reuse strategy to determine the found deleted data, puts valid deleted data back into the training set of the nearest neighbor model, and retrains the nearest neighbor model based on the current training set; searches for nearest neighbor gas data based on the features of the latest hour's processed gas data, and predicts gas concentration by weighted summation; the present invention is more reliable and advantageous for predicting gas concentration results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer / coal mine gas, and in particular relates to a method and a device for predicting gas concentration based on an adaptive nearest neighbor model. Background Art

[0002] The coal mining industry is the industry with the most serious safety accidents in my country. With the strengthening of my country's supervision and inspection of the coal mining industry, the comprehensive use of cloud computing, artificial intelligence, machine learning and other technologies to conduct detection and early warning of the coal mining industry has become an important solution to ensure safe production in coal mines.

[0003] In the process of informatization construction of coal mine production, a large amount of coal mine safety production data has been accumulated, among which the most important is gas concentration monitoring data. However, there are very few models based on gas concentration prediction that can achieve satisfactory results in actual production environments. This is because coal mine gas mining is a dynamic process. The geological environment, the most critical factor affecting gas, is constantly changing, which causes the distribution of gas data to change continuously, and ultimately leads to the inability of gas concentration prediction models to effectively complete predictions in dynamic environments. Current prediction methods all use static environments to describe changes in gas concentration, and how to predict gas concentration in such a dynamic geological environment is an urgent problem that needs to be solved. Summary of the invention

[0004] To solve the above problems, the present invention provides a method and device for predicting gas concentration based on an adaptive nearest neighbor model. The invention fully considers the dynamic change characteristics of gas distribution and can effectively achieve high-precision gas concentration prediction in a dynamically changing geological environment.

[0005] In a first aspect, the present invention provides a method for predicting gas concentration based on an adaptive nearest neighbor model, comprising the following steps:

[0006] S1. Setting monitoring points in the mine, obtaining initial gas data of the monitoring points, preprocessing the initial gas data to obtain gas training data, and using the gas training data as a training set for the nearest neighbor model for training;

[0007] S2. Real-time collection of gas data within one hour of the monitoring point, and constructing it into a binary sample set including a feature value and a concentration value, wherein the gas data includes the gas recording time and gas concentration value of the monitoring point;

[0008] S3. Obtain the binary sample set of the latest hour, normalize the binary sample set of the latest hour to obtain the gas sample set of the latest hour, and add the gas sample set of the latest hour to the training set of the nearest neighbor model;

[0009] S4. using a sample denoising method to process the training set of the nearest neighbor model, wherein the processed training set of the nearest neighbor model includes the gas sample set of the latest hour;

[0010] S5. Find the deleted data based on the latest one-hour gas sample set in the training set, and use the sample reuse strategy to judge the found deleted data, put the deleted data judged to be valid back into the training set of the nearest neighbor model, and retrain the current nearest neighbor model based on the current training set;

[0011] S6. According to the characteristics of the gas sample set in the latest hour, find K nearest gas data in the training set of the nearest neighbor model and perform weighted sum prediction to obtain the gas concentration to be predicted for one hour; and return to step S3.

[0012] Furthermore, the process of processing the training set of the nearest neighbor model using the sample denoising method in step S4 includes:

[0013] S41. Initialize the old sample set to make it an empty set; and use the latest one-hour gas sample set in the training set as the new sample set;

[0014] S42. Calculate the Euclidean distance between each gas sample in the new sample set and all gas samples in the training set one hour ago, and add the gas samples one hour ago corresponding to the Euclidean distance value less than the preset value d to the old sample set;

[0015] S43. Calculate the maximum difference between the cumulative empirical distribution of the new sample set and the old sample set;

[0016] S44. Given a preset confidence α, if the maximum difference of the cumulative empirical distribution is greater than the preset confidence α, all gas samples before the latest hour in the training set are deleted; if the maximum difference of the cumulative empirical distribution is not greater than the preset confidence α, the current training set is maintained.

[0017] Furthermore, the maximum difference D of the cumulative experience distribution label The calculation formula is:

[0018]

[0019] Among them, Y N (y) represents the cumulative empirical distribution of the gas concentration value in the latest hour, Y(y) represents the cumulative empirical distribution of the gas concentration value one hour ago, and sup represents the upper bound.

[0020] Furthermore, the specific process of step S5 includes:

[0021] S51. Find N nearest deleted data according to the gas sample set of the latest hour, where the deleted data is one of all data deleted by the denoising algorithm;

[0022] S52. Set a weight threshold, calculate the weight of each nearest deleted data, and put the nearest deleted data with a weight greater than the weight threshold back into the training set of the nearest neighbor model.

[0023] Furthermore, the weight calculation formula for each nearest deleted data s=(x, y) is:

[0024]

[0025] Among them, x is the characteristic value of the deleted data, y is the concentration value of the deleted data, N c (s) represents the old sample set, y i N c (s) represents the concentration value of the i-th old sample, w(s,t) represents the weight of the deleted data s at time t, and Δ represents the indicator function.

[0026] Furthermore, the formula for weighted summation according to the distances of K nearest gas data in step S6 is:

[0027]

[0028] in, represents the predicted value, z t Represents the characteristics of the gas sample in the latest hour, z i represents the feature of the ith nearest gas sample, K represents the number of preset nearest gas samples, Represents the concentration value of the i-th nearest gas sample.

[0029] In a second aspect, based on the method of the first aspect, the present invention further provides a gas concentration prediction device based on an adaptive nearest neighbor model, comprising:

[0030] Data storage module, used to store the training set of the nearest neighbor model;

[0031] A real-time data acquisition module is used to acquire the to-be-processed gas data within one hour in real time and construct it into a two-tuple sample set including a feature value and a concentration value, wherein the gas data includes the gas recording time and the gas concentration value of the monitoring point;

[0032] A preprocessing module, used for performing maximum and minimum value normalization on a binary sample set of the to-be-processed gas data to obtain a gas sample set of the to-be-processed gas data, and adding the gas sample set to the data storage module;

[0033] A denoising module, used for performing data deletion processing on the data storage module of the gas sample set to which the gas data to be processed is added according to the sample denoising method;

[0034] An old data reuse module is used to search for deleted data according to a gas sample set of gas data to be processed, and to judge the found deleted data using a sample reuse strategy, and to put the deleted data judged to be valid back into the data storage module;

[0035] The prediction module is used to search for K nearest gas data in the data storage module according to the characteristics of the gas sample set of the gas data to be processed, and to perform weighted sum prediction to obtain the gas concentration to be predicted for one hour.

[0036] Beneficial effects of the present invention:

[0037] The present invention adopts an adaptive nearest neighbor model, and by continuously screening new gas concentration data during the prediction process and updating the training set of the nearest neighbor model, the nearest neighbor model can adapt to the changing geological environment under the coal mine and the changing gas data distribution. The model prediction output of the present invention has a better fit with the actual gas concentration change, and the prediction output has good stability without large fluctuations, that is, the present invention is more reliable and more advantageous in predicting coal mine gas concentration. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The figure is a flow chart of the method for predicting gas concentration based on the adaptive nearest neighbor model of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] The present invention provides a gas concentration prediction method based on an adaptive nearest neighbor model, which is mainly based on an incremental learning machine to train a gas concentration prediction model (nearest neighbor model), obtain gas data in real time and preprocess it, input the preprocessed gas data into the trained gas concentration prediction model, and obtain the gas concentration prediction result within the prediction time.

[0041] In one embodiment, if Figure 1 As shown, the gas concentration prediction method based on the incremental learning machine can be divided into an initialization stage and an incremental stage, including the following steps:

[0042] Initialization phase:

[0043] S1. Setting monitoring points in the mine, obtaining initial gas data of the monitoring points, preprocessing the initial gas data to obtain gas training data, and using the gas training data as a training set of the nearest neighbor model for training.

[0044] Specifically, the location in the working face of the mine that needs to be predicted is called a monitoring point. A gas sensor is set at the monitoring point to collect gas concentration data. The gas concentration data collected by the gas sensor is collected through an existing data acquisition system and stored in a database; the gas sensor can be a methane sensor.

[0045] S2. Collect gas data of the monitoring point within one hour in real time, and construct each data in the gas data into a two-tuple sample including its characteristic value and concentration value, and finally obtain a two-tuple sample set within the hour, wherein the gas data includes the gas recording time and gas concentration value of the monitoring point.

[0046] Incremental phase:

[0047] S3. Obtain the binary sample set of the latest hour, normalize the binary sample set of the latest hour to obtain the gas sample set of the latest hour, and add the sample set to the training set of the nearest neighbor model.

[0048] Specifically, the gas data collected by the gas sensor is collected according to a fixed time granularity, and the gas data is constructed into a binary sample set and then normalized to the maximum and minimum values.

[0049] S4. A sample denoising method is used to process the training set of the nearest neighbor model. The processed training set of the nearest neighbor model includes the gas sample set of the latest hour.

[0050] Specifically, the process of processing the training set of the nearest neighbor model using the sample denoising method in step S4 includes:

[0051] S41. Initialize the old sample set to make it an empty set; and use the latest one-hour gas sample set in the training set as the new sample set;

[0052] S42. Calculate the Euclidean distance value between each gas sample in the new sample set and all gas samples in the training set before the latest hour. For each Euclidean distance value, determine whether it is less than the preset value d. If so, add the gas sample before the latest hour corresponding to the Euclidean distance value to the old sample set. If the gas sample before the latest hour already exists in the old sample set, it does not need to be added again.

[0053] S43. Calculate the maximum difference between the cumulative empirical distribution of the new sample set and the old sample set;

[0054] Specifically, the maximum difference D of the cumulative empirical distribution label The calculation formula is:

[0055] D label =sup|Y N (y)-Y(y)|

[0056] Among them, Y N (y) represents the cumulative empirical distribution of the gas concentration value in the latest hour, Y(y) represents the cumulative empirical distribution of the gas concentration value one hour ago, and sup represents the upper bound.

[0057] S44. Given a preset reliability α, if the maximum difference in the cumulative empirical distribution is greater than the preset reliability α, all gas samples before the latest hour in the training set are deleted; if the maximum difference in the cumulative empirical distribution is not greater than the preset reliability α, the current training set is maintained. In this embodiment, α=0.2.

[0058] S5. Search for deleted data based on the latest one-hour gas sample set in the training set, and use the sample reuse strategy to judge the deleted data found, put the deleted data judged to be valid back into the training set of the nearest neighbor model, and retrain the current nearest neighbor model based on the current training set.

[0059] Specifically, as mining progresses, gas geology changes continuously, and data distribution also changes, that is, gas data may be valid for a period of time and invalid for another period of time, so reusing valid old data can enrich the current sample set and make the trained prediction model more stable and accurate; the specific process of step S5 includes:

[0060] S51. Find N nearest deleted data according to the gas sample set of the latest hour, where the deleted data is one of all data deleted by the denoising algorithm;

[0061] S52. Set a weight threshold, calculate the weight of each nearest deleted data, and put the nearest deleted data with a weight greater than the weight threshold back into the training set of the nearest neighbor model; the weight calculation formula for each nearest deleted data s = (x, y) is:

[0062]

[0063] Among them, x is the characteristic value of the deleted data, y is the concentration value of the deleted data, N c (s) represents the old sample set, y i N c (s) represents the concentration value of the i-th old sample, w(s,t) represents the weight of the deleted data s at time t, and Δ represents the indicator function.

[0064] S6. According to the characteristics of the gas sample set in the latest hour, find the K nearest gas data in the gas training data and perform weighted sum prediction to obtain the gas concentration to be predicted for one hour; if you want to continue the prediction, return to step S3.

[0065] Specifically, the formula for weighted summation according to the distances of K nearest gas data in step S6 is:

[0066]

[0067] in, represents the predicted value, z t Represents the characteristics of the gas sample in the latest hour, z i represents the feature of the ith nearest gas sample (nearest gas data), K represents the number of preset nearest gas samples, Represents the concentration value of the i-th nearest gas sample.

[0068] In one embodiment, the prediction effect of the current prediction model proposed by the present invention is shown in Table 1. Two gas concentration prediction experiments were carried out for 15 consecutive days on two working faces A and B of a mine. The prediction average error (MAE) and the prediction mean error (MSE) were both lower than 0.05, which means that the prediction of the present invention is relatively accurate and has the ability to provide production guidance for coal mines. At the same time, the root mean square error (RMSE) is also relatively small, indicating that the model can achieve relatively stable prediction effects when dealing with high-concentration and low-concentration gas situations, and is suitable for guiding high-concentration gas mines and low-concentration gas mines.

[0069] Table 1 Prediction results of the model at each working face

[0070]

[0071] In one embodiment, the present invention further provides a gas concentration prediction device based on an adaptive nearest neighbor model, comprising:

[0072] Data storage module, used to store the training set of the nearest neighbor model;

[0073] A real-time data acquisition module is used to acquire the to-be-processed gas data within one hour in real time and construct it into a two-tuple sample set including a feature value and a concentration value, wherein the gas data includes the gas recording time and the gas concentration value of the monitoring point;

[0074] A preprocessing module, used for performing maximum and minimum value normalization on a binary sample set of the to-be-processed gas data to obtain a gas sample set of the to-be-processed gas data, and adding the gas sample set to the data storage module;

[0075] A denoising module, used for performing data deletion processing on the data storage module of the gas sample set to which the gas data to be processed is added according to the sample denoising method;

[0076] The old data reuse module is used to search for deleted data according to the gas sample set of the gas data to be processed, and use the sample reuse strategy to judge the found deleted data, and put the deleted data judged to be valid back into the data storage module;

[0077] The prediction module is used to search for K nearest gas data in the data storage module according to the characteristics of the gas sample set of the gas data to be processed, and to perform weighted sum prediction to obtain the gas concentration to be predicted for one hour.

[0078] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.

[0079] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting gas concentration based on an adaptive nearest neighbor model, characterized in that: The following steps are involved: S1. Setting monitoring points in the mine, obtaining initial gas data of the monitoring points, preprocessing the initial gas data to obtain gas training data, and using the gas training data as a training set for the nearest neighbor model for training; S2. Real-time collection of gas data within one hour of the monitoring point, and constructing it into a binary sample set including a feature value and a concentration value, wherein the gas data includes the gas recording time and gas concentration value of the monitoring point; S3. Obtain the binary sample set of the latest hour, normalize the binary sample set of the latest hour to obtain the gas sample set of the latest hour, and add the gas sample set of the latest hour to the training set of the nearest neighbor model; S4. using a sample denoising method to process the training set of the nearest neighbor model, wherein the processed training set of the nearest neighbor model includes the gas sample set of the latest hour; The process of processing the training set of the nearest neighbor model using the sample denoising method in step S4 includes: S41. Initialize the old sample set to make it an empty set; and use the latest one-hour gas sample set in the training set as the new sample set; S42. Calculate the Euclidean distance between each gas sample in the new sample set and all gas samples in the training set one hour ago, and add the gas samples one hour ago corresponding to the Euclidean distance value less than the preset value d to the old sample set; S43. Calculate the maximum difference between the cumulative empirical distribution of the new sample set and the old sample set; S44. Given a preset confidence α, if the maximum difference in the cumulative experience distribution is greater than the preset confidence α, all gas samples before the latest hour in the training set are deleted; if the maximum difference in the cumulative experience distribution is not greater than the preset confidence α, the current training set is maintained; Maximum difference in cumulative experience distribution D label The calculation formula is: D label =sup|Y N (y)-Y(y)| Among them, Y N (y) represents the cumulative empirical distribution of the gas concentration value in the latest hour, Y(y) represents the cumulative empirical distribution of the gas concentration value one hour ago, and sup represents the upper bound; S5. Find the deleted data based on the latest one-hour gas sample set in the training set, and use the sample reuse strategy to judge the found deleted data, put the deleted data judged to be valid back into the training set of the nearest neighbor model, and retrain the current nearest neighbor model based on the current training set; The specific process of step S5 includes: S51. Find N nearest deleted data according to the gas sample set of the latest hour, where the deleted data is all data deleted by the sample denoising method; S52. Set a weight threshold, calculate the weight of each of the nearest deleted data, and put the nearest deleted data with a weight greater than the weight threshold back into the training set of the nearest neighbor model; The weight calculation formula for each nearest deleted data s = (x, y) is: Among them, x is the characteristic value of the deleted data, y is the concentration value of the deleted data, N c (s) represents the old sample set, y i N c (s) represents the concentration value of the i-th old sample, w(s,t) represents the weight of the deleted data s at time t, and Δ represents the indicator function; S6. According to the characteristics of the gas sample set in the latest hour, find the K nearest gas data in the training set of the nearest neighbor model and perform weighted sum prediction to obtain the gas concentration to be predicted in one hour.

2. A method for predicting gas concentration based on an adaptive nearest neighbor model according to claim 1, characterized in that: The formula for weighted summation based on the distances of the K nearest gas data in step S6 is: in, represents the predicted value, z t Represents the characteristics of the gas sample in the latest hour, z i represents the feature of the i-th nearest gas sample, K represents the number of preset nearest gas samples, y i * Represents the concentration value of the i-th nearest gas sample.

3. A gas concentration prediction device for implementing a gas concentration prediction method based on an adaptive nearest neighbor model as described in any one of claims 1 to 2, characterized in that: include: Data storage module, used to store the training set of the nearest neighbor model; A real-time data acquisition module is used to acquire the to-be-processed gas data within one hour in real time and construct it into a two-tuple sample set including a feature value and a concentration value, wherein the gas data includes the gas recording time and the gas concentration value of the monitoring point; A preprocessing module, used for performing maximum and minimum value normalization on a binary sample set of the to-be-processed gas data to obtain a gas sample set of the to-be-processed gas data, and adding the gas sample set to the data storage module; A denoising module, used for performing data deletion processing on the data storage module of the gas sample set to which the gas data to be processed is added according to the sample denoising method; An old data reuse module is used to search for deleted data according to a gas sample set of gas data to be processed, and to judge the found deleted data using a sample reuse strategy, and to put the deleted data judged to be valid back into the data storage module; The prediction module is used to search for K nearest gas data in the data storage module according to the characteristics of the gas sample set of the gas data to be processed, and to perform weighted sum prediction to obtain the gas concentration to be predicted for one hour.

Citation Information

Patent Citations

  • Text abstract extraction method and device, computer equipment and storage medium

    CN110750637A

  • Coal mine gas concentration prediction method combining ensemble learning and weighted extreme learning machine

    CN112712192A