Industrial waste gas identification method inspired by visual characteristics
A neural network model leveraging human visual characteristics addresses the limitations of existing waste gas identification methods by enhancing feature extraction and real-time accuracy for multi-type gas recognition in industrial settings.
Patent Information
- Application Number
- CN202510362667.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
The existing vision-based industrial waste gas identification method has poor recognition of multiple industrial waste gases in complex environments, and the model is lightweight and real-time, making it difficult to meet the rapid detection needs of actual industrial scenarios.
An industrial waste gas recognition method inspired by visual characteristics was designed. By constructing a visual characteristic module, fusing visual characteristics and building a feature regression module, the convolutional neural network is used to simulate the visual physiological characteristics of the human eye, combined with the European-style distance constraint convolution kernel size, and improving the feature extraction ability.
It realizes efficient and accurate identification of various industrial waste gases in industrial scenarios with resource constraints, and shows excellent real-time and robustness.
Smart Images

Figure FT_1 
Figure SMS_1 
Figure SMS_2
Abstract
Description
Technical Field
[0001] The present invention belongs to both the field of pollution prevention and control and the field of artificial intelligence. Based on convolutional neural networks and inspired by visual characteristics, an industrial waste gas recognition method is designed, which is applicable to scenarios such as industrial safety monitoring and environmental quality detection. Background Art
[0002] With the acceleration of the industrialization process, industrial waste gas emissions have posed a serious threat to the environment and human health. Traditional industrial waste gas recognition methods mainly rely on chemical sensors and physical detection and recognition devices. Although these methods can identify the components and concentrations of waste gases, they have limitations in terms of real-time performance and accuracy. In recent years, with the development of computer vision and deep learning technologies, industrial waste gas recognition methods based on visual characteristics have gradually become a research hotspot. Convolutional neural networks (CNNs) have achieved great success in the field of image recognition and processing, and their powerful feature extraction capabilities make them an ideal choice for dealing with complex visual tasks. In the field of industrial waste gas recognition, researchers have begun to attempt to use CNNs to analyze and recognize waste gas images. For example, Tao et al. proposed a spatio-temporal network based on channel enhancement (CENet) for identifying industrial smoke emissions. This network significantly improves the accuracy and robustness of smoke detection through multi-layer supervision information and channel enhancement modules. In addition, some studies have combined electronic nose systems and deep learning algorithms, and processed and classified the signals collected by the electronic nose through convolutional neural networks to achieve efficient recognition of industrial waste gases. However, most of the existing vision-based industrial waste gas recognition methods focus on specific scenarios or single-type waste gas detection, and there are still challenges in recognizing multiple industrial waste gases in complex environments. In addition, existing methods still need to be further optimized in terms of model lightweight and real-time performance to meet the rapid detection requirements in actual industrial scenarios. For this reason, the present invention designs an industrial waste gas recognition method inspired by visual characteristics, which is used to efficiently and accurately recognize multiple industrial waste gases, and has important practical significance and broad application prospects. Summary of the Invention
[0003] The present invention designs an industrial waste gas recognition method inspired by visual characteristics. Inspired by visual characteristics, a neural network model for simulating the physiological and psychological reactions of the human eye vision is built, which is realized through three steps: constructing a visual characteristic module, fusing visual features, and building a feature regression module. It can efficiently and accurately recognize multiple industrial waste gases, and shows excellent real-time performance, accuracy, and robustness in resource-constrained industrial scenarios.
[0004] The present invention is realized through the following technical solutions, including the following steps:
[0005] The first step: construct a visual characteristic module;
[0006] Step 2: Integrate visual features;
[0007] Step 3: Build a feature regression module.
[0008] The creativity of the present invention is mainly reflected in:
[0009] Inspired by two characteristics of the human eye visual system (1. The human eye has high sensitivity to horizontal and vertical stimuli; 2. The human eye has center-surround inhibition characteristics), the present invention designs a new visual characteristic module and constrains the size of the convolution kernels in the module through the Euclidean distance. Description of the Drawings
[0010] Figure 1 is the flowchart of the industrial waste gas recognition neural network designed by the present invention inspired by visual characteristics. Detailed Implementation Manner
[0011] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0012] Embodiment:
[0013] Step 1: Build a visual characteristic module;
[0014] Based on two important discoveries in neuroscience (1. The human eye has high sensitivity to horizontal and vertical stimuli; 2. The human eye has center-surround inhibition characteristics), and combining the advantage that simulating the human eye perception behavior can improve the monitoring performance in visual monitoring tasks, the present invention designs a horizontal bar module, a vertical bar module and a square module based on the Euclidean distance and connects them in parallel to construct a visual characteristic module (as Figure 1 shown) to improve the effectiveness of feature extraction. The specific implementation manner is as follows.
[0015] First, construct a square convolution with a convolution kernel of n×n. The output O s of this square convolution can be expressed as:
[0016] O s = Cov s (X)
[0017] where X is the input industrial waste gas image, and Cov s (·) is the convolution operation of the square convolution.
[0018] Secondly, construct a horizontal bar convolution with a convolution kernel of 1×m and a vertical bar convolution with a convolution kernel of m×1, where m = 2k + 1 and k ∈ Ν (Ν is a natural number). The output O hThe output O of the convolution with the vertical bar v can be respectively expressed as:
[0019] O h = Cov h (X)
[0020] and
[0021] O v = Cov v (X)
[0022] where Cov h (·) and Cov v (·) are respectively the convolution operations of the horizontal bar convolution and the vertical bar convolution.
[0023] Finally, it is learned from neuroscience that the stimulated retinal cells of the human eye will produce a strong neuronal response, enhancing the central receptive field and suppressing the surrounding environment. Inspired by this, the present invention uses the Euclidean distance to establish the relationship between n and m, and then determines the kernel sizes of the three convolution kernels to reflect the finiteness of the human eye receptive field. The binary (such as a and b) Euclidean distance can be expressed as:
[0024]
[0025] where a = {a1, a2}, b = {b1, b2}. Considering that the above three convolutions are simultaneously used to simulate the central receptive field of the human eye vision, the maximum distance from the center of each convolution kernel to any other position is the same. Assuming that the centers of the above three convolution kernels are the origin, that is, b1 = b2 = 0, then the relationship between m and n can be deduced as:
[0026]
[0027] where is the ceiling operation. In the present invention, d = 2, n = 3, m = 5.
[0028] Step 2: Fuse visual features;
[0029] The present invention first combines the outputs of the above three convolutions through a concatenation operation, and then performs a feature transformation through a 1×1 convolution to extract the fused visual feature F v :
[0030]
[0031] where is the concatenation operation, and Cov 11 (·) is the 1×1 convolution operation.
[0032] Step 3: Build a feature regression module;
[0033] To obtain accurate recognition results, the present invention builds a feature regression module. This module consists of a global average pooling layer (GAP) and a fully connected layer (FC), which map the above-mentioned fused visual feature F v to the final recognition result.
Claims
1. Construct a visual feature module; Based on two important findings in optic nerve science (1. The human eye has high sensitivity to horizontal and vertical stimuli; 2. The human eye has center-surround inhibition characteristics), and combining the advantage that simulating the human eye's perception behavior can improve the monitoring performance in visual monitoring tasks, the present invention designs a horizontal bar module, a vertical bar module, and a square module based on the Euclidean distance and connects them in parallel to construct a visual feature module to improve the effectiveness of feature extraction. The specific implementation method is as follows. First, construct a square convolution with a convolution kernel of n×n, and the output O of this square convolution s can be expressed as: O s = Cov s (X) Among them, X is the input industrial waste gas image, Cov s (·) is the convolution operation of square convolution. Secondly, construct a horizontal bar convolution with a convolution kernel of 1×m and a vertical bar convolution with a convolution kernel of m×1, where m = 2k + 1 and k ∈ Ν (Ν is a natural number). The output O h of the horizontal bar convolution and the output O v of the vertical bar convolution can be respectively expressed as: O h = Cov h (X) and O v = Cov v (X) Among them, Cov h (·) and Cov v (·) are the convolution operations of horizontal bar convolution and vertical bar convolution, respectively. Finally, it is learned from optic nerve science that the stimulated human eye retinal cells will generate strong neuronal responses, enhancing the central receptive field and suppressing the surrounding environment. Inspired by this, the present invention uses the Euclidean distance to establish the relationship between n and m, and then determines the kernel sizes of the three convolution kernels to reflect the finiteness of the human eye's receptive field. The binary (such as a and b) Euclidean distance can be expressed as: where a = {a1, a2}, b = {b1, b2}. Considering that the above three convolutions are simultaneously used to simulate the central receptive field of the human eye vision, the maximum distance from the center of each convolution kernel to any other position is the same. Assuming that the centers of the above three convolution kernels are the origin, that is, b1 = b2 = 0, then the relationship between m and n can be deduced as: Among them, is the ceiling operation. In the present invention, d = 2, n = 3, and m = 5.
2. Fuse visual features; The present invention first combines the outputs of the above three convolutions through a concatenation operation, and then performs a feature transformation through a 1×1 convolution to extract the fused visual feature F v : Among them is the concatenation operation, Cov 11 (·) is a 1×1 convolution operation.
3. Build a feature regression module; To obtain accurate recognition results, the present invention constructs a feature regression module. This module consists of one global average pooling layer (GAP) and one fully connected layer (FC), and maps the above-mentioned fused visual feature F v to the final recognition result.