Gas identification method based on deep separable convolutional network

By constructing a deep separable convolution network FRCNN and a hybrid quantization training algorithm, combined with the Winograd convolution engine, the existing gas recognition algorithm has solved the problems of high cost and high power consumption, and achieved low-cost and low-power gas recognition effect.

CN120387064APending Publication Date: 2025-07-29CHONGQING PERKINS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510331278.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing gas recognition algorithms require high-performance hardware support, which are costly and have high power consumption, and cannot meet the low cost and low power consumption requirements of portable electronic noses.

Method used

The gas recognition method based on the deep separable convolution network is adopted. By constructing the depth separable convolution network FRCNN, combining the channel residual allocation module, the depth separable convolution module and the depth aggregation module, the feature extraction and classification are performed, and the hybrid quantization training algorithm and the Winograd convolution engine are used for edge acceleration.

Benefits of technology

On the premise of ensuring the accuracy of gas classification, gas identification with low cost and low power consumption is achieved, and is suitable for edge equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387064A_ABST
    Figure CN120387064A_ABST
Patent Text Reader

Abstract

The invention discloses a gas recognition method based on a deep separable convolutional network. The method comprises the following steps: constructing a gas recognition system based on the deep separable convolutional network; the gas collection module collects gas characteristic data and converts the gas characteristic data into two-dimensional matrix data; obtaining two-dimensional matrix data by an input layer of the depth separable convolutional network; the first feature extraction block performs feature extraction operation on the two-dimensional matrix data; the first maximum pooling layer performs maximum pooling operation on the first feature data; the second feature extraction block performs feature extraction operation on the first pooling data; the second maximum pooling layer performs maximum pooling operation on the second feature data; a third feature extraction block performs feature extraction operation on the second pooling data; a fourth feature extraction block performs feature extraction operation on the third feature data; and the full connection layer performs full connection operation on the fourth characteristic data, and then outputs a gas identification result through a softmax activation function. The method has the effects that the requirements of low cost and low power consumption can be met on the premise of ensuring the gas classification precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gas recognition, and particularly to a gas recognition method based on a depthwise separable convolutional network. Background Art

[0002] An electronic nose generates response signals to different gases through a sensor array. After these signals are preprocessed, they are analyzed and processed through a pattern recognition algorithm to identify the gas type. Currently, electronic noses usually use machine learning algorithms for pattern recognition, such as random forest (RF), k-nearest neighbor (KNN), support vector machine (SVM), decision tree (DT), and multi-layer perceptron neural network (MLPs). However, during the data acquisition, processing, and algorithm prediction processes, the interference of factors such as sensor noise and environmental changes will cause the response baseline value of the electronic nose to drift, resulting in a significant decrease in the recognition accuracy of machine learning algorithms.

[0003] On the one hand, with the development of deep learning, many researchers use deep learning algorithms for gas pattern recognition. Research shows that deep learning algorithms can effectively improve the recognition performance of electronic noses and have more advantages in terms of robustness and generalization ability. However, these algorithms all focus on how to improve the gas recognition accuracy and do not consider the hardware implementation of the electronic nose. In practical applications, complex neural networks with a large number of parameters require high-performance hardware support, including high-end processing and large-capacity memory, etc., which is contradictory to portable electronic noses.

[0004] On the other hand, field programmable gate array (FPGA) is a popular choice in edge artificial intelligence solutions. It has unique advantages such as parallel computing, low power consumption, and high programmability, and can effectively overcome the problems brought by algorithms. However, most artificial intelligence accelerators are implemented based on high-resource FPGAs, and there are differences in computing modes and conflicts in storage requirements between them, which are unbearable for electronic noses that are extremely sensitive to cost and power consumption.

[0005] There are conflicts between traditional deep learning algorithms and edge devices in terms of computational load and computational resources, which severely limit the application of electronic noses in edge devices. Facing these problems, previous research mainly optimized the algorithms by introducing model compression methods. However, there are still two challenges, namely redundant parameters and high power consumption.

[0006] Disadvantages of the prior art: Existing gas recognition algorithms require high-performance hardware support, with high costs and high power consumption, which are unbearable for electronic noses that are extremely sensitive to cost and power consumption. Summary of the Invention

[0007] A gas recognition method based on a depthwise separable convolutional network provided by the present invention can meet the requirements of low cost and low power consumption on the premise of ensuring the gas classification accuracy.

[0008] To achieve the above object, a key aspect of a gas recognition method based on a depthwise separable convolutional network provided by the present invention includes the following steps:

[0009] Step 1: Construct a gas recognition system based on a depthwise separable convolutional network. The gas recognition system is provided with a gas collection module and a depthwise separable convolutional network FRCNN connected in sequence. The depthwise separable convolutional network FRCNN is provided with an input layer, a first feature extraction block, a first max-pooling layer, a second feature extraction block, a second max-pooling layer, a third feature extraction block, a fourth feature extraction block, a fully connected layer, and an output layer connected in sequence. The output end of the fully connected layer is built-in with a softmax activation function;

[0010] The first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block have the same structure and the same working logic, and are all provided with a channel residual allocation module CRA, a depthwise separable convolution module DSConv, a depth aggregation module DA, and a multiplication unit. The depthwise separable convolution module DSConv is provided with a depthwise convolution layer and a pointwise convolution layer connected in sequence;

[0011] Step 2: The gas collection module collects gas feature data a with time series attributes in real time, converts it into two-dimensional matrix data b, and transmits it to the depthwise separable convolutional network FRCNN;

[0012] Step 3: The input layer of the depthwise separable convolutional network FRCNN obtains the two-dimensional matrix data b and transmits it to the first feature extraction block;

[0013] Step 4: The first feature extraction block performs feature extraction operations on the two-dimensional matrix data b to obtain first feature data c and transmits it to the first max-pooling layer;

[0014] Step 5: The first max-pooling layer performs max-pooling operations on the first feature data c to obtain first pooled data d and transmits it to the second feature extraction block;

[0015] Step 6: The second feature extraction block performs feature extraction operations on the first pooled data d to obtain second feature data e and transmits it to the second max-pooling layer;

[0016] Step 7: The second max-pooling layer performs max-pooling operations on the second feature data e to obtain second pooled data f and transmits it to the third feature extraction block;

[0017] Step 8: The third feature extraction block performs a feature extraction operation on the second pooled data f to obtain third feature data g, and transfers it to the fourth feature extraction block;

[0018] Step 9: The fourth feature extraction block performs a feature extraction operation on the third feature data g to obtain fourth feature data h, and transfers it to the fully connected layer;

[0019] Step 10: The fully connected layer performs a fully connected operation on the fourth feature data h, and then outputs the gas recognition result through the softmax activation function.

[0020] The overall structure of the depthwise separable convolutional network FRCNN mainly consists of three parts: the channel residual distribution module, the depthwise separable convolution module, and the depth aggregation module.

[0021] First, the input two-dimensional data map enters the depthwise separable convolution module to extract image features. Since the depthwise separable convolution decomposes the traditional convolution into depth convolution and pointwise convolution, it not only greatly reduces the amount of computation and the number of parameters, but also can achieve efficient feature extraction on devices with limited resources, meeting the requirements of low cost and low power consumption. In addition, it can focus on learning the specific features of each channel and the cross-channel correlation.

[0022] Subsequently, the data enters the depth aggregation module to perform four iterative accumulations on the feature maps of the previous layer.

[0023] Subsequently, the channel residual distribution module is used to generate weighted feature maps, continuously enhancing the feature expression ability. This weighted operation can be adjusted according to the importance of different features, making the final feature map more accurately reflect the essential features of the image, and effectively improving the gas classification progress.

[0024] Finally, the Softmax function is used for feature classification to achieve the gas recognition effect.

[0025] Preferably: The i-th feature extraction block performs a feature extraction operation on its input feature X-i, and the process is as follows:

[0026] S1: The depth convolution layer in the depthwise separable convolution module DSConv uses a depth convolution kernel of size 3×3×1 to perform convolution operations on the data of each input channel in the input feature x to obtain the i-th depth convolution data x1-i, and transfers it to the pointwise convolution layer;

[0027] The pointwise convolution layer uses a 1×1 convolution kernel to perform pointwise convolution operations on the i-th depth convolution data x1-i, combines the i-th depth convolution data x1-i in the channel dimension, completes the information fusion between different channels, obtains the i-th pointwise convolution data x2-i, and transfers it to the depth aggregation module DA;

[0028] S2: The depth aggregation module DA performs four iterative accumulations on the i-th pointwise convolution data x2-i to obtain the i-th depth aggregation data x3-i, and transmits it to the multiplication unit;

[0029] S3: The channel residual allocation module CRA uses a 1×1 convolutional layer to weight the input feature X mapping in the channel dimension, then performs a non-linear mapping through the ReLU activation function to generate an attention map, automatically learns the importance weights of different channels, generates the i-th attention weight w-i, and transmits it to the multiplication unit;

[0030] S4: The multiplication unit multiplies the i-th depth aggregation data x3-i and the i-th attention weight w-i, and outputs the i-th feature data, where i ∈ [1, 4].

[0031] The input feature X-i is either two-dimensional matrix data b, or first pooling data d, or second pooling data f, or third feature data g.

[0032] The channel residual allocation block CRA uses a 1x1 convolutional layer to weight the input feature mapping in the channel dimension. A non-linear mapping is performed through the activation function to generate an attention map, and the importance weights of different channels are automatically learned. The generated attention weights reflect the importance of different positions or channels in the residual connection. The residual connection multiplies the input feature by the feature after the channel dimension weighted operation to form a new feature representation, reducing the risk of model gradient disappearance.

[0033] The depthwise separable convolution block DSConv consists of two parts: depthwise convolution and pointwise convolution. Depthwise convolution: Use a depthwise convolution kernel of size 3x3x1 to perform convolution operations on each input channel respectively. Each convolution kernel is only responsible for one input channel and does not perform channel mixing. Pointwise convolution: Use a 1x1 convolution kernel to perform pointwise convolution operations to combine the outputs of the depthwise convolution in the channel dimension to achieve information fusion between different channels. Due to the small number of parameters, the depthwise separable convolution block makes the model lighter and easier to deploy on devices with limited resources.

[0034] The depth aggregation block DA is a powerful feature fuser. It performs repeated iterative accumulation operations on the feature maps of the first four layers. Each iterative summation is a depth fusion of features from different layers. The multi-layer feature maps are fused with the feature map of the current layer, continuously enriching the feature information and ensuring that important features are not missed.

[0035] For the entire network structure of FRCNN, except that the number of channels in the last depthwise separable convolution block is reduced to 2, the number of channels output by other depthwise convolution blocks is 4. Before performing the 3x3 depthwise-separable convolution, zero-padding is first performed with a padding value of 1, and the stride of all convolutional layers is set to 1 to ensure that the sizes of the input and output images remain unchanged. For the channel residual allocation block, a 1x1 convolutional layer is used to perform a weighted operation on the input feature map in the channel dimension, and then a non-linear mapping is generated through an activation function to produce an attention map of the same size as the feature map. The non-linear activation function Rel is selected for all convolutional layers, and the maximum pooling layer is set to 2.

[0036] In FRCNN, the 1x1 convolution that performs the pseudo-attention mechanism operation is essentially a linear transformation with a computational complexity of O(n) level, which only requires convolution and element-wise multiplication. Compared with traditional self-attention mechanisms, such as the Q-K-V dot product calculation in transformers, which calculates the correlations between positions in the input sequence, especially for long sequences, its complexity is O(n 2 ) level, and a large number of matrix operations and normalization operations are required. In practical applications, the simplicity and regularity of the calculation process of this algorithm enable it to meet the deployment requirements of FPGAs while maintaining a low computational complexity and memory footprint.

[0037] Preferably: The gas collection module is a gas converter array, which is composed of 15 different metal oxide semiconductor gas sensors, an analog-to-digital converter, and peripheral circuits.

[0038] The gas reacts with the metal oxide sensor to generate an analog signal. Subsequently, the analog signal is converted into a digital signal through an analog-to-digital converter.

[0039] Preferably: The depthwise separable convolutional network FRCNN is quantized and trained using a hybrid quantization training algorithm, and its training process is as follows:

[0040] First, the initial learning rate of the depthwise separable convolutional network FRCNN model is set to 0.01, the number of training iterations is set to 1000, the cross-entropy loss function is used to calculate the gradient, and the Adaptive Moment Estimation AdamW optimizer with weight decay is used to update the model parameters; subsequently, training is performed using input data with quantization bits of 8 bits, 16 bits, and 32 bits;

[0041] The QAT algorithm is used to quantize and dequantize the weights and activation values of the algorithm during the training process to simulate the quantization effect during the inference process; at the same time, the training learning rate is adjusted to make the algorithm converge to the global optimal solution.

[0042] The original gas data obtained by the analog-to-digital converter (ADC) device and the weight activation value parameters saved during the subsequent algorithm training process are both in the float32 format. This high-precision format brings a heavy burden to the storage space and computing resources of edge devices. Quantization can effectively solve this problem. However, converting floating-point numbers to fixed-point numbers or integers will result in information loss and reduced precision. The hybrid quantization training algorithm designed in the present invention is based on the QAT algorithm, which pre-quantizes the input data to adapt to the quantization process and combines a learning scheduler to optimize the learning rate.

[0043] Preferably: At the beginning of training, the weights and activation values of the algorithm are initialized using floating-point numbers, and then the weights and activation values are quantized to 8-bit integers. The quantization formula is as follows:

[0044]

[0045] where X is the floating-point number to be quantized, X q is the quantized integer, S is the scale factor, Z is the zero point, and round is the rounding operation;

[0046] The expression for the scale factor S is as follows:

[0047]

[0048] where, and represent the minimum and maximum values of int8; for example, for 8-bit integers, the range is usually [-128, 127].

[0049] After forward propagation, the quantized values are de-quantized back to the floating-point representation, and this process is expressed as:

[0050] X = S(X q - Z)

[0051] After calculating the loss using the de-quantized weights and activation values, backpropagation is performed to calculate the gradients, and then the weights and activation values in the floating-point representation are updated.

[0052] After every 10 training sessions, a learning rate scheduler is used to adjust the learning rate, and the quantization, de-quantization, forward propagation, backpropagation, and parameter update steps are repeated until the training converges.

[0053] At this time, during multiple quantization simulation training processes, the neural network has gradually adapted to the int8 representation and can still maintain a high precision after quantization. Therefore, finally, the weights and activation values are represented as int8.

[0054] Preferably, the gas recognition system based on the depthwise separable convolutional network is further provided with an electronic nose edge device, and the electronic nose edge device is provided with a processing system PS and a programmable logic module PL;

[0055] The processing system PS is provided with an ARM processor, a DDR3 random access memory, and a UART serial port; the programmable logic module PL is provided with a hardware accelerator, an on-chip DRAM memory, a DSP computing unit, and an AXI bus;

[0056] After the ARM processor initializes the system, it reads the gas two-dimensional matrix data collected by the gas collection module from the DDR3 random access memory, then uses the AXI bus to complete the data interaction between the processing system PS and the programmable logic module PL, and transmits the collected gas two-dimensional matrix data to the hardware accelerator of the programmable logic module PL. The reconfigurable hardware accelerator uses the Winograd convolution engine to perform edge acceleration on the depthwise separable convolutional network FRCNN. Finally, the calculation result is transmitted back to the processing system PS through the AXI bus, and then the recognition result is output to the serial port through the UART controller for real-time monitoring on the computer.

[0057] The electronic nose edge device combines the ARM processor "ARM Cortex-A9" at the PS end with the hardware accelerator "XC7Z020" ZYNQ at the PL end to achieve cooperative acceleration.

[0058] Preferably, the Winograd convolution engine is an algorithm for accelerating the calculation of convolutional neural networks. It decomposes the convolution process into a series of matrix multiplication operations to reduce the number of multiplications in the convolution; the Winograd algorithm has different calculation processes for different sizes of convolution kernels.

[0059] The Winograd convolution engine nests the one-dimensional algorithm F(m, r) with itself to obtain the two-dimensional algorithm F(m×m, r×r) for representing two-dimensional convolution, expressed as:

[0060] Y = A T [[GgG T ⊙ [B T dB]]A

[0061] Among them, the symbol "⊙" represents element-wise multiplication, g is a filter of size r×r, and d is an image block of size (m + r - 1)×(m + r - 1);

[0062] There are r - 1 overlapping elements between adjacent blocks, and the boundary blocks are filled with zeros. For F(2×2, 3×3), its transformation matrix is:

[0063]

[0064] The Winograd convolution engine is used to accelerate the depth convolution kernel of size 3×3 in the depth convolution layer using the Winograd algorithm of F(2×2, 3×3).

[0065] Preferably, a cyclic tiling strategy is adopted to enhance the parallel processing ability of the Winograd convolution engine.

[0066] The data transmission delay problem is the most important factor affecting the model calculation time. However, the on-chip storage resources are often insufficient to store the parameters and caches of the entire network. Therefore, a cyclic tiling strategy is adopted to solve this problem.

[0067] Preferably, the ARM processor is specifically an ARM Cortex-A9 chip, and the hardware accelerator is specifically an XC7Z020 ZYNQ hardware acceleration chip.

[0068] The beneficial effects of the present invention: A lightweight and efficient focus residual convolution FRCNN based on a depthwise separable convolution framework is proposed to reduce parameters and balance the computational efficiency and accuracy of traditional depthwise separable convolutions. In addition, in order to implement a low-power and low-resource edge-type electronic nose, a pipeline-form reconfigurable convolution acceleration engine is constructed based on the Winograd algorithm, and a hybrid quantization training method is designed to reduce the FPGA deployment resources to about one-fourth of the original. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a flowchart of the present invention;

[0070] Figure 2 It is a distribution diagram of the response data of the ten industrial harmful gas sensor arrays in the embodiment;

[0071] Figure 3 It is a schematic diagram of the structure of the depthwise separable convolution network FRCNN in the embodiment;

[0072] Figure 4 It is a block diagram of the structure of the electronic nose edge device in the embodiment;

[0073] Figure 5 It is a schematic diagram of the Winograd convolution engine algorithm in the embodiment;

[0074] Figure 6 It is a schematic diagram of the row buffer structure in the embodiment;

[0075] Figure 7 It is a graph of the evaluation indexes (a) accuracy and (b) mean absolute error MAE of FRCNN in the embodiment;

[0076] Figure 8 Quantitative heat maps with different precisions in the embodiments;

[0077] Figure 9 It is a training result graph of the mixed-precision quantization algorithm in the embodiments. Specific implementation manners

[0078] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following embodiments or drawings are used to illustrate the present invention, but not to limit the scope of the present invention.

[0079] As Figure 1 shown: A gas recognition method based on a depthwise separable convolutional network includes the following steps:

[0080] Step 1: Construct a gas recognition system based on a depthwise separable convolutional network. The gas recognition system is provided with a gas collection module and a depthwise separable convolutional network FRCNN connected in sequence. As Figure 3 shown, the depthwise separable convolutional network FRCNN is provided with an input layer, a first feature extraction block, a first max pooling layer, a second feature extraction block, a second max pooling layer, a third feature extraction block, a fourth feature extraction block, a fully connected layer, and an output layer connected in sequence. The output end of the fully connected layer is internally provided with a softmax activation function;

[0081] Step 2: The gas collection module collects gas feature data a with time series attributes in real time, converts it into two-dimensional matrix data b, and transmits it to the depthwise separable convolutional network FRCNN;

[0082] Step 3: The input layer of the depthwise separable convolutional network FRCNN obtains the two-dimensional matrix data b and transmits it to the first feature extraction block;

[0083] Step 4: The first feature extraction block performs a feature extraction operation on the two-dimensional matrix data b to obtain first feature data c and transmits it to the first max pooling layer;

[0084] Step 5: The first max pooling layer performs a max pooling operation on the first feature data c to obtain first pooled data d and transmits it to the second feature extraction block;

[0085] Step 6: The second feature extraction block performs a feature extraction operation on the first pooled data d to obtain second feature data e and transmits it to the second max pooling layer;

[0086] Step 7: The second max pooling layer performs a max pooling operation on the second feature data e to obtain second pooled data f and transmits it to the third feature extraction block;

[0087] Step 8: The third feature extraction block performs feature extraction operations on the second pooled data f to obtain third feature data g, and transfers it to the fourth feature extraction block;

[0088] Step 9: The fourth feature extraction block performs feature extraction operations on the third feature data g to obtain fourth feature data h, and transfers it to the fully connected layer;

[0089] Step 10: The fully connected layer performs a fully connected operation on the fourth feature data h, and then outputs the gas recognition result through the softmax activation function.

[0090] The first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block have the same structure, and are all provided with a channel residual allocation module CRA, a depthwise separable convolution module DSConv, a depth aggregation module DA, and a multiplication unit. The depthwise separable convolution module DSConv is provided with a depthwise convolution layer and a pointwise convolution layer connected in sequence.

[0091] The first feature extraction block performs feature extraction operations on the two-dimensional matrix data b, and the process is as follows:

[0092] S1: The depthwise convolution layer in the depthwise separable convolution module DSConv uses a depthwise convolution kernel of size 3×3×1 to perform convolution operations on the input channel data in each of the two-dimensional matrix data b to obtain first depthwise convolution data x1-1, and transfers it to the pointwise convolution layer;

[0093] The pointwise convolution layer uses a 1×1 convolution kernel to perform pointwise convolution operations on the first depthwise convolution data x1-1, combines the first depthwise convolution data x1-1 in the channel dimension, completes information fusion between different channels, obtains first pointwise convolution data x2-1, and transfers it to the depth aggregation module DA;

[0094] S2: The depth aggregation module DA performs four iterative accumulations on the first pointwise convolution data x2-1 to obtain first depth aggregation data x3-1, and transfers it to the multiplication unit;

[0095] S3: The channel residual allocation module CRA uses a 1×1 convolution layer to weight the mapping of the two-dimensional matrix data b in the channel dimension, then performs a non-linear mapping through the ReLU activation function to generate an attention map, automatically learns the importance weights of different channels, generates first attention weights w-1, and transfers it to the multiplication unit;

[0096] S4: The multiplication unit multiplies the first depth aggregation data x3-1 and the first attention weights w-1, and outputs first feature data c, i∈[1,4].

[0097] The depthwise separable convolutional network FRCNN is quantized and trained using a hybrid quantization training algorithm, and its training process is as follows:

[0098] First, set the initial learning rate of the depthwise separable convolutional network FRCNN model to 0.01, set the number of training iterations to 1000, calculate the gradient using the cross-entropy loss function, and update the model parameters using the Adaptive Moment Estimation AdamW optimizer with weight decay; subsequently, use input data with quantization bits of 8 bits, 16 bits, and 32 bits for training;

[0099] The QAT algorithm is used to quantize and dequantize the weights and activation values of the algorithm during training to simulate the quantization effect during the inference process; at the same time, adjust the training learning rate to make the algorithm converge to the global optimal solution.

[0100] At the beginning of training, initialize the weights and activation values of the algorithm using floating-point numbers, and then quantize the weights and activation values to 8-bit integers. The quantization formula is as follows:

[0101]

[0102] where X is the floating-point number to be quantized, X q is the quantized integer, S is the scale factor, Z is the zero point, and round is the rounding operation;

[0103] The expression for the scale factor S is as follows:

[0104]

[0105] where, and represent the minimum and maximum values of int8; for example, for 8-bit integers, the range is usually [-128, 127].

[0106] After forward propagation, the quantized values are dequantized back to floating-point representation, and this process is expressed as:

[0107] X = S(X q - Z)

[0108] After calculating the loss using the dequantized weights and activation values, perform backpropagation to calculate the gradient, and then update the weights and activation values in floating-point representation.

[0109] After every 10 training sessions, use the learning rate scheduler to adjust the learning rate, and repeat the steps of quantization, dequantization, forward propagation, backpropagation, and parameter update until the training converges.

[0110] At this time, during multiple quantization simulation training processes, the neural network has gradually adapted to the int8 representation and can still maintain high precision after quantization. Therefore, the weights and activation values are finally represented as int8.

[0111] The gas recognition system based on the depthwise separable convolutional network is also provided with an electronic nose edge device. The electronic nose edge device is provided with a processing system PS and a programmable logic module PL.

[0112] The structure of the electronic nose edge device is as Figure 4 shown. The processing system PS is provided with an ARM processor, a DDR3 random access memory, and a UART serial port. The programmable logic module PL is provided with a hardware accelerator, an on-chip DRAM memory, a DSP computing unit, and an AXI bus.

[0113] After the ARM processor initializes the system, it reads the gas two-dimensional matrix data collected by the gas collection module from the DDR3 random access memory, then uses the AXI bus to complete the data interaction between the processing system PS and the programmable logic module PL, and transmits the collected gas two-dimensional matrix data to the hardware accelerator of the programmable logic module PL. The reconfigurable hardware accelerator uses the Winograd convolution engine to perform edge acceleration on the depthwise separable convolutional network FRCNN. Finally, the calculation result is sent back to the processing system PS through the AXI bus, and the recognition result is output to the serial port through the UART controller for real-time monitoring on the computer.

[0114] The Winograd convolution engine is an algorithm used to accelerate the calculation of convolutional neural networks. It decomposes the convolution process into a series of matrix multiplication operations to reduce the number of multiplications in the convolution. The Winograd algorithm has different calculation processes for different sizes of convolution kernels.

[0115] The Winograd convolution engine nests the one-dimensional algorithm F(m,r) with itself to obtain the two-dimensional algorithm F(m×m,r×r) used to represent two-dimensional convolution, expressed as:

[0116] Y = A T [[GgG T ⊙ [B T dB]]A

[0117] where the symbol "⊙" represents element-wise multiplication, g is a filter of size r×r, and d is an image block of size (m + r - 1)×(m + r - 1).

[0118] There are r - 1 overlapping elements between adjacent blocks, and zero padding is used for the boundary blocks. For F(2×2,3×3), its transformation matrix is:

[0119]

[0120] The Winograd convolution engine is used to accelerate the 3×3 depth convolution kernel in the depth convolution layer using the Winograd algorithm of F(2×2, 3×3).

[0121] The Winograd convolution engine stores the 4×4 input buffer g in the FIFO data buffer, which follows the principle of first in first out. And uses the transformation matrix G to transform g to obtain a new matrix g'. At the same time, the convolution buffer d is stored in another FIFO data buffer, and d is transformed using the transformation matrix B to obtain a new matrix d'. These two transformation steps are the core of the Winograd algorithm. They improve the computational efficiency of the engine by reducing the number of multiplications and enhancing data reuse. The transformed matrices g' and d' are both in the form of 4×4. The calculation uses 16 DSPs. The matrix g' is multiplied point by point with the matrix d' to obtain the output matrix Y'. Finally, the output matrix Y' is inverse-transformed to obtain a 2×2 matrix Y. The standard algorithm uses 2×2×3×3 = 36 multiplications, and F(2×2, 3×3) uses 4×4 = 16 multiplications, reducing the algorithm complexity by 2.25 times. It is very suitable for FPGA chips with scarce DSP resources. The structural design of the Winograd convolution engine is as Figure 5 shown.

[0122] The data transfer delay problem is the most important factor affecting the model calculation time. However, the on-chip storage resources are often insufficient to store the parameters and caches of the entire network. Therefore, a cyclic tiling strategy is adopted to solve this problem.

[0123] In the cyclic tiling strategy, in order to enhance the parallel processing ability of the engine, the N-channel input feature map is divided into M×4×4 image blocks d every 4 rows, and the missing elements are filled with 0. These are tiled into the FIFO data buffer in a first in first out manner. Each input buffer row contains M×4×4 elements, and there are a total of M + N rows. M rows are used to buffer the read data, and N rows are used to directly output the data to the Winograd convolution engine for calculation. When outputting data, the data is read into the m-row buffer. At the same time, in order to reuse the data of the convolution kernel, the convolution kernel of size 3×3×N×L is cached in the on-chip memory and repeatedly input to the Winograd convolution engine for calculation every clock cycle. The N rows of results calculated by the Winograd convolution engine are stored in the FIFO data buffer. Each input buffer row contains 2×M×2×M elements until all M rows of input data are successfully calculated. The row buffer structure is as Figure 6 shown, Figure 6In the text, WCE represents the Winograd convolution engine. Since the data cached each time is adjacent feature map data of size 2×2, a separate maximum pooling calculation engine is no longer designed. Instead, the calculation of the maximum pooling layer is directly embedded, and caching and calculation are executed in parallel in a pipelined form, improving the throughput of the system.

[0124] In this embodiment, the ARM processor uses an ARM Cortex-A9 chip, and the hardware accelerator uses an XC7Z020 ZYNQ hardware acceleration chip.

[0125] Next, the technical solution of the present invention is verified through specific experiments:

[0126] In this embodiment, a comprehensive gas data acquisition system is built, which includes a gas supply unit and a sensing and conversion unit. The gas supply unit includes ten gas cylinders, namely sulfur dioxide (SO2), nitrogen dioxide (NO2), nitrous oxide (NO), ammonia (NH3), hydrogen chloride (HCl), hydrogen sulfide (H2S), hydrogen (H2), carbon monoxide (CO), methyl mercaptan (CH4S), and acetaldehyde (C2H4O). The sensing and conversion unit consists of a mass flow controller (MFC), an air chamber, and a gas converter array. The gas converter array consists of 15 different metal oxide semiconductor gas sensors, an analog-to-digital converter, and peripheral circuits.

[0127] The operation steps are as follows: First, open the gas cylinder valves of the gas supply unit. The gas is controlled by the MFC and quantitatively transported to the air chamber of the sensing and conversion unit. In the air chamber, the gas reacts with the metal oxide sensors to generate an analog signal. Subsequently, the analog signal is converted into a digital signal through the analog-to-digital converter. These digital signals are transmitted to a computer for data acquisition and processing.

[0128] In the experiment, the instantaneous response data of 15 sensors to a gas is defined as a gas response data unit. Figure 2 The response data distribution of the sensor array for ten industrial harmful gases is described. In one-dimensional data, the data points are usually linearly arranged. On the contrary, in two-dimensional data, the data points can usually be organized in the form of a matrix or a table. In this arrangement, each row or each column can represent a different feature or variable. This organization method can better capture the correlation between data.

[0129] To obtain a more comprehensive gas response data unit, the acquisition frequency is set to 10 Hz, and the acquisition time for each time is set to 1.5 seconds. In one experiment, 15 gas response data units with different concentrations can be collected. Through the two-dimensional matrix mapping of the time series, these data are converted into a 15×15 two-dimensional matrix diagram. A total of 6000 two-dimensional matrix diagram data containing ten industrial harmful gases are collected, providing data support for the training and verification of the gas detection algorithm.

[0130] The experimental platform used for the training of the present invention is as follows: the CPU is i913900HX, and the GPU is NVIDIA GeForce RTX4060. First, the data set is divided into a training set, a validation set, and a test set. 4800 of the 6000 industrial gas two-dimensional matrix diagrams collected are used as the training set, 600 as the validation set, and 600 as the test set. Then, a gas recognition system based on a depthwise separable convolutional network is constructed.

[0131] For this classification task, the cross-entropy loss function is used to measure the difference between the probability distribution predicted by the model and the true label. Then, a specific initial learning rate of 0.01 is set, and an AdamW optimizer is created. Since AdamW is an adaptive optimization algorithm that combines the Adam algorithm with weight decay, the AdamW optimizer is used to optimize the parameters of the model to minimize the loss function during the training process.

[0132] At the beginning of the experiment, the images in the training set are input into the neural network, and the prediction results are calculated through forward propagation. Then, the loss value between the prediction result and the true label is calculated according to the loss function. Then, the gradients of each parameter are calculated according to the loss value through the backpropagation algorithm. Finally, the optimization algorithm is used to update the weights and biases of the network according to the gradients until the loss value converges or reaches the predetermined number of training epochs. Based on multiple experiments, the number of training epochs is set to 1000, and the entire training set is traversed in each epoch.

[0133] To compare the energy acceleration effect of this algorithm on different platforms, the following experimental deployments are carried out. First, the proposed network algorithm is deployed on three mainstream hardware acceleration platforms: CPU, GPU, and FPGA.

[0134] On the CPU platform, the Intel Core i7-8700 processor with 16GB RAM is selected as the CPU experimental environment, which has high computing performance and stability and can provide certain guarantee for the operation of the algorithm. The Imtel Core i7-8700 processor has high computing performance and stability and can provide a certain basic guarantee for the operation of the algorithm.

[0135] For the GPU platform, NVIDIA Jetson Xavier NX is selected as the experimental environment. NVIDIA Jetson Xavier NX has powerful graphics processing capabilities and parallel computing capabilities, making it very suitable for tasks that require a large amount of computing such as deep learning. The algorithm efficiency on the GPU can be further improved through the in-depth optimization of the CUDA and cuDNN libraries. The CUDA and cuDNN libraries are NVIDIA's parallel computing platforms and programming models that allow developers to utilize the parallel computing capabilities of the GPU to perform efficient computations. cuDNN is an acceleration library provided by NVIDIA for deep learning, which can accelerate operations such as convolution and pooling in deep learning algorithms.

[0136] On the FPGA platform, the Zyq XC7Z020 FPGA development board with a clock frequency of 150 MHz is selected as the experimental platform. The network algorithm is mapped to the FPGA through the advanced synthesis tool Xilinx Vivado HS to achieve customized hardware acceleration. FPGAs are programmable and low-power, allowing for customized designs according to specific application requirements. Using Xilinx Vivado HS, algorithms described in high-level languages can be converted into hardware description languages and then implemented on the FPGA. This can make full use of the parallel computing capabilities and flexibility of the FPGA to improve the computing efficiency and energy efficiency of the algorithm.

[0137] The purpose of this experiment is to comprehensively evaluate the performance of the proposed algorithm. The algorithm is tested on the collected industrial gas dataset and public dataset. To verify the effectiveness of the collected gas sensor response data, the public dataset "Chemical gas sensor array dataset" is used. A set of 16 metal oxide gas sensors are exposed to 6 different concentrations of volatile organic compounds, and the dynamic response of each sensor is recorded at a sampling rate of 100 Hz. The dynamic response of each sensor is recorded at a sampling rate of 100 Hz, and each measurement generates a 16-channel time series sequence. Therefore, this dataset is mapped into a 16x16 two-dimensional matrix. At the same time, the processed public data is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. In this experiment, MAE and Accuracy are selected as the main evaluation metrics to comprehensively measure the performance of the algorithm on different datasets. The experimental results are as Figure 7 shown.

[0138] Accuracy, as an important evaluation metric to explain the performance of the model, intuitively reflects the correctness of the algorithm prediction. Experiments are conducted on the collected industrial gas dataset and public dataset. In this experiment, the abscissa represents 1000 epochs during the training process, and the ordinate represents the accuracy, as shown in Figure 7As shown in a. At the beginning of training, the accuracy is relatively low. As training progresses, the accuracy increases rapidly. After a certain number of training times, the accuracy can reach about 95%. When approaching 1000 times, the accuracy tends to stabilize at about 99%.

[0139] The mean absolute error (MAE) is another important evaluation metric. It represents the average of the absolute errors between the algorithm's predicted values and the true values. The smaller the MAE value, the closer the algorithm's prediction results are to the true values, and the higher the accuracy. As Figure 7 shown in b, in the collected industrial gas dataset, after 1000 iterations and optimizations, the MAE value of the algorithm is 0.0049. In this dataset, the algorithm can approximate the true values more accurately, providing a reliable reference for decision-making in practical applications. When conducting experiments on the public dataset, it can be found that the MAE value of the algorithm on this dataset is 0.0064, indicating that the prediction deviation of the algorithm is small, and it can effectively minimize the impact of errors, further verifying the effectiveness and reliability of the algorithm.

[0140] Table 1 Performance comparison of different gas opening models

[0141]

[0142] To fully demonstrate the performance of the depthwise separable convolutional network FRCNN in industrial harmful gas detection, comparative experiments were conducted with public gas detection networks such as the classical convolutional neural network VGG-16, the 18-layer Residual Network ResNet18, the deep convolutional neural network DCNN, and the electronic nose pattern recognition engine LHF-GCNN. As shown in Table 1, in terms of the number of parameters, there are significant differences among different models. The number of parameters of VGG-16 is as high as 138.36 million, which means it requires a large amount of storage and computing resources for hardware deployment. In contrast, FRCNN only has 826 parameters, significantly fewer than other models. This low number of parameters gives FRCNN an advantage in low-power embedded devices. In terms of accuracy, FRCNN achieved an accuracy of 99.71%, proving that its efficient design for feature extraction and classification tasks not only has a small number of parameters but also can achieve high accuracy performance with limited computing resources.

[0143] The purpose of this experiment is to compare the FRCNN detection accuracy of the optimized hybrid quantization training algorithm for the collected industrial gas dataset under different quantization methods and precision conditions. The experiment covers two quantization methods, PTQ training quantization (quantizing the dataset, weights, and activation values after model training) and QAT quantization-aware training (simulating the quantization effect during the training process to keep the trained model at a high precision after quantization, and quantizing the model to int8 after model training), and tests are carried out under three precision conditions of float32, float16, and int8. The experimental results are as follows Figure 8 shown. First, when the input data and weights are at float32 precision, the model has the highest detection accuracy for the weight parameters and gas data, reaching 99.71%. This indicates that the model processes data most precisely under full-precision conditions and has the strongest detection ability for gas types. Second, at float16 precision, using QAT training quantization, the detection accuracy is 97.82%. Compared with float32 precision, the accuracy has decreased, but the decrease amplitude is relatively small, and the accuracy loss is only about 2%. Finally, when the input data is int8, the detection accuracy after QAT quantization-aware training is 85.28%, and the decrease in accuracy is more obvious. This may be because the model's ability to represent data is greatly limited at a very low accuracy, and although QAT tries its best to optimize the training process, it is still difficult to completely make up for the loss of accuracy.

[0144] The training results of the mixed-precision quantization algorithm used in this paper are as follows Figure 9 shown. As can be seen from Figure 9 a, the accuracy rises rapidly at the beginning and then gradually stabilizes, indicating that the model has a fast learning speed in the initial stage of training, and the accuracy becomes higher and higher as the training progresses. As can be seen from Figure 9 b, the loss is very high at the beginning and then drops rapidly, finally converging to a very low level, indicating that the prediction error of the model is very large at the beginning of training, and the prediction error gradually decreases as the training progresses. Compared with the 32-bit floating-point model, QAT quantization reduces the model size to 1 / 4 of the original, and float16 precision reduces the input data storage to 1 / 2 of the original, which is extremely important for FPGA devices with limited memory resources.

[0145] This experiment uses the ZymgXC7Z020 FPGA development board with a clock frequency of 150 MHz as the platform, and successfully maps the network algorithm to the FPGA through the advanced synthesis tool Xilinx Vivado HS, realizing customized hardware acceleration.

[0146] Table 2 Comparison of different platforms

[0147]

[0148] By comparing different hardware acceleration platforms, as shown in Table 2, it can be found that the GPU is not suitable for the hardware acceleration of such tasks, while the optimized FPGA chip performs well. First, in terms of power consumption, the FPGA is only 1.7W, far lower than 20W of the CPU and 10W of the GPU, which is crucial for realizing a low-power embedded system. Second, in terms of speed, the FPGA is as fast as the CPU and performs well in actual tasks. Although the CPU has a clock frequency as high as 2.20 GHz, the clock frequency of the FPGA is relatively low. Finally, in terms of energy efficiency, the FPGA far exceeds 6.0 s / W of the CPU and 0.7 fps / W of the GPU with a performance of 58.8 s / W. This shows that the FPGA has obvious advantages in achieving higher performance with lower power consumption when processing this task.

[0149] Table 3 Resource Utilization Overview

[0150]

[0151] Table 3 shows the resource utilization of the algorithm deployed on the Xilinx mid-to-high-end FPGA chip xc7z020clg400. It can be seen from the table that the resource utilization of LUT is 7.25%, LUTRAM is 1.88%, FF is 4.85%, BRAM is 6.07%, DSP is 3.63%, and BUFG is 3.12%. Generally speaking, the resource utilization is very low. However, it should be clear that the resource utilization is a relative concept. For some cheaper FPGA chips, such resource utilization may still seem high. Generally, the cost of an FPGA is usually proportional to the amount of resources it has. Therefore, by optimizing FPGA resources, the hardware cost of the FPGA can be effectively reduced. This enables many cheaper FPGA chips to implement the electronic nose system. At the same time, the remaining resources can also more flexibly meet the needs of future chip updates and iterations, providing a good foundation for the upgrade and expansion of the system.

[0152] The present invention establishes a comprehensive gas data acquisition system, collects the response data of ten toxic and harmful gases, maps them into a two-dimensional matrix of time series, and uses the matrix diagram for algorithm training and prediction. At the same time, an efficient and lightweight gas detection neural network is proposed for gas classification, adopting different precision quantization strategies to optimize the quantization scheme of the gas detection algorithm to meet the deployment requirements of FPGAs. Finally, based on the Winograd convolution algorithm, a reconfigurable FPGA convolution engine is designed, and the algorithm is deployed on the FPGA to achieve a portable electronic nose with low power consumption, low latency, and low resources. Experiments show that the recognition accuracy of this algorithm on the FPGA reaches 99.71%, and the accuracy loss after quantization is less than 2%. The acceleration effect reaches 100 fps, and the energy efficiency reaches 58.8 fps / W. The resource utilization rate of the deployment on the xc7z020clg400 chip does not exceed 10%. A portable electronic nose with high precision, low latency, low power consumption, and low cost is realized.

[0153] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A gas recognition method based on a depthwise separable convolutional network, characterized in that, It includes the following steps: Step 1: Construct a gas recognition system based on a depthwise separable convolutional network. The gas recognition system is provided with a gas collection module and a depthwise separable convolutional network FRCNN connected in sequence. The depthwise separable convolutional network FRCNN is provided with an input layer, a first feature extraction block, a first max pooling layer, a second feature extraction block, a second max pooling layer, a third feature extraction block, a fourth feature extraction block, a fully connected layer, and an output layer connected in sequence. The output end of the fully connected layer is built-in with a softmax activation function; The first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block have the same structure, and are all provided with a channel residual allocation module CRA, a depthwise separable convolution module DSConv, a depth aggregation module DA, and a multiplication unit. The depthwise separable convolution module DSConv is provided with a depth convolution layer and a pointwise convolution layer connected in sequence; Step 2: The gas collection module collects gas feature data a with time series attributes in real time, converts it into two-dimensional matrix data b, and transmits it to the depthwise separable convolutional network FRCNN; Step 3: The input layer of the depthwise separable convolutional network FRCNN obtains the two-dimensional matrix data b and transmits it to the first feature extraction block; Step 4: The first feature extraction block performs a feature extraction operation on the two-dimensional matrix data b to obtain first feature data c and transmits it to the first max pooling layer; Step 5: The first max pooling layer performs a max pooling operation on the first feature data c to obtain first pooled data d and transmits it to the second feature extraction block; Step 6: The second feature extraction block performs a feature extraction operation on the first pooled data d to obtain second feature data e and transmits it to the second max pooling layer; Step 7: The second max pooling layer performs a max pooling operation on the second feature data e to obtain second pooled data f and transmits it to the third feature extraction block; Step 8: The third feature extraction block performs a feature extraction operation on the second pooled data f to obtain third feature data g and transmits it to the fourth feature extraction block; Step 9: The fourth feature extraction block performs a feature extraction operation on the third feature data g to obtain fourth feature data h and transmits it to the fully connected layer; Step 10: The fully connected layer performs a fully connected operation on the fourth feature data h, and then outputs the gas recognition result through the softmax activation function.

2. The gas recognition method based on a depthwise separable convolutional network according to claim 1, wherein: The i-th feature extraction block performs a feature extraction operation on its input feature X, and the process is as follows: S1: The depth convolution layer in the depthwise separable convolution module DSConv uses a depth convolution kernel of size 3×3×1 to perform a convolution operation on the data of each input channel in the input feature x to obtain depth convolution data x1 and transmits it to the pointwise convolution layer; The pointwise convolution layer uses a 1×1 convolutional kernel to perform pointwise convolution operations on the depth convolution data x1, combines the depth convolution data x1 in the channel dimension, completes information fusion between different channels, obtains pointwise convolution data x2, and transmits it to the depth aggregation module DA; S2: The depth aggregation module DA performs four iterative accumulations on the pointwise convolution data x2 to obtain depth aggregation data x3, and transmits it to the multiplication unit; S3: The channel residual allocation module CRA uses a 1×1 convolutional layer to weight the input feature X mapping in the channel dimension, then performs non-linear mapping through the ReLU activation function to generate an attention map, automatically learns the importance weights of different channels, generates the attention weight w, and transmits it to the multiplication unit; S4: The multiplication unit multiplies the depth aggregation data x3 and the attention weight w, and outputs the i-th feature data, where i ∈ [1, 4].

3. The gas recognition method based on a depthwise separable convolutional network according to claim 1, wherein: The gas collection module is a gas converter array, which is composed of 15 metal oxide semiconductor gas sensors, an analog-to-digital converter, and peripheral circuits.

4. The gas recognition method based on a depthwise separable convolutional network according to claim 1, characterized in that: The depthwise separable convolutional network FRCNN is quantized and trained using a hybrid quantization training algorithm, and its training process is as follows: First, set the initial learning rate of the depthwise separable convolutional network FRCNN model to 0.01, set the training iteration times to 1000, calculate the gradient using the cross-entropy loss function, and update the model parameters using the adaptive moment estimation AdamW optimizer with weight decay; subsequently, use input data with 8-bit, 16-bit, and 32-bit quantization bits for training; The QAT algorithm is used to quantize and dequantize the weights and activation values of the algorithm during training to simulate the quantization effect during the inference process; at the same time, adjust the training learning rate to make the algorithm converge to the global optimal solution.

5. The gas recognition method based on a depthwise separable convolutional network according to claim 4, wherein: At the beginning of training, initialize the weights and activation values of the algorithm using floating-point numbers, and then quantize the weights and activation values to 8-bit integers. The quantization formula is as follows: where, X is the floating-point number to be quantized, X q is the quantized integer, S is the scale factor, Z is the zero point, and round is the rounding operation; The expression of the scale factor S is as follows: Among them, and represent the minimum and maximum values of int8; After forward propagation, the quantized values are dequantized back to floating-point representation, and this process is expressed as: X = S(X q -Z) After calculating the loss using the dequantized weights and activation values, perform backpropagation to calculate the gradient, and then update the weights and activation values in floating-point representation.

6. The gas recognition method based on the depthwise separable convolutional network according to claim 2, characterized in that: The gas recognition system based on the depthwise separable convolutional network is also provided with an electronic nose edge device, and the electronic nose edge device is provided with a processing system PS and a programmable logic module PL; The processing system PS is provided with an ARM processor, a DDR3 random access memory, and a UART serial port; the programmable logic module PL is provided with a hardware accelerator, an on-chip DRAM memory, a DSP computing unit, and an AXI bus; After the ARM processor initializes the system, it reads the two-dimensional matrix data of the gas collected by the gas collection module from the DDR3 random access memory, then uses the AXI bus to complete the data interaction between the processing system PS and the programmable logic module PL, and transmits the collected two-dimensional matrix data of the gas to the hardware accelerator in the programmable logic module PL. The reconfigurable hardware accelerator uses the Winograd convolution engine to perform edge acceleration on the depth separable convolution network FRCNN. Finally, the calculation result is sent back to the processing system PS through the AXI bus, and then the recognition result is output to the serial port through the UART controller for real-time monitoring on the computer.

7. The gas recognition method based on a depthwise separable convolutional network according to claim 6, characterized in that: The Winograd convolution engine is an algorithm used to accelerate the calculation of convolutional neural networks. It decomposes the convolution process into a series of matrix multiplication operations to reduce the number of multiplications in the convolution. The Winograd convolution engine nests the one-dimensional algorithm F(m,r) with itself to obtain the two-dimensional algorithm F(m×m,r×r) for representing two-dimensional convolution, expressed as: Y = A T [[GgG T ⊙[B T dB]]A where the symbol "⊙" represents element-wise multiplication, g is a filter of size r×r, and d is an image block of size (m+r-1)×(m+r-1). There are r-1 overlapping elements between adjacent blocks, and the boundary blocks are padded with zeros. For F(2×2,3×3), its transformation matrix is: The Winograd convolution engine is used to accelerate the 3×3 depth convolution kernel in the depth convolution layer using the Winograd algorithm of F(2×2,3×3).

8. The gas recognition method based on a depthwise separable convolutional network according to claim 7, characterized in that: The parallel processing ability of the Winograd convolution engine is enhanced by adopting a loop tiling strategy.

9. The gas recognition method based on a depthwise separable convolutional network according to claim 6, wherein: The ARM processor is specifically the ARM Cortex-A9 chip, and the hardware accelerator is specifically the XC7Z020 ZYNQ hardware acceleration chip.