Improved pooling method based on pixel state coding and neural network impression unit
Through an improved pooling method based on pixel state encoding and neural network impression units, the impression unit matrix group is used to fit the pixel gradient distribution state, which solves the problem of gradient information loss in the pooling layer of the convolutional neural network and improves the effect of feature extraction and image reconstruction.
Patent Information
- Application Number
- CN202510968045.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
The pooling layer of existing convolutional neural networks causes the loss of pixel gradient information during the feature extraction process, affecting the accuracy of tasks such as feature point positioning detection and semantic segmentation.
An improved pooling method based on pixel state encoding and neural network impression unit is adopted. The impression unit matrix group is combined with maximum pooling to fit the pixel gradient distribution state and compensate for the gradient distribution information loss caused by pooling sampling.
It effectively preserves pixel gradient distribution information, enhances the network's ability to represent the saliency of local features, improves feature channel activity and data loss, and is suitable for image reconstruction and feature extraction of complex tasks.
Smart Images

Figure CN120806015A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computer vision, and mainly relates to an improved pooling method based on pixel state coding and neural network impression unit. BACKGROUND
[0002] Convolutional Neural Networks (CNNs) as the core architecture of deep learning in computer vision, its development is inspired by neuroscience, and with the improvement of computing power and the evolution of task demand is constantly optimized. The core advantage of CNN is its local receptive field and hierarchical feature extraction mechanism, and the pooling module further enhances the translation invariance and computational efficiency of the model.
[0003] Early CNN models (such as LeNet-5) are inspired by Hubel & Wiesel's research on visual cortex, using convolutional layers to simulate the local feature detection ability of simple cells (S-cells), and pooling layers to simulate the position invariance response of complex cells (C-cells). In 2012, AlexNet model through the combination of ReLU activation function, GPU parallel computing and Max Pooling, made a breakthrough in ImageNet competition, laid the foundation of modern CNN. Subsequently, VGGNet model proved that small convolution kernel (3x3) deep stack can improve the feature expression ability, and ResNet model through the residual connection (Skip Connections) solves the gradient vanishing problem of deep network, makes CNN can train hundreds of layers of network, effectively improves the adaptability of network for complex task reasoning.
[0004] Wherein the pooling operation is initially used to reduce the spatial dimension of the feature map, reduce the amount of calculation and enhance the robustness. Max pooling (Max Pooling) is the mainstream choice of early CNN because it can retain significant features and suppress part of the noise. However, the pooling layer can cause spatial information loss, so when facing the needs of more complex tasks (such as target detection and semantic segmentation), the related research field makes corresponding improvements to the pooling structure. Among them, in order to meet the input needs of different scales, spatial pyramid pooling (Spatial Pyramid Pooling, SPP) is introduced, which allows fixed-length features to be generated for inputs of arbitrary size, effectively improving the input universality of the model. In the task of target detection, the region of interest pooling (RoI Pooling) method uses precise pooling of the region of interest (Region of Interest, RoI) to retain spatial details, effectively improving the target detection accuracy of the region of interest. Global average pooling (Global Average Pooling, GAP) replaces the fully connected layer, effectively reducing the model parameter amount while enhancing the network interpretability (such as Class Activation Mapping). In recent years, with the rise of attention mechanism and dynamic calculation, adaptive pooling (such as learnable pooling, attention-guided pooling) has gradually become a research hotspot, to balance the computational efficiency and feature preservation in a more flexible way. In summary, the CNN backbone network in the process of feature extraction, in order to ensure the calculation efficiency, through the step-by-step calculation sampling of data to realize the compression and simplification of features, but the existing method tends to use a single pixel sample obtained by sampling the region to complete the representation of the information of the region, which leads to the complete loss of the gradient information of the region pixels. In the calculation process of feature point positioning detection, semantic segmentation and other tasks, this loss will lead to defects in the integrity of the global features guaranteed by the neural network model and insufficient ability to represent the local feature saliency. In order to solve this problem, the present application designs an improved pooling method based on auto-encoding structure and neural network impression unit, which effectively realizes the fitting of the pixel gradient distribution state by combining the impression unit matrix group with the maximum pooling, and then completes the spatial dimension reduction of the sampling region and the effective extraction of the pixel distribution state information. SUMMARY
[0005] In order to overcome the shortcomings of the prior art, in order to solve the problem of loss of pixel gradient information in the existing maximum pooling layer, and to improve the local feature saliency representation ability of the neural network, the present application provides an improved pooling method based on pixel state coding and neural network impression unit.
[0006] An improved pooling method based on pixel state encoding and neural network impression unit, comprising an impression unit-based enhanced pooling layer; the input of the impression unit-based enhanced pooling layer is an input feature map and a pooling parameter; the pooling parameter comprises a sampling window size K and a step size; the output of the impression unit-based enhanced pooling layer comprises a sampling result feature map M and an impression encoding feature map C; the sampling result feature map M is consistent with the output result of the existing maximum pooling layer; The impression unit-based enhanced pooling layer comprises an impression unit initialization module, an impression unit feature pooling encoding module and an impression unit updating module; the impression unit matrix group comprises B impression unit matrices; the size of each impression unit matrix is the same as the size of the sampling window; the impression unit initialization module initializes the impression unit matrix group; the impression unit feature pooling encoding module processes the initialized impression unit matrix group and the input feature map to obtain the impression encoding feature map C; the impression unit updating module counts the elements of the impression encoding feature map C, and updates the impression unit matrix group according to the percentage of the corresponding number of the impression unit matrix in the impression encoding feature map C; the impression unit matrix group is used to fit the gradient distribution state of the sampling window input feature map matrix, forms a semantic representation of the pixel distribution state in the sampling area, ensures the sampling of the maximum value of the pixels in the sampling area, and realizes the fitting of the gradient distribution state of the pixels to compensate for the loss of gradient distribution information caused by the pooling sampling.
[0007] Further, the impression unit initialization module initializes the impression unit matrix group; the impression unit feature pooling encoding module processes the initialized impression unit matrix group and the input feature map to obtain the impression encoding feature map C. The step of initializing the impression unit matrix group I is:
[0008] Wherein, I represents the impression unit matrix group, B is the number of impression unit matrices, is the zth impression unit matrix of the impression unit matrix group, wherein the qth row and pth column element is , bin(z)[qK+l] represents the qK+pth digit of the result after binaryzation of the integer z.
[0009] Further, the step of processing the initialized impression unit matrix group and the input feature map by the impression unit feature pooling encoding module to obtain the impression encoding feature map C is: According to the sampling window size K, the pixels in the sampling area of the input feature map X are extracted to obtain sampling matrices ; The range of ; Initialize the sampling result feature map M; all elements in the sampling result feature map M are empty; the size of the sampling result feature map M is times the size of the input feature map X; Initialize the impression encoding feature map C, all elements in the impression encoding feature map C are empty; the size of the impression encoding feature map C is the same as the size of the sampling result feature map M; ; Extract the maximum value element of the sampling matrix , and obtain the maximum pooling sampling value ; The calculation process of the impression encoding feature map C is as follows: Subtract each impression unit matrix in the impression unit matrix group I from the sampling matrix , and select the number of the impression unit matrix with the minimum L1 norm in the difference value result matrix as the encoding calculation result ; The impression encoding calculation process is represented as:
[0010] Where I represents the impression unit matrix group, X is the input feature map, is the z-th sampling matrix, argmin is the minimum value index calculation, is the impression encoding result of , is the impression encoding result of , is the impression encoding feature map; is the element value of the maximum pooling sampling; Further, the impression unit updating module counts the elements of the impression encoding feature map C and calculates the percentage of the number corresponding to each impression unit matrix in C n; marks the impression unit matrix with the lowest percentage and its index z, and uses the percentage of the remaining impression unit matrices as weights to perform weighted sum processing on the remaining impression unit matrices to obtain the updated impression unit matrix ; Replace the impression unit matrix with the lowest percentage with the updated impression unit matrix , and complete the update of the impression unit matrix group; The processing steps of the impression unit updating module are as follows:
[0011] Where I represents the impression unit matrix group, N is the impression unit matrix encoding probability set, and the element is the number corresponding to the z-th impression unit matrix in the impression encoding feature map argmin is the index of the minimum value, and the percentage of the index in the total number of indexes in the matrix is calculated, is the index of the impression unit matrix with the lowest usage rate, is the updated impression unit matrix.
[0012] Further, the image feature decoding reconstruction method based on the enhanced impression unit pooling layer is: According to the sampling window size K, the up-sampling result matrix Y is initialized; the value of the up-sampling result matrix Y is null; the size of the up-sampling result matrix Y is K times the size of the input feature map X’; 2 ; The up-sampling result matrix Y is uniformly divided into up-sampling matrices to be sampled; According to the sampling window size K, the input feature map X’ sampling area pixel extraction is completed, and up-sampling elements are obtained; According to the index of the up-sampling element in , the corresponding impression encoding result in the impression encoding feature map C is extracted, and the impression encoding result is used as an index to retrieve the corresponding impression matrix in the impression unit matrix group I; the up-sampling element is multiplied by the corresponding impression matrix , and the obtained result matrix is placed in the up-sampling matrix , and the impression unit decoding reconstruction calculation of a single sampling window is completed; The above process is repeated times, and the obtained up-sampling matrices are merged into the up-sampling result matrix Y, and the up-sampling result matrix Y and the compression ratio are input into the interpolation function to calculate the output decoding reconstruction feature map corresponding to the input ; The calculation method of the process can be summarized as follows: (4) Where I represents the impression unit matrix group, C is the impression encoding feature map, is the input feature map corresponding to the network level, Interpolate is the interpolation function, is the compression ratio parameter used to control the scale of the reconstruction result .
[0013] The beneficial effects of the present application are: the present application aims at the problem of serious loss of gradient distribution state information during calculation of traditional pooling structure, and proposes a feature encoding representation scheme for gradient distribution state of the pooling region. The scheme uses a group of impression unit matrix groups combined with the classic max pooling to effectively realize the fitting of the pixel gradient distribution state while ensuring the sampling of the maximum value of the pixels in the region, and finally outputs the maximum pooling result obtained by sampling calculation and the gradient distribution fitting information obtained by encoding calculation. Compared with the classic pooling structure, the average information entropy (IE) of the network model bottom layer feature channel in the image reconstruction task is 1.32 times that of the classic pooling structure network, and the average gradient (AG) is 1.57 times that of the classic scheme, which compensates for the loss of gradient distribution information caused by pooling sampling, effectively improves the channel activity and data loss of the network feature, and helps the network model to complete more complex fitting tasks with a lighter depth structure. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 It is an improved pooling method flow chart based on pixel state encoding and neural network impression unit; Figure 2 The internal data structure diagram of the encoding pooling layer of the present application; Figure 3 It is an impression unit pooling encoding-reconstruction principle diagram; Figure 4 It is an impression unit encoding pooling layer in the deployment state diagram of the network structure; Figure 5 It is a feature extraction reconstruction result and a traditional scheme comparison diagram of the present application; wherein (a) is the reconstruction result of the autoencoder image reconstruction network without any pooling compensation; (b) is the reconstruction result of the autoencoder image reconstruction network with skip connection compensation; (c) is the reconstruction result of the autoencoder image reconstruction network with pooling position compensation; (d) is the reconstruction result of the autoencoder image reconstruction network deploying the pooling scheme of the present application; Figure 6 It is a feature thermal visualization comparison result diagram of the present application and the traditional scheme; wherein (a) is the visualization comparison result diagram of the autoencoder image reconstruction network and its bottom layer feature map without any pooling compensation; (b) is the visualization comparison result diagram of the autoencoder image reconstruction network and its bottom layer feature map with skip connection compensation; (c) is the visualization comparison result diagram of the autoencoder image reconstruction network and its bottom layer feature map with pooling position compensation; (d) is the visualization comparison result diagram of the autoencoder image reconstruction network and its bottom layer feature map deploying the pooling scheme of the present application. DETAILED DESCRIPTION
[0015] The technical scheme adopted by the present application to solve its technical problems is an improved pooling method based on pixel state coding and neural network impression unit, comprising an enhanced impression unit-based pooling layer; the enhanced impression unit-based pooling layer adds an impression unit matrix group on the basis of an existing maximum pooling layer, the impression unit matrix group is used to fit the gradient distribution state of a sampling window input feature map matrix, forms a semantic representation of the pixel distribution state of the sampling area, and realizes fitting of the pixel gradient distribution state while ensuring sampling of the maximum value of the pixels in the sampling area, so as to compensate for the loss of gradient distribution information caused by pooling sampling; The input of the enhanced impression unit-based pooling layer comprises an input feature map and a pooling parameter; the pooling parameter comprises a sampling window size K and a step length; The output of the enhanced impression unit-based pooling layer comprises a sampling result feature map M and an impression coding feature map C; The enhanced impression unit-based pooling layer adds an impression unit initialization module, an impression unit feature pooling coding module and an impression unit updating module on the basis of the existing maximum pooling layer; The impression unit matrix group comprises B impression unit matrices; the size of each impression unit matrix is the same as that of the sampling window; The impression unit initialization module initializes the impression unit matrix group; the impression unit feature pooling coding module processes the initialized impression unit matrix group and the input feature map to obtain the impression coding feature map C; The step of initializing the impression unit matrix group I is:
[0016] Wherein, I represents the impression unit matrix group, B is the number of impression unit matrices, is the zth impression unit matrix of the impression unit matrix group, wherein the element in the qth row and the pth column is bin(z)[qK+l] represents the qK+pth digit of the result obtained by binaryzing the integer z; This link is manifested in the implementation link as an impression unit matrix group initialization module function, the input of the impression unit matrix group initialization module function is a pooling parameter, the data format of the pooling parameter is a dictionary dict, the output of the impression unit matrix group initialization module function is an impression unit matrix group, the data format of the impression unit matrix group is a tensor type, the size of the impression unit matrix group is a multi-dimensional matrix of B*K*K, K is the size of the sampling window, and the calculation logic of the number B of impression unit matrices is:
[0017] Wherein, B is the number of impression unit matrices, and denotes the number of cases of selecting j elements from K elements; In a specific embodiment, the above initialization process is exemplified by setting the window size K as 2, the impression unit matrix group is composed of 2*2 size matrices, using 0, 1 to represent all possible states of the unit pixel array gradient distribution in the 2*2 size region; The steps of the impression unit feature pooling encoding module for processing according to the initialized impression unit matrix group and the input feature map to obtain the impression encoding feature map C are: According to the sampling window size K, the input feature map X sampling region pixel extraction is completed, and The number of to-be-sampled matrices is The range of ; Initialize the sampling result feature map M; the elements in the sampling result feature map M are all empty; the size of the sampling result feature map M is times the size of the input feature map X; Initialize the impression encoding feature map C; the elements in the impression encoding feature map C are all empty; the size of the impression encoding feature map is the same as the size of the sampling result feature map M; Extract the maximum value element of the sampling matrix , and obtain the maximum pooling sampling value ; The calculation process of the impression encoding feature map C is: Subtract each impression unit matrix in the impression unit matrix group I from the sampling matrix , and select the number of the impression unit matrix with the minimum L1 norm of the difference value result matrix L as the encoding calculation result ; The impression encoding calculation process is represented as:
[0018] Where I represents the impression unit matrix group, X is the input feature map, is the th sampling matrix, argmin is the minimum value index calculation, is the impression encoding result of , C is the impression encoding feature map; is the element value of the maximum pooling sampling; At the algorithm level, this link is represented as a pooling layer function, inputting deep feature data with height H, width W, channel number CH, and batch size b into the encoding pooling layer. Assuming that the kernel size is 2 and the step is 2 in the parameters, the output is the maximum pooling data with height H / 2, width W / 2, channel number CH, and batch size b, and the pixel distribution impression encoding data with height H / 2, width W / 2, channel number CH, and batch size b. The impression unit update module counts the elements of the impression encoding feature map C and calculates the percentage n of the number corresponding to each impression unit matrix in C. The impression unit matrix with the lowest percentage is marked and its index z, and the percentage of the remaining impression unit matrices is used as a weight to perform weighted summation processing on the remaining impression unit matrices to obtain the updated impression unit matrix The updated impression unit matrix replaces the impression unit matrix with the lowest percentage , completing the update of the impression unit matrix group. The processing steps of the impression unit update module are as follows:
[0019] where I represents the impression unit matrix group, N is the impression unit matrix encoding probability set, and the element is the percentage of the number corresponding to the zth impression unit matrix in the impression encoding feature map , argmin is the minimum value index calculation, is the index of the impression unit matrix with the lowest usage rate, is the updated impression unit matrix. The image feature decoding reconstruction method based on the enhanced impression unit-based pooling layer is as follows: Initialize the up-sampling result matrix Y according to the sampling window size K. The value of the up-sampling result matrix Y is empty, and the size of the up-sampling result matrix Y is K 2 times the size of the input feature map ; Uniformly divide the up-sampling result matrix Y into up-sampling matrices to be sampled; Extract the pixel of the sampling area of the input feature map X' according to the sampling window size K to obtain up-sampling elements ; According to the index of the up-sampling element in , extract the corresponding impression encoding result from the impression encoding feature map C, and use the impression encoding result Use it as an index to retrieve the corresponding impression matrix in the impression unit matrix group I ; The elements to be upsampled And the corresponding impression matrix Multiply, and the resulting matrix is placed into the matrix to be upsampled , complete the decoding and reconstruction calculation of the impression unit of a single sampling window; Repeat the above process times, the resulting upsampling matrix Merge into upsampling result matrix Y, and combine upsampling result matrix Y and compression ratio Input to the interpolation function and calculate the corresponding input The output decoding reconstructs the feature map ; The calculation method of this process can be summarized as follows: (4) Among them, I represents the impression unit matrix group, C is the impression coding feature map, is the input feature map of the corresponding network layer, Interpolate is the interpolation function, Compression ratio parameter used to control the reconstruction results Scale; This link is represented as a pooling layer model function at the algorithm level. The input is feature data with height H / 2, width W / 2, number of channels CH, batch size b and image encoding data with height H / 2, width W / 2, number of channels CH, batch size b. The built-in compression ratio parameter , the output is high H, width W, number of channels CH, and deep feature data with batch size b.
[0020] The present invention will be further described below with reference to the accompanying drawings and examples. like Figure 1 This is a schematic diagram comparing the core algorithm of the present invention with the traditional method. The present method uses a set of impression unit matrices combined with classic maximum pooling to effectively fit the pixel gradient distribution state while ensuring the sampling of the maximum value of pixels in the area. It ultimately outputs two sets of sampling data of equal size, namely the maximum pooling result obtained by sampling calculation and the gradient distribution fitting information obtained by encoding calculation. The pooling result still participates in the pooling calculation of the classic backbone network as normal, and the gradient distribution information is input into the decoding calculation module downstream of the network (such as upsampling processing) to compensate for the loss of gradient distribution information caused by pooling sampling. The following takes the image feature extraction and reconstruction tasks as an example to further illustrate the specific implementation steps of the present invention.
[0021] (1) Network building and method deployment; This process involves a feature extraction network. Assume that the input image size is , that is, the image P has a width of W, a length of H, and a channel number of C. A CNN autoencoder network is used as the method deployment platform. The model has 4 sampling layers in the encoder and decoder stages (pooling layers in the encoder and up-sampling layers in the decoder). Therefore, the feature data F size of the model at the i-th sampling layer is The deployment work of the present application is divided into impression unit initialization and sampling layer deployment. The impression unit initialization needs to initialize the calculation of the impression unit matrix group at each level according to the preset parameters (such as the sampling region size, the step length, etc.) of the pooling sampling at each level. The calculation process can be seen in the foregoing, and will not be repeated here. The sampling layer deployment work needs to replace the pooling layers and the up-sampling layers in the autoencoder network. The pooling layers are changed to the encoding pooling layers, the up-sampling layers are changed to the decoding reconstruction layers, and the corresponding impression unit matrices are loaded. The principle of this process can be seen in Figure 3 .
[0022] (2) Network training; This step first pre-processes the input image (such as normalization, blocking), and then extracts multi-level features through the convolution layer with a fixed 3x3 convolution kernel combined with the encoding pooling layer of the present application; in the reconstruction stage, the decoding reconstruction layer of the present application is used for up-sampling, and finally the end-to-end training is performed through the pixel-level mean square error loss function (MSE). The data set is COCO2024, the training round number is 20, and the learning rate is 0.001.
[0023] In this training process, the model uses the back propagation function built in the impression unit initialization module to count the encoding index usage frequency of each impression unit matrix in the encoding result matrix obtained by each level of encoding pooling, and uses the initial impression unit group to obtain a new impression unit matrix through the matrix operation of step 4 in the foregoing, to replace the unit with the lowest encoding calculation participation degree in the initial impression unit matrix group, and participate in the next round of training and reasoning.
[0024] (3) Image feature reconstruction reasoning; The reasoning process of this image feature extraction and reconstruction network first needs to pre-process the input image through histogram equalization, and then input it into the trained reconstruction network model (usually truncated to a certain intermediate layer) to extract deep features. These multi-level features are encoded and pooled through the pre-trained reconstruction module, and finally a high-resolution or repaired output image is generated. The entire process is completed in one time using a feedforward method without iterative optimization. The output result maintains the high-level semantic features consistent with the feature space, and is suitable for visual tasks such as super-resolution reconstruction, style transfer generation, and other tasks that require consistent semantics.
[0025] The technical effects of the application are illustrated through feature extraction and reconstruction experiments of aerial images.
[0026] The aerial image acquisition device is DJI Spark 4 Pro unmanned aerial vehicle, carrying a self-made strapdown camera. The image acquisition site is Shizuishan City, Ningxia, the aerial vehicle flies at a height of 210 m, and the pixel size of a single aerial image is 1024x1024. The corresponding ground coverage longitude and latitude coordinate ranges are: longitude 106.333351°E-106.349786°E, latitude 39.207651°N-39.222293°N. Ten groups of test sample images are formed by multiple shootings.
[0027] The experimental results are shown in Figure 5 As shown in the figure, the classical autoencoder network and its classical improvement schemes (pooling sampling position compensation and jump connection compensation) are selected as control items, and are trained together with the classical autoencoder network deployed by the application for 20 rounds on the COCO dataset (loss function is mean square error loss, learning rate is 0.001, and batchsize is 128). Finally, the above aerial image samples are tested to verify the effectiveness of the application method. The network structure using the scheme and the corresponding reconstructed image intuitively show the superiority of the application in maintaining the effective information of the image, effectively solve the grid distortion phenomenon in the classical scheme, and effectively reduce the detail distortion compared with the other two compensation schemes, so that the reconstructed image has clearer and more accurate global distribution and local details.
[0028] To intuitively show the unique advantages of the application in maintaining the distribution state information of feature data, the feature maps of each scheme in the decoder stage in the above test process are selected for visualization comparison, as shown in Figure 6 Among them, the upper left is the result of the image reconstruction network of the autoencoder without any pooling compensation, as shown in the figure, the features of each channel are generally distorted and the activity distribution of each channel is not uniform; the lower left is the result of the image reconstruction network of the autoencoder with pooling position compensation, as shown in the figure, the distortion of the features of each channel is effectively improved, but there are significant distortion points in the 6th and 7th rows of channels, which further leads to significant errors in the corresponding positions of the image reconstruction result; the upper right is the result of the image reconstruction network of the autoencoder with jump connection compensation, as shown in the figure, the image reconstruction result is excellent, but the activity of the feature channel is significantly inhibited, which is not conducive to the secondary development of the model for high-dimensional semantic features; the lower right is the result of the application, as shown in the figure, the global distribution of the reconstruction result is correct and each local detail is clear, the activity distribution of each feature channel is uniform, and the feature state in each channel shows a high similarity to the semantic distribution state of the original image, which is conducive to the subsequent research on the explainability of the model.
[0029] In summary, the experiment shows that the application can effectively realize the fitting of the pixel gradient distribution state while ensuring the maximum value sampling of the pixels in the region, wherein the pooling result still normally participates in the pooling calculation of the backbone network, and the gradient distribution information is input to the decoding calculation module (such as the up-sampling processing) downstream of the network to compensate for the loss of the gradient distribution information caused by the pooling sampling, effectively improving the channel activity and data loss of the network features. These advantages make the application have important application value and broad market prospect in the field of artificial neural network infrastructure.
[0030] The improved pooling method based on pixel state coding and neural network impression unit and the deployment mode thereof provided by the application are described in detail above, but obviously the specific implementation form of the application is not limited thereto. Various obvious changes made to the application by those skilled in the art without departing from the scope of the claims of the application are within the protection scope of the application.
Claims
1. An improved pooling method based on pixel state encoding and neural network impression unit, characterized in that: The invention comprises an enhanced pooling layer based on an impression unit; the input of the enhanced pooling layer based on the impression unit is an input feature map and a pooling parameter; the pooling parameter includes a sampling window size and a step size; the output of the enhanced pooling layer based on the impression unit includes a sampling result feature map and an impression encoding feature map; the sampling result feature map is consistent with the output result of the existing maximum pooling layer; The enhanced pooling layer based on impression units includes an impression unit initialization module, an impression unit feature pooling encoding module and an impression unit update module; the impression unit matrix group includes B impression unit matrices; the size of each impression unit matrix is the same as the size of the sampling window; the impression unit initialization module initializes the impression unit matrix group; the impression unit feature pooling encoding module processes the initialized impression unit matrix group and the input feature map to obtain the impression coding feature map; the impression unit update module counts the elements of the impression coding feature map, and updates the impression unit matrix group according to the percentage of the corresponding number of the impression unit matrix in the impression coding feature map; the impression unit matrix group is used to fit the gradient distribution state of the input feature map matrix of the sampling window to form a semantic representation of the pixel distribution state of the sampling area, and while ensuring the sampling of the maximum value of the pixels in the sampling area, realizes the fitting of the pixel gradient distribution state to compensate for the gradient distribution information loss caused by pooling sampling.
2. The improved pooling method based on pixel state encoding and neural network impression unit according to claim 1, characterized in that: The steps of initializing the impression unit matrix group I by the impression unit initialization module are as follows: Where I represents the impression unit matrix group, B is the number of impression unit matrices, is the zth impression unit matrix of the impression unit matrix group, The element in row q and column p is , bin(z)[qK+l] means taking the qK+pth digit of the result after binarizing the integer z.
3. The improved pooling method based on pixel state encoding and neural network impression unit according to claim 1, characterized in that: The impression unit feature pooling encoding module processes the initialized impression unit matrix group and the input feature map to obtain the impression coding feature map C as follows: According to the sampling window size K, the pixels of the input feature map X sampling area are extracted to obtain The matrix to be sampled ; The range is ; Initialize the sampling result feature map M; the elements in the sampling result feature map M are all empty; the size of the sampling result feature map M is the size of the input feature map X times; Initialize the impression coding feature map C, the elements in the impression coding feature map C are all empty; the impression coding feature map The size of is the same as the size of the sampling result feature map M; Extract sampling matrix The maximum value element of the maximum pooling sample value is obtained ; The calculation process of the impression coding feature map C is: The sampling matrix Difference is made with each impression unit matrix in the impression unit matrix group I. In the obtained difference result matrix, the number of the impression unit matrix with the smallest L1 norm of the difference result matrix is selected as the encoding calculation result ; The impression coding calculation process is expressed as: Where I represents the impression unit matrix group, X is the input feature map, For the Sampling matrix, argmin is the minimum index calculation, for The impression coding results, Encode feature maps for impressions; The element value sampled for max pooling.
4. The improved pooling method based on pixel state encoding and neural network impression unit according to claim 1, characterized in that: The impression unit update module counts the elements of the impression coding feature map C and calculates the percentage n of the number corresponding to each impression unit matrix in C; marks the impression unit matrix with the lowest percentage and its index z, and use the percentage of the remaining impression unit matrix as the weight to perform weighted summation on the remaining impression unit matrix to obtain the updated impression unit matrix , using the updated impression unit matrix Impression unit matrix with the lowest replacement markup percentage , complete the update of the impression unit matrix group; The processing steps of the impression unit update module are: Where I represents the impression unit matrix group, N is the impression unit matrix encoding probability set, and the elements The corresponding number of the z-th impression unit matrix in the impression encoding feature map The percentage of , argmin is the minimum index calculation, is the matrix index of the impression unit with the lowest usage rate, is the updated impression unit matrix.
5. The improved pooling method based on pixel state encoding and neural network impression unit according to claim 1, characterized in that: The image feature decoding and reconstruction method of the enhanced pooling layer based on the impression unit includes a reconstruction process; The reconstruction process is as follows: initialize the upsampling result matrix Y according to the sampling window size K; the value of the upsampling result matrix Y is a null value; the size of the upsampling result matrix Y is the input feature map K 2 times; divide the upsampling result matrix Y evenly into Upsampling matrix to be sampled ; Complete the input feature map according to the sampling window size K Extract the pixels in the sampling area and obtain elements to be upsampled ; According to the element to be upsampled exist The index in the image is used to extract the corresponding impression coding result in the impression coding feature map C. , with impression coding results Use it as an index to retrieve the corresponding impression matrix in the impression unit matrix group I ; The elements to be upsampled And the corresponding impression matrix Multiply, and the resulting matrix is placed into the matrix to be upsampled , complete the decoding and reconstruction calculation of the impression unit of a single sampling window; Repeat the rebuild process times, the resulting upsampling matrix Merge into upsampling result matrix Y, and combine upsampling result matrix Y and compression ratio Input to the interpolation function and calculate the corresponding input The output decoding reconstructs the feature map ; The mathematical expression of the image feature decoding and reconstruction method of the enhanced pooling layer based on the impression unit is: (4) Among them, I represents the impression unit matrix group, C is the impression coding feature map, is the input feature map of the corresponding network layer, Interpolate is the interpolation function, Compression ratio parameter used to control the reconstruction results scale.
6. A computer-readable storage medium storing a computer program; wherein: When the computer program is executed by a processor, it implements an improved pooling method based on pixel state encoding and neural network impression unit according to any one of claims 1 to 5.