A refractory brick shape classification and recognition method based on deep learning
Through the deep learning-based refractory brick shape classification and identification method, the problem of low detection efficiency of refractory bricks in the construction site is solved, and the rapid and accurate identification of refractory bricks is achieved, which improves construction efficiency and reduces costs.
Patent Information
- Application Number
- CN202310216732.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-03-08
AI Technical Summary
The detection efficiency of refractory bricks in the construction site is low and cannot meet the needs of automatic production lines. Especially in high-risk environments, it is difficult to complete this work with artificial vision, resulting in delayed construction progress.
The refractory brick shape classification recognition method is adopted based on deep learning. By obtaining the refractory brick video information, using random mask removal and sequential embedding encoding to process video images, combining the multi-head attention mechanism network and the parameter adaptive adjustment network, a refractory brick depth convolution classification network is built to achieve automatic recognition of shape and quality.
The rapid shape and quality classification of refractory bricks is realized, the accuracy and efficiency of inspection are improved, labor costs are reduced, intelligent unmanned inspection is realized in high-risk environments, and the construction efficiency of the construction site is improved.
Smart Images

Figure CN116385924B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image analysis, and in particular relates to a refractory brick shape classification and recognition method based on deep learning. Background Art
[0002] In existing construction sites, refractory bricks have been widely used as the main building material in different buildings. Different types of refractory bricks provide a variety of needs for different occasions. However, since the quality of refractory bricks will change during the production and transportation process, in order to evaluate the quality of refractory bricks, it is necessary to identify the refractory bricks in different factory environments.
[0003] At present, on the production line, the inspection of refractory bricks is mainly done by manpower. This process is labor-intensive and inefficient, and cannot meet the current implementation of automatic production lines in enterprises. In addition, in some high-risk environments, it is difficult for artificial vision to complete this task, which can easily delay the progress of construction on site. Therefore, in the face of the above-mentioned problems, it is urgent to propose a high-precision and intelligent method to automatically identify refractory bricks of different qualities in different operating environments. Summary of the invention
[0004] The purpose of the present invention is to solve the above-mentioned problems in the prior art and provide a refractory brick shape classification and recognition method based on deep learning, which can classify the quality of refractory bricks in different engineering environments according to their shapes. The classification and recognition method specifically includes the following steps:
[0005] Step S100: Acquire refractory brick video information, wherein the video information contains refractory bricks of different shapes; process the acquired refractory brick video using a random mask removal method; and sequentially embed code the video image after the mask is removed to obtain a masked video image;
[0006] Step S200: Taking the mask video encoded in step S100 as input, the network first performs feature learning by a multi-head attention mechanism network including two multi-head attention modules to obtain an output vector;
[0007] Step S300: The output vector of the multi-head attention mechanism network is first transformed in dimension through a dimension transformation layer to obtain a transformed feature vector;
[0008] Step S400: performing feature learning and extraction on the feature vector transformed in step 300 through the constructed M parameter adaptive adjustment networks in turn, wherein each parameter adaptive adjustment network specifically comprises a convolution layer, a normalization layer, a fully connected layer, a gating module, a channel grouping module, a channel information screening module, a spatial information screening module and an information splicing module;
[0009] The specific processing process is as follows: after the information after dimensional transformation is convolutional, the normalization layer is used to process the features. The information processed by the normalization layer is processed by the grouping function and the channel screening module to obtain the channel information; the information processed by the normalization layer is input into the channel grouping module through the full connection layer and the gating module, and then the spatial information is obtained by the spatial information screening module. The output channel information and spatial information are spliced to obtain the refractory brick shape and quality classification features;
[0010] After feature learning and extraction through the constructed M parameter adaptive adjustment networks, the refractory brick deep convolution classification network is finally constructed through the fully connected layer and the refractory brick shape classification layer;
[0011] Step S500: Based on the deep convolution classification network constructed in step S400, the refractory brick shape classification model is trained using an optimization function, wherein the classification task output uses a softmax function; the training set of the model is 1000 video images containing refractory bricks shot in step S100, and the video is marked with label information containing different shapes and qualities;
[0012] Step S600: Based on the model trained in the above step S500, the refractory bricks in the actual scene are classified according to their shape and quality; the refractory brick shape and quality classification results of all images will be divided into four corresponding levels: 1) complete shape and excellent quality; 2) relatively complete shape and good quality; 3) missing shape and average quality; 4) incomplete shape and defective product.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: the method proposed by the present invention can realize rapid classification of refractory brick shapes and quality by real-time acquisition of refractory brick images in actual scenes. On the production line, the present invention proposes a method based on deep learning to screen deep features and adaptively adjust training parameters, and the obtained deep model greatly improves classification accuracy. At the same time, the method can perform uninterrupted refractory brick shape detection tasks, reduce the labor cost of refractory brick detection work, and realize intelligent unmanned refractory brick shape classification tasks in some high-risk environments. This allows the factory to effectively improve construction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 The present invention discloses a deep learning network structure diagram for refractory brick shape classification based on deep learning. DETAILED DESCRIPTION
[0015] The following is combined with Figure 1 The present invention is further described in detail. The refractory brick shape classification and recognition method based on deep learning provided by the present invention specifically comprises the following steps:
[0016] Step S100: Acquire refractory brick video information, wherein the video information contains refractory bricks of different shapes; the present invention collects 1000 videos containing refractory bricks at different angles and focal lengths, wherein each video contains 100-1000 frames of images;
[0017] Furthermore, in order to improve the learning efficiency of the model and reduce the space occupied by the input image, the present invention processes the input information by random mask removal for the data of the training set, and the processing method includes steps S110-S120;
[0018] Step S110: Using a random mask removal method, set some pixel values of each frame of the acquired refractory brick video to zero, and for the continuous K frame image vectors in the Nth video The corresponding K mask operator vectors are defined as: Each operator First, use the mask matrix for multiplication to achieve the zeroing operation of some pixels. The mask matrix only contains values 0 and 1, and its size is consistent with each frame of the image;
[0019] The specific implementation of the random mask removal method is to perform dot product of the elements at the corresponding positions of the image and the mask. The dot product is represented by *. Then the video X after the mask is removed is used. N Calculated by the following formula:
[0020]
[0021] Step S120: performing sequential embedding coding on the video image after the mask is removed to obtain a masked video image; the sequential embedding coding in the present invention is implemented by alternating coding of sine function and cosine function.
[0022] Step S200: Taking the mask video encoded in step S120 as input, the network first performs feature learning by a multi-head attention mechanism network including two multi-head attention modules to obtain an output vector;
[0023] Step S300: The output vector of the multi-head attention mechanism network is first transformed in dimension through a dimension transformation layer to obtain a transformed feature vector;
[0024] Step S400: performing feature learning and extraction on the feature vector transformed in step 300 through the constructed M parameter adaptive adjustment networks in turn, wherein M=4 is preferably used in the present invention;
[0025] For each parameter adaptive adjustment network, see Figure 1Each parameter adaptive adjustment network specifically includes a convolution layer, a normalization layer, a fully connected layer, a gating module, a channel grouping module, a channel information screening module, a spatial information screening module and an information splicing module; the specific processing process is as follows: after the information after dimensional transformation is subjected to convolution operation, the normalization layer is used to process the features, and the information processed by the normalization layer is processed by the grouping function and the channel screening module to obtain the channel information; the information processed by the normalization layer is input into the channel grouping module through the fully connected layer and the gating module, and then the spatial information is obtained by the spatial information screening module, and the output channel information and spatial information are spliced to obtain the refractory brick shape and quality classification features.
[0026] Specifically, the activation function used in the above convolutional layer is LeakyReLU, and the normalization operation is calculated by the following formula:
[0027]
[0028] In the formula, I i Represents the feature map of the normalization layer input, and σ 2 I i The mean and variance of ξ are 10 by default. -5 .
[0029] For the channel grouping module, in order to adjust the channel parameters, the local classification prediction value is output through the full connection in each convolution module. The mean average accuracy (mAP) output by the current convolution module is calculated in the gating module and compared with the mAP values of the previous iterations, so as to adjust the parameters of the module. The parameter adjustment parameter mAP value is set to three thresholds of 0.75, 0.85 and 0.9 respectively. Starting from the first iteration, the network channel number grouping function is adaptively adjusted according to the following rules:
[0030]
[0031] Where N group is the number of groups into which the feature map in the current module is divided, N channel is the number of feature maps contained in the current channel, N epoch Indicates the number of iterations during training. Indicates that within the set number of iterations, the module gate output classification prediction indicator has the corresponding value. If the number of iterations of model training is 50<N epoch ≤100, when mAP reaches 0.9 or above, adjust N channel / N groupIf the number of model training iterations is 100<N, it means that the model can learn key information quickly, and the number of model learning parameters can be reduced by reducing the number of groups. epoch ≤200, when mAP reaches 0.85 or above, N channel / N group The value of is adjusted to 10. When the number of iterations has been between 200 and 300, it is found that the mAP index is still below 0.85, then the network parameters should be increased and N should be set. group =1, improving the learning ability of the model.
[0032] For the channel screening module, the present invention proposes a channel maximum correlation minimum redundancy screening method. The screening method first defines the normalized output two-dimensional feature image as I i For the sake of convenience, the feature image is referred to as image in the following text. In this image, the probability of different pixel values appearing in the middle represents the information relevance it contains. In order to calculate the information relevance between different images, the present invention calculates the probability p of a pixel point with gray value i in the image appearing in the image. i Defined as:
[0033]
[0034] Where h i Represents the number of pixels with gray value i in the image, N 0 Represents the grayscale number, and the pixel grayscale in the present invention is 256. This probability value characterizes the relationship between local information and overall information in the image. Then perform the following steps:
[0035] Step Sa: Based on the probability value of the pixel point appearing in the image, the image I i The information entropy of is defined as:
[0036]
[0037] Where N 0 Indicates the number of gray levels, p i is the probability that a pixel with gray value i appears in the image.
[0038] Step Sb: Based on the above definition of single image information entropy, the joint information entropy of two images is defined as:
[0039]
[0040] Where p I (x, y) indicates that the same pixel point is in I i The gray value is x i , and Figure I j The gray value in is y jThe probability of the three images can be obtained by the joint distribution histogram of the two conventional images. Similarly, based on the above joint information entropy, the joint information entropy H(I i ;I j ;I k ) can be obtained through three pictures I i , I j , I k The joint distribution histogram of is obtained.
[0041] Step Sc: Based on the above mutual information calculation method, the mutual information calculation formula between two images is: MI(I i ,I j )=H(I i )+H(I j )-H(I i ,I j ), where H(I i ) represents image I i The information entropy, H(I j ) represents image I j Information entropy, H(I i ,I j ) Two pictures I i and I j The joint information entropy of .
[0042] Step Sd: Based on the above entropy value of a single image and the mutual information calculation method of two images, the mutual information calculation method between three images is defined as follows:
[0043] H 1 (I i,j,k )=H(I i )+H(I j )+H(I k ),
[0044] MI 1 (I i,j,k )=MI(I i ;I k )+MI(I i ;I j )+MI(I j ;I k ),
[0045] MI(I i ;I j ;I k )=H 1 (I i,j,k )-MI 1 (I i,j,k )-H(I i ;I j ;I k),
[0046] Where H 1 (I i,j,k ) represents the sum of the information entropy of the three images, MI 1 (I i,j,k ) represents the sum of mutual information of the three images after pairwise combination, MI(I i ;I j ;I k ) represents the mutual information of the three images, H(I i ;I j ;I k ) represents the joint information entropy of the three images.
[0047] Step Se: Based on the above mutual information calculation rules, the present invention proposes a channel information screening function to screen the information of the images in the group. The present invention performs information screening by calculating the information entropy of each feature map and the feature combination map in each iteration, so that the mutual information of each feature map is maximized and the feature map with the minimum redundancy between feature maps is retained. The channel information screening function is defined as follows:
[0048]
[0049]
[0050] Where S represents different feature sets, |S| is the number of sets corresponding to the feature maps of different outputs in each iteration, and its value is greater than N channel / N group The smallest integer, N channel is the number of feature map channels, N group is the number of groups divided in different channels.
[0051] Regarding the spatial information screening module, the spatial information screening of the present invention specifically includes the following processing process: Assuming that the output has N channel feature maps, the number of channels in each group after division is N channel / N group , N group The number of groups divided in different channels is calculated as follows for each group of feature maps:
[0052]
[0053] In the above formula, I i ,I j ,I k Represents different feature maps, max(I i ;I j ;I k ) means taking the maximum value of the pixel value at the same position in the three feature maps, and then forming a new feature map with the maximum values of all positions; mean(Ii ;I j ;I k ) means taking the average of the pixel values at the same position in the three feature maps, and then forming a new feature map from the average values of all positions. For each group, its channel information is mainly obtained by obtaining the maximum value and average value of the corresponding positions of all feature maps in the group to obtain two feature maps after spatial information fusion.
[0054] Furthermore, the present invention performs spatial information screening on the grouped feature maps while calculating the mutual information on the feature maps, and the screening of the spatial information of the feature maps and the channel information are processed in parallel.
[0055] After feature learning and extraction through the constructed M parameter adaptive adjustment networks, the refractory brick deep convolution classification network is finally constructed through the fully connected layer and the refractory brick shape classification layer;
[0056] Step S500: Based on the deep convolution classification network constructed in step S400, the refractory brick shape classification model is trained using an optimization function, wherein the classification task output uses a softmax function; the training set of the model is 1,000 video images containing refractory bricks shot in step S100, and the video is marked with label information containing different shapes and qualities.
[0057] Step S600: Based on the model trained in the above step S500, the refractory bricks in the actual scene are classified according to their shape and quality; the refractory brick shape and quality classification results of all images will be divided into four corresponding levels: 1) complete shape and excellent quality; 2) relatively complete shape and good quality; 3) missing shape and average quality; 4) incomplete shape and defective product.
[0058] In addition, the present application also provides a computing device and a computer-readable storage medium corresponding to a refractory brick shape classification and recognition method based on deep learning, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned refractory brick shape classification and recognition method based on deep learning.
[0059] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0060] In the description of the present invention, unless otherwise specified, the terms "upper", "lower", "left", "right", "inside", "outside", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they cannot be understood as limitations on the present invention.
[0061] Finally, it should be explained that the above technical solution is only one implementation method of the present invention. For those skilled in the art, it is easy to make various types of improvements or modifications based on the application methods and principles disclosed in the present invention, and it is not limited to the method described in the above specific implementation method of the present invention. Therefore, the method described above is only preferred and does not have a restrictive meaning.
Claims
1. A refractory brick shape classification and recognition method based on deep learning, characterized in that The following steps are included: Step S100: Acquire refractory brick video information, wherein the video information contains refractory bricks of different shapes; and process the acquired refractory brick video using a random mask removal method; For the video image after the mask is removed, sequential information marking encoding is performed on it to obtain a masked video image; The sequential information mark coding is implemented by alternating coding of sine function and cosine function; Step S200: Taking the mask video encoded in step S100 as input, the network first performs feature learning by a multi-head attention mechanism network including two multi-head attention modules to obtain an output vector; Step S300: The output vector of the multi-head attention mechanism network is first transformed in dimension through a dimension transformation layer to obtain a transformed feature vector; Step S400: performing feature learning and extraction on the feature vector transformed in step 300 through the constructed M parameter adaptive adjustment networks in turn, wherein each parameter adaptive adjustment network specifically comprises a convolution layer, a normalization layer, a fully connected layer, a gating module, a channel grouping module, a channel information screening module, a spatial information screening module and an information splicing module; The specific processing process is as follows: after the information after dimensional transformation is convolutional, the normalization layer is used to process the features. The information processed by the normalization layer is processed by the channel grouping module and the channel screening module to obtain channel information; the information processed by the normalization layer is input into the channel grouping module through the fully connected layer and the gating module, and then the spatial information is obtained by the spatial information screening module. The output channel information and spatial information are spliced to obtain the refractory brick shape and quality classification features; After feature learning and extraction through the constructed M parameter adaptive adjustment networks, the refractory brick deep convolution classification network is finally constructed through the fully connected layer and the refractory brick shape classification layer; Step S500: training the deep convolution classification network for refractory bricks constructed in step S400 using an optimization function; Step S600: Based on the network trained in step S500, the refractory brick images in the actual scene are classified according to the shape and quality of the refractory bricks.
2. The refractory brick shape classification and recognition method based on deep learning according to claim 1 is characterized in that: In the M parameter adaptive adjustment network, M=4.
3. The refractory brick shape classification and recognition method based on deep learning according to claim 1 is characterized in that The normalization operation LN is calculated by the following formula: In the formula, I i Represents the feature map of the normalization layer input, and σ2 I i The mean and variance of ξ are 10 by default. -5 .
4. The refractory brick shape classification and recognition method based on deep learning according to claim 2 is characterized in that: In the parameter adaptive adjustment network module, the mean average accuracy mAP of the current convolution module output is calculated in the gating module and compared with the mAP values of the previous iterations, so as to adjust the parameters of the module. The parameter adjustment parameter mAP value is set to three thresholds, namely 0.75, 0.85 and 0.
9. Starting from the first iteration, the network channel grouping module is adaptively adjusted according to the following rules: Where N group is the number of groups into which the feature map in the current module is divided, N channel is the number of feature maps contained in the current channel, N epoch represents the number of iterations during training. It means that within the set number of iterations, the corresponding value of the module gated output classification prediction indicator appears.
5. The refractory brick shape classification and recognition method based on deep learning according to claim 4 is characterized in that: The channel information screening function performs information screening of the images within the group by calculating the information entropy of each feature map and the feature combination map in each iteration, so that the mutual information of each feature map retained is maximized and the feature map with the smallest redundancy between feature maps is retained; the channel information screening function is defined as follows: Where S represents different feature sets, |S| is the number of sets corresponding to the feature maps of different outputs in each iteration, and its value is greater than N channel / N group The smallest integer, N channel is the number of feature map channels, N group is the number of groups divided in different channels.
6. The refractory brick shape classification and recognition method based on deep learning according to claim 1 is characterized in that The classification task output in step S500 uses a softmax function; the training set of the model is 1,000 video images containing refractory bricks taken in step S100, and the video is marked with label information containing different shapes and qualities.
7. The refractory brick shape classification and recognition method based on deep learning according to claim 1 is characterized in that In step S600, the refractory brick shape and quality classification results are divided into four corresponding levels: 1) complete shape and excellent quality; 2) relatively complete shape and good quality; 3) missing shape and average quality; 4) incomplete shape and defective product.
8. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
K-means hyperspectral image wave band clustering method based on mutual information
CN109583469A
An industrial control system intrusion detection method based on integrated learning
CN109861988A