Image quality evaluation method for smelting furnace front operation and related equipment
By extracting smoke and strong light features in the smelting furnace operation and performing multi-scale deep feature fusion and encoding, the problem of low accuracy of image quality assessment in the smelting furnace operation is solved, and higher image quality assessment accuracy is achieved.
Patent Information
- Application Number
- CN202510808636.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing no-reference image quality assessment algorithm has low accuracy in image data of the smelting furnace operation site due to visual noise interference such as smoke and strong light.
By obtaining the smoke and strong light features of the image in front of the smelting furnace, multi-scale deep feature extraction is performed, and these features are fused and pre-self-attention operated. Finally, they are encoded and spliced to obtain the image quality assessment result.
The information richness and accuracy of image quality assessment are improved, and the accuracy of image quality assessment in complex visual noise environments is effectively improved.
Smart Images

Figure CN120689320A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image quality assessment, and in particular to an image quality assessment method and related equipment for smelting furnace operations. Background Art
[0002] With the advancement of intelligent manufacturing, industrial robots are gradually replacing manual operations in heavy industries like metallurgy. Automation and intelligentization have become the future development trend of the metallurgical industry. This shift not only improves production efficiency but also allows workers to avoid dusty and high-temperature environments.
[0003] The intelligentization of robots operating at smelting furnaces relies on machine vision. Industrial robots need to obtain target location information and status information about working objects, such as the smelting furnace, from the vision system. However, variable visual interference noise, such as dust and lighting, at metallurgical work sites can affect machine vision performance. Assessing the level of visual noise interference is the first step in subsequent processing. Therefore, accurately evaluating the image quality of worksite images captured by industrial camera vision systems is a key issue in the intelligentization of metallurgical robots.
[0004] In the field of image quality assessment, early research focused on subjective evaluation methods, where humans act as observers to make subjective assessments of image quality. With the development of computer vision technology, objective image quality evaluation methods using algorithms have become a research focus. Based on whether or not they rely on reference images, objective image quality evaluation methods can be divided into three categories: full reference (FR), reduced reference (RR), and no reference (NR). Because reference images are difficult to obtain in reality, no-reference image quality assessment algorithms are more widely used in practice. However, existing no-reference image quality assessment algorithms mostly focus on the quality assessment of everyday or field images. However, image data from smelting furnace operations is subject to complex and variable visual noise interference such as smoke, dust, and strong light, resulting in low accuracy in image quality assessment of smelting furnace images. Summary of the Invention
[0005] The present application provides an image quality assessment method and related equipment for smelting furnace operations, which can solve the problem of low accuracy in image quality assessment of smelting furnace images.
[0006] In a first aspect, an embodiment of the present application provides an image quality assessment method for smelting furnace operations, the image quality assessment method comprising:
[0007] Acquire an image in front of a smelting furnace, extract smoke and dust features from the image in front of the smelting furnace to obtain smoke and dust features of the image in front of the smelting furnace, and extract strong light features from the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace;
[0008] Perform multi-scale deep feature extraction on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales;
[0009] For each deep feature, the smoke feature, strong light feature and deep feature are fused to obtain the fused feature, and the fused feature is pre-self-attention operated to obtain the final fused feature;
[0010] Encode each final fusion feature to obtain the encoding feature corresponding to each final fusion feature, and splice all the encoding features to obtain the spliced feature;
[0011] Image quality assessment is performed based on the stitching features to obtain the image quality assessment results of the smelting furnace front image.
[0012] Optionally, smoke and dust features are extracted from the image in front of the smelting furnace to obtain smoke and dust features, including:
[0013] By formula:
[0014]
[0015] Calculate the smoke feature T(x) of pixel x in the image in front of the smelting furnace;
[0016] Where ω represents the smoke residue coefficient, α represents the balance weight, θ represents the linear mapping parameter of the color attenuation prior, A represents the global atmospheric light coefficient, D(x) represents the pixel value of pixel x in the dark channel image of the smelting furnace image, x∈X, X represents the number of all pixels in the smelting furnace image, and C(x) represents the depth of field model of the smelting furnace image:
[0017]
[0018] C(x)=I V (x)-I S (x)
[0019] Among them, Ω(x) represents the local window centered at pixel x, R represents the red channel, G represents the green channel, B represents the blue channel, and I V (x) and I s (x) represents the pixel value of pixel x in V and S channels after the image in front of the smelting furnace is converted to HSV space, I c (y) represents the pixel value of pixel y in the image in front of the smelting furnace under channel c;
[0020] The smoke features of all pixels are integrated to obtain the smoke feature T of the smelting furnace image.
[0021] Optionally, the strong light feature of the image in front of the smelting furnace is extracted to obtain the strong light feature, including:
[0022] By formula:
[0023]
[0024] Calculate the strong light feature S(x) of pixel x in the image in front of the smelting furnace;
[0025] Among them, m s (x) represents the mirror coefficient, m b (x) represents the diffuse reflectance:
[0026]
[0027] Among them, R(x), G(x) and B(x) represent the pixel values of pixel x in the R, G and B channels in the image in front of the smelting furnace, respectively. s Indicates the chromaticity of the light source of the red channel, g s Indicates the chromaticity of the light source of the green channel, b s Represents the chromaticity of the light source of the blue channel, r b Indicates the diffuse chromaticity of the object in the red channel, g b Indicates the diffuse reflection chromaticity of the green channel object, b b Represents the diffuse chromaticity of the object in the blue channel;
[0028] The strong light features of all pixels are integrated to obtain the strong light feature S of the image in front of the smelting furnace.
[0029] Optionally, multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales, including:
[0030] Perform multi-scale adjustment on the image in front of the smelting furnace to obtain the adjusted image in front of the smelting furnace at each scale;
[0031] For each scale, deep features are extracted from the scale-adjusted smelting furnace image to obtain the deep features at that scale.
[0032] Optionally, deep feature extraction is performed on the scale-adjusted smelting furnace front image to obtain deep features at the scale, including:
[0033] By formula:
[0034]
[0035]
[0036] Calculate the deep features at the i-th scale
[0037] in, It represents the initial features of the smelting furnace image adjusted at the i-th scale after convolution and maximum pooling operations. represents the secondary convolution feature, represents the cubic convolution feature, represents the cubic convolution feature, Both represent convolution, f ds Represents a downsampling operation.
[0038] Optionally, smoke features, glare features, and deep features are fused to obtain fused features, including:
[0039] The smoke and dust features and the strong light features are combined to obtain the apparent features;
[0040] The surface features and deep features are fused to obtain fused features.
[0041] Optionally, the smoke and dust features and the glare features are combined to obtain apparent features, including:
[0042] By formula:
[0043]
[0044] Calculate the apparent feature F h ;
[0045] Among them, concat represents the splicing operation, T represents the smoke feature, S represents the strong light feature, R 2×H×W Indicates the dimension of the apparent feature, H represents the height of the image in front of the smelting furnace, and W represents the width of the image in front of the smelting furnace;
[0046] The surface features and deep features are integrated to obtain fused features, including:
[0047] By formula:
[0048]
[0049] Calculate the fusion features at the i-th scale
[0050] Among them, reshape represents the size adjustment operation, represents the deep features at the i-th scale, R 2050 ×(H / 32)×(W / 32) Dimensions that represent apparent features.
[0051] Optionally, all coded features are concatenated to obtain concatenated features, including:
[0052] All coding features are spliced together to obtain the initial spliced features;
[0053] Add an embedding sequence before the initial concatenated feature to obtain the embedded feature, perform multi-layer encoding on the embedded feature to obtain the final encoded feature;
[0054] The feedforward network is used to perform linear transformation and nonlinear activation on the final encoding features to obtain feedforward features;
[0055] Perform residual connection on the feedforward features and the encoding features to obtain the concatenated features.
[0056] Optionally, the image quality assessment result is an image quality score;
[0057] Image quality assessment is performed based on the stitching features to obtain the image quality assessment results of the smelting furnace image, including:
[0058] By formula:
[0059] score=Linear2(GELU(Linear1(v)))
[0060] Calculate the image quality score;
[0061] Among them, Linear1 and Linear2 represent linear mappings, GELU represents the GELU activation function, and v represents the embedded sequence in the splicing feature.
[0062] In a second aspect, an embodiment of the present application provides an image quality assessment device for smelting furnace operations, comprising:
[0063] A feature extraction module is used to obtain an image in front of a smelting furnace, perform smoke and dust feature extraction on the image in front of the smelting furnace to obtain smoke and dust features of the image in front of the smelting furnace, and perform strong light feature extraction on the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace;
[0064] The multi-scale deep feature extraction module is used to extract multi-scale deep features of the smelting furnace image to obtain the deep features of the smelting furnace image at multiple scales;
[0065] The fusion module is used to fuse the smoke feature, the strong light feature and the deep feature for each deep feature to obtain the fused feature, and perform a pre-self-attention operation on the fused feature to obtain the final fused feature;
[0066] The encoding module is used to encode each final fusion feature to obtain the encoding feature corresponding to each final fusion feature, and to splice all the encoding features to obtain the spliced feature;
[0067] The image quality assessment module is used to perform image quality assessment based on the stitching features to obtain the image quality assessment result of the smelting furnace front image.
[0068] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned image quality assessment method for smelting furnace operations when executing the above-mentioned computer program.
[0069] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned image quality assessment method for smelting furnace operations.
[0070] The above solution of the present application has the following beneficial effects:
[0071] In some embodiments of the present application, by acquiring an image in front of a smelting furnace, smoke and dust part feature extraction is performed on the image in front of the smelting furnace to obtain smoke features of the image in front of the smelting furnace, and strong light part feature extraction is performed on the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace, and then multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales, and then for each deep feature, the smoke feature, strong light feature and deep feature are fused to obtain fused features, and a pre-self-attention operation is performed on the fused features to obtain final fused features, and then each final fused feature is encoded to obtain a coding feature corresponding to each final fused feature, and all the coding features are spliced to obtain a spliced feature, and finally, image quality assessment is performed based on the spliced feature to obtain an image quality assessment result of the image in front of the smelting furnace. Among them, the smoke and strong light features of the image in front of the smelting furnace are extracted, which can capture the visual noise features such as smoke and light, perform multi-scale feature extraction on the image in front of the smelting furnace, and express the feature information of the image at multiple scales. The image quality of the image in front of the smelting furnace is evaluated based on the smoke and strong light features and deep features at multiple scales, which improves the information richness of the image quality evaluation and effectively improves the accuracy of the image quality evaluation.
[0072] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0074] Figure 1 A flowchart of a method for evaluating image quality of smelting furnace operations provided in one embodiment of the present application;
[0075] Figure 2 A schematic diagram of smoke characteristics provided in one embodiment of the present application;
[0076] Figure 3 A schematic diagram of a strong light feature provided in an embodiment of the present application;
[0077] Figure 4 A schematic diagram of image quality assessment results of a simulated environment image provided in one embodiment of the present application;
[0078] Figure 5 A schematic diagram of image quality assessment results of a real environment image provided by an embodiment of the present application;
[0079] Figure 6 A schematic diagram of the structure of an image quality assessment device for smelting furnace operations provided in one embodiment of the present application;
[0080] Figure 7 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0081] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0082] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0083] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0084] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0085] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0086] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0087] In response to the problem of low accuracy in image quality assessment of existing smelting furnace front images, an embodiment of the present application provides an image quality assessment method for smelting furnace front operations. The image quality assessment method obtains a smelting furnace front image, extracts smoke and dust features of the smelting furnace front image, and extracts strong light features of the smelting furnace front image to obtain strong light features of the smelting furnace front image. Then, multi-scale deep feature extraction is performed on the smelting furnace front image to obtain deep features of the smelting furnace front image at multiple scales. Then, for each deep feature, the smoke feature, strong light feature and deep feature are fused to obtain a fused feature. The fused feature is pre-self-attention-operated to obtain a final fused feature. Then, each final fused feature is encoded to obtain a coding feature corresponding to each final fused feature. All coding features are spliced to obtain a spliced feature. Finally, image quality assessment is performed based on the spliced feature to obtain an image quality assessment result of the smelting furnace front image. Among them, the smoke and strong light features of the image in front of the smelting furnace are extracted, which can capture the visual noise features such as smoke and light, perform multi-scale feature extraction on the image in front of the smelting furnace, and express the feature information of the image at multiple scales. The image quality of the image in front of the smelting furnace is evaluated based on the smoke and strong light features and deep features at multiple scales, which improves the information richness of the image quality evaluation and effectively improves the accuracy of the image quality evaluation.
[0088] Next, an example description is given of the image quality assessment method for smelting furnace operations provided in this application.
[0089] like Figure 1 As shown, the image quality assessment method for smelting furnace operations provided by this application includes the following steps:
[0090] Step 11, obtain the image in front of the smelting furnace, extract the smoke and dust features of the image in front of the smelting furnace to obtain the smoke and dust features of the image in front of the smelting furnace, and extract the strong light features of the image in front of the smelting furnace to obtain the strong light features of the image in front of the smelting furnace.
[0091] The above-mentioned image in front of the smelting furnace is an image of the smelting furnace operation site.
[0092] In some embodiments of the present application, a camera mounted on a metallurgical robot operating in front of a smelting furnace may be used to obtain an image of the smelting furnace at a smelting furnace operation site. The steps of extracting smoke and dust features from the image in front of the smelting furnace to obtain smoke and dust features of the image in front of the smelting furnace, and extracting strong light features from the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace include:
[0093] The first step is to extract the smoke and dust features of the image in front of the smelting furnace to obtain the smoke and dust features of the image in front of the smelting furnace.
[0094] First, through the formula:
[0095]
[0096] Calculate the smoke feature T(x) of pixel x in the image in front of the smelting furnace.
[0097] Where ω represents the smoke residue coefficient, α represents the balance weight, θ represents the linear mapping parameter of the color attenuation prior, A represents the preset global atmospheric light coefficient, which is used to describe the overall bias of the smoke ambient light on the image brightness, D(x) represents the pixel value of pixel x in the dark channel image of the smelting furnace image, x∈X, X represents the number of all pixels in the smelting furnace image, and C(x) represents the depth of field model of the smelting furnace image:
[0098]
[0099] C(x)=I V (x)-I S (x)
[0100] Among them, Ω(x) represents the local window centered at pixel x, R represents the red channel, G represents the green channel, B represents the blue channel, and I V (x) and I s (x) represents the pixel value of pixel x in V and S channels after the image in front of the smelting furnace is converted to HSV space, I c (y) represents the pixel value of pixel y in the image in front of the smelting furnace under channel c.
[0101] Then, the smoke features of all pixels are integrated to obtain the smoke feature T of the smelting furnace image.
[0102] Specifically, the smoke features of all pixels are integrated into one image according to the positions of the pixels in the image in front of the smelting furnace, so as to obtain the smoke features of the image in front of the smelting furnace.
[0103] The second step is to extract the strong light feature of the image in front of the smelting furnace to obtain the strong light feature of the image in front of the smelting furnace.
[0104] First, through the formula:
[0105]
[0106] Calculate the strong light feature S(x) of pixel x in the image in front of the smelting furnace.
[0107] Among them, m s (x) represents the mirror coefficient, m b (x) represents the diffuse reflectance:
[0108]
[0109] Among them, R(x), G(x) and B(x) represent the pixel values of pixel x in the R, G and B channels in the image in front of the smelting furnace, respectively. s Indicates the chromaticity of the light source of the red channel, g s Indicates the chromaticity of the light source of the green channel, b s Represents the chromaticity of the light source of the blue channel, r b Indicates the diffuse chromaticity of the object in the red channel, g b Indicates the diffuse reflection chromaticity of the green channel object, b b Represents the object's diffuse chromaticity in the blue channel.
[0110] Then, the strong light features of all pixels are integrated to obtain the strong light feature S of the image in front of the smelting furnace.
[0111] Specifically, the strong light features of all pixels are integrated into one picture according to the positions of the pixels in the image in front of the smelting furnace, so as to obtain the strong light features of the image in front of the smelting furnace.
[0112] This step is illustrated below with reference to a specific example.
[0113] The original image (i.e. the image in front of the smelting furnace) and the smoke inspection image (i.e. the smoke characteristics) are as follows Figure 2 As shown, Figure 2 a is the original image, Figure 2 b is the smoke inspection image, the original image and the strong light inspection image (i.e., strong light feature) are as follows Figure 3 As shown, Figure 3 a is the original image, Figure 3 b is the strong light test image.
[0114] Step 12: Perform multi-scale deep feature extraction on the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales.
[0115] In some embodiments of the present application, the step of extracting multi-scale deep features from the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales includes:
[0116] In the first step, the smelting furnace front image is adjusted at multiple scales to obtain the adjusted smelting furnace front image at each scale.
[0117] The above scale is the resolution of the image. The resolution of the smelting furnace front image is adjusted to each scale to obtain the adjusted smelting furnace front image at each scale.
[0118] It should be noted that the scale of the original smelting furnace front image can also be used as a scale, and the smelting furnace front image is directly used as the smelting furnace front image adjusted at the scale.
[0119] For example, there are three scales: original scale, first scale, and second scale. The original image of the smelting furnace is H is the image height, W is the image width, and the two adjusted images in front of the smelting furnace are H1 is the image height at the first scale, W1 is the image width at the first scale, H2 is the image height at the second scale, and W2 is the image width at the second scale, to obtain For example, the width of the adjusted image is set to s1, the scaling factor is s1 / W, and the calculated H1 is H*(s1 / W). This method can keep the original aspect ratio unchanged. At the same time, the original smelting furnace front image is used as the adjusted smelting furnace front image, and finally the adjusted smelting furnace front image at three scales is obtained:
[0120] In the second step, for each scale, deep features are extracted from the scale-adjusted smelting furnace image to obtain the deep features at that scale.
[0121] Specifically, through the formula:
[0122]
[0123] Calculate the deep features at the i-th scale
[0124] in, It represents the initial features of the smelting furnace image adjusted at the i-th scale after convolution and maximum pooling operations. represents the secondary convolution feature, represents the cubic convolution feature, represents the cubic convolution feature, Both represent convolution, f ds Indicates a downsampling operation with a stride of 2.
[0125] It should be noted that the number of channels and convolution kernels between the above convolutions are different. The upper corner indicates the number of channels of the convolution, and the lower corner indicates the convolution kernel. For example, It is a 1×1 convolution with 64 channels.
[0126] For example, For example, the image can be input into the normalization layer, and then processed by 7×7 convolution (channel number 64, step size 2), and then calculated by the maximum pooling layer to obtain the initial features The dimension of the initial feature is The dimension of the final deep feature is
[0127] In step 13, for each deep feature, the smoke feature, the strong light feature and the deep feature are fused to obtain a fused feature, and a pre-self-attention operation is performed on the fused feature to obtain the final fused feature.
[0128] In some embodiments of the present application, the steps of fusing the smoke feature, the glare feature, and the deep feature to obtain a fused feature, and performing a pre-self-attention operation on the fused feature to obtain the final fused feature include:
[0129] The first step is to combine the smoke and dust features with the strong light features to obtain the apparent features.
[0130] Specifically, through the formula:
[0131]
[0132] Calculate the apparent feature F h .
[0133] Among them, concat represents the splicing operation, T represents the smoke feature, S represents the strong light feature, R 2×H×W Indicates the dimension of the apparent feature, H represents the height of the image in front of the smelting furnace, and W represents the width of the image in front of the smelting furnace.
[0134] In the second step, the surface features and deep features are fused to obtain fused features.
[0135] Specifically, through the formula:
[0136]
[0137] Calculate the fusion features at the i-th scale
[0138] Among them, reshape represents the size adjustment operation, represents the deep features at the i-th scale, R 2050 ×(H / 32)×(W / 32) Dimensions that represent apparent features.
[0139] In the third step, pre-self-attention operation is performed on the fused features to obtain the final fused features.
[0140] Specifically, the fused features are flattened to obtain a flattened feature map, and then attention calculation is performed on it to obtain a weighted feature map. The weighted feature map is then residually connected with the flattened feature map to obtain a residually connected feature map. The residually connected feature map is recalibrated to obtain a recalibrated feature sequence, and finally restored to the original feature map spatial dimension to obtain the final fused features.
[0141] For example, the input feature map (i.e., fusion feature) is first flattened to For example, it can be shortened to C is the number of feature map channels, H F and W F are the height and width of the feature map respectively.
[0142] The flattened features are Generate query Q using three fully connected layers F , key K F Sum V F :
[0143] Q F =W q ·F flat
[0144] K F =W k ·F flat
[0145] V F =W v ·F flat
[0146] in, is a learnable parameter matrix used to map the input feature map to the attention matrix space;
[0147] Calculate attention weights:
[0148]
[0149] in, Softmax is a normalization function that makes the sum of weights equal to 1.
[0150] Use attention weight to V F Perform a weighted sum:
[0151] Y F =A·V F
[0152] in, Represents the weighted feature map.
[0153] In order to preserve the information of the original feature map, the obtained Y F With the feature map F flat Make residual connection:
[0154] Z=Y F +F flat
[0155] in, is the feature map after residual connection.
[0156] In order to strengthen the model's response to important feature channels, the obtained feature map is recalibrated through the SE channel attention module. The specific implementation process is as follows:
[0157] For Z in the sequence dimension H F W F Perform global average pooling on , and get channel statistics:
[0158]
[0159] Where Z[i,:] represents each row vector of Z.
[0160] After that, s is transformed through two layers of full connection to obtain the channel weights:
[0161] w=σ(W2·ReLU(W1s))
[0162] Among them, ReLU is the ReLU activation function, σ is the Sigmoid activation function, and W1 and W2 represent the parameters of the two fully connected layers respectively.
[0163] Then, each channel in Z is recalibrated using the above channel weights:
[0164]
[0165] in, Represents the recalibrated feature sequence.
[0166] Finally, Restore the original feature map spatial dimensions from the sequence format:
[0167]
[0168] in, Reshape is a feature map adjustment operation used to map the sequence dimension to the original feature map space dimension.
[0169] In step 14, each final fusion feature is encoded to obtain the encoding feature corresponding to each final fusion feature, and all the encoding features are spliced to obtain the spliced feature.
[0170] In some embodiments of the present application, the steps of encoding each final fused feature to obtain a coding feature corresponding to each final fused feature, and concatenating all coding features to obtain a concatenated feature include:
[0171] The first step is to encode each final fusion feature to obtain the encoding feature corresponding to each final fusion feature.
[0172] Specifically, we first add scale embedding to the final fusion feature, and then add sinusoidal position encoding to obtain the encoded feature.
[0173] Exemplary, learnable scale embeddings Size S to H S ×W S Consistent with the size of the embedded feature map, the scale embedding method is pixel-by-pixel addition, adding the scale embedding to the corresponding feature in the final fusion feature for each pixel. The expression of the above sinusoidal position encoding is:
[0174]
[0175] Where PE is the obtained position encoding, pos is the position index, i is the dimension index, and d is the total dimension of the embedding vector. The sinusoidal position encoding is also added pixel by pixel, adding the sinusoidal position encoding to the corresponding feature in the final fused feature for each pixel.
[0176] In the second step, all the encoding features are spliced to obtain the initial spliced features.
[0177] Specifically, each coding feature is flattened, and then all the flattened coding features are concatenated to obtain an initial concatenated feature.
[0178] For example, the features at each scale are flattened to For example, the flattened feature is The features after flattening at multiple scales are spliced in the sequence dimension, i.e., the HW dimension, to obtain the initial spliced features. N total is the total length of the initial splicing feature.
[0179] The third step is to add the embedding sequence before the initial splicing feature to obtain the embedded feature, and perform multi-layer encoding on the embedded feature to obtain the final encoded feature.
[0180] The above embedding sequence is used for quality score prediction. The embedding sequence is learnable and reflects the feature information in the feature where the embedding sequence is located. When the feature where the embedding sequence is located is fused, spliced, etc., the information contained in the embedding sequence is learned as the feature changes. It is usually a preset token sequence, such as the token sequence cls token in the Transformer architecture. For example, the embedding sequence added for the splicing feature is a vector It will be added at the forefront of the concatenated features, and during the training process it will pay attention to other sequences in the concatenated features to learn the summary of the overall features.
[0181] For example, the embedded features can be calculated using multiple stacked coding layers to obtain the final coding features. The multiple coding layers are connected in sequence, and the operation of each coding layer is:
[0182] Through linear transformation, we get the Q, K and V of the input data and calculate the multi-head self-attention:
[0183]
[0184] Among them, d k is the dimension of each head in the multi-head attention.
[0185] The above calculation results are used as the output of the current coding layer after residual connection and layer normalization. The output after stacking multiple layers is
[0186] The fourth step is to use the feedforward network to perform linear transformation and nonlinear activation on the final encoding features to obtain feedforward features.
[0187] For example, the nonlinear activation can use the GELU function, and the linear transformation can use the fully connected layer.
[0188] The fifth step is to perform residual connection on the feedforward features and the encoding features to obtain the splicing features.
[0189] Step 15: Perform image quality assessment based on the stitching features to obtain an image quality assessment result of the smelting furnace front image.
[0190] The above image quality evaluation result is an image quality score. The higher the image quality score, the better the quality of the image in front of the smelting furnace, and the lower the image quality score, the worse the quality of the image in front of the smelting furnace.
[0191] Specifically, through the formula:
[0192] score=Linear2(GELU(Linear1(v)))
[0193] Calculate the image quality score.
[0194] Among them, Linear1 and Linear2 represent linear mappings, GELU represents the GELU activation function, and v represents the embedded sequence in the splicing feature.
[0195] It should be noted that, through the processing of features in step 14, the embedded sequence added therein will be updated and learned with the feature processing of each step, and finally capture and integrate the feature information in the splicing feature, and provide the feature information of the splicing feature for the image quality score prediction.
[0196] It is worth mentioning that extracting the smoke and strong light features of the image in front of the smelting furnace can capture the visual noise features such as smoke and light, perform multi-scale feature extraction on the image in front of the smelting furnace, and express the feature information of the image at multiple scales. The image quality of the image in front of the smelting furnace is evaluated based on the smoke features, strong light features and deep features at multiple scales, which improves the information richness of the image quality evaluation and effectively improves the accuracy of the image quality evaluation.
[0197] The method of the present application is illustrated below with reference to a specific example.
[0198] A comparative experiment was conducted using multiple existing methods and the method of this application. The Spearman Rank Correlation Coefficient (SRCC) and the Pearson Linear Correlation Coefficient (PLCC) were selected as performance indicators. The experimental results are shown in Table 1.
[0199] Table 1
[0200] Model SRCC PLCC BRISQUE 0.665 0.681 ILNIQE 0.507 0.523 HOSA 0.671 0.694 QUR 0.797 0.805 PQR 0.880 0.884 DBCNN 0.875 0.884 MetaIQA 0.850 0.887 BIQA 0.906 0.917 MUSIQ 0.905 0.919 Method of the present application 0.920 0.939
[0201] Among them, BRISQUE is a blind / referenceless image spatial quality evaluator, ILNIQE is an image localized naturalness image quality evaluator, HOSA is a higher order statistics-based image quality assessment, WaDIQaM is a wavelet domain image quality assessment method, PQR is a perceptual quality rating method, DBCNN is a deep bilateral convolutional neural network, MetaIQA is a meta image quality assessment method, BIQA is a blind image quality assessment method, and MUSIQ is a multiscale image quality assessment method.
[0202] The results of the ablation experiment on the method of the present application are shown in Table 2.
[0203] Table 2
[0204]
[0205] It can be seen that the accuracy of image quality assessment using the method of the present application is higher than that of other existing methods.
[0206] The results of using the method of this application to evaluate the quality of images in a simulation environment are as follows: Figure 4 As shown, Figure 4 a is the Gaussian noise of the image. When Gauss_noise_sigma=0, the image quality score Iqa=0.6216, Figure 4 b is the Gaussian noise of the image. When Gauss_noise_sigma=5, the image quality score Iqa=0.5750, Figure 4 c is the Gaussian noise of the image. When Gauss_noise_sigma=10, the image quality score Iqa=0.4728, Figure 4 When d is the Gaussian noise of the image and Gauss_noise_sigma=15, the image quality score Iqa=0.4203.
[0207] The results of using the method of this application to evaluate the quality of images in a real working environment are as follows: Figure 5 As shown, Figure 5 a represents the image quality score of the first image in the real working environment Iqa=0.3148, Figure 5 a represents the image quality score of the second image in the real working environment Iqa=0.3174, Figure 5 a represents the image quality score Iqa=0.4197 of the third image in the real working environment.
[0208] It can be seen that the method of the present application can accurately evaluate the image quality of images in both simulation and real working environments.
[0209] The following is an exemplary description of the image quality assessment device for smelting furnace operations provided in this application.
[0210] like Figure 6 As shown, an embodiment of the present application provides an image quality assessment device for smelting furnace operations, and the image quality assessment device 600 for smelting furnace operations includes:
[0211] The feature extraction module 601 is used to obtain an image in front of a smelting furnace, perform smoke and dust feature extraction on the image in front of the smelting furnace to obtain smoke and dust features of the image in front of the smelting furnace, and perform strong light feature extraction on the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace;
[0212] The multi-scale deep feature extraction module 602 is used to perform multi-scale deep feature extraction on the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales;
[0213] The fusion module 603 is used to fuse the smoke feature, the strong light feature and the deep feature for each deep feature to obtain a fused feature, and perform a pre-self-attention operation on the fused feature to obtain a final fused feature;
[0214] The encoding module 604 is used to encode each final fusion feature to obtain an encoding feature corresponding to each final fusion feature, and to splice all the encoding features to obtain a spliced feature;
[0215] The image quality assessment module 605 is used to perform image quality assessment based on the splicing features to obtain an image quality assessment result of the smelting furnace front image.
[0216] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0217] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0218] like Figure 7 As shown, an embodiment of the present application provides a terminal device, and the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above method embodiments when executing the computer program D102.
[0219] Specifically, when the processor D100 executes the computer program D102, it obtains the image in front of the smelting furnace, extracts the smoke and dust features of the image in front of the smelting furnace, and obtains the smoke features of the image in front of the smelting furnace. It also extracts the strong light features of the image in front of the smelting furnace to obtain the strong light features of the image in front of the smelting furnace. Then, it extracts multi-scale deep features of the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales. Then, for each deep feature, it fuses the smoke features, strong light features and deep features to obtain a fused feature. It performs a pre-self-attention operation on the fused feature to obtain a final fused feature. Then, it encodes each final fused feature to obtain a coding feature corresponding to each final fused feature. It splices all the coding features to obtain a spliced feature. Finally, it performs image quality assessment based on the spliced feature to obtain an image quality assessment result of the image in front of the smelting furnace. Among them, the smoke and strong light features of the image in front of the smelting furnace are extracted, which can capture the visual noise features such as smoke and light, perform multi-scale feature extraction on the image in front of the smelting furnace, and express the feature information of the image at multiple scales. The image quality of the image in front of the smelting furnace is evaluated based on the smoke and strong light features and deep features at multiple scales, which improves the information richness of the image quality evaluation and effectively improves the accuracy of the image quality evaluation.
[0220] The processor D100 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0221] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card, etc. equipped on the terminal device D10. Furthermore, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0222] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0223] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0224] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the image quality assessment method device / terminal device for smelting furnace operation, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0225] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0226] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0227] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for evaluating image quality of smelting furnace operations, characterized in that: include: Acquire an image in front of a smelting furnace, perform smoke and dust feature extraction on the image in front of the smelting furnace to obtain smoke and dust features of the image in front of the smelting furnace, and perform strong light feature extraction on the image in front of the smelting furnace to obtain strong light features of the image in front of the smelting furnace; Performing multi-scale deep feature extraction on the image in front of the smelting furnace to obtain deep features of the image in front of the smelting furnace at multiple scales; For each deep feature, the smoke feature, the strong light feature and the deep feature are fused to obtain a fused feature, and a pre-self-attention operation is performed on the fused feature to obtain a final fused feature; Encode each final fusion feature to obtain the encoding feature corresponding to each final fusion feature, and splice all the encoding features to obtain the spliced feature; An image quality assessment is performed based on the stitching features to obtain an image quality assessment result of the smelting furnace front image.
2. The image quality assessment method according to claim 1, wherein: The extracting of smoke features from the image in front of the smelting furnace to obtain smoke features includes: By formula: Calculating the smoke feature T(x) of pixel x in the image in front of the smelting furnace; Wherein, ω represents the smoke residue coefficient, α represents the balance weight, θ represents the linear mapping parameter of the color attenuation prior, A represents the preset global atmospheric light coefficient, D(x) represents the pixel value of pixel x in the dark channel image of the image in front of the smelting furnace, x∈X, X represents the number of all pixels in the image in front of the smelting furnace, and C(x) represents the depth of field model of the image in front of the smelting furnace: C(x)=I V (x)-I S (x) Among them, Ω(x) represents the local window centered at pixel x, R represents the red channel, G represents the green channel, B represents the blue channel, and I V (x) and I s (x) represents the pixel value of pixel x in V and S channels after the image in front of the smelting furnace is converted into HSV space, I c (y) represents the pixel value of pixel y in the image in front of the smelting furnace under channel c; The smoke features of all pixels are integrated to obtain the smoke feature T of the image in front of the smelting furnace.
3. The image quality assessment method according to claim 1, wherein: The step of extracting the strong light feature of the image in front of the smelting furnace to obtain the strong light feature includes: By formula: Calculating the strong light feature S(x) of pixel x in the image in front of the smelting furnace; Among them, m s (x) represents the mirror coefficient, m b (x) represents the diffuse reflectance: Among them, R(x), G(x) and B(x) represent the pixel values of pixel x in the R, G and B channels in the image in front of the smelting furnace, respectively. s Indicates the chromaticity of the light source of the red channel, g s Indicates the chromaticity of the light source of the green channel, b s Represents the chromaticity of the light source of the blue channel, r b Indicates the diffuse chromaticity of the object in the red channel, g b Indicates the diffuse reflection chromaticity of the green channel object, b b Represents the diffuse chromaticity of the object in the blue channel; The strong light features of all pixels are integrated to obtain the strong light feature S of the image in front of the smelting furnace.
4. The image quality assessment method according to claim 1, wherein: The multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales, including: Performing multi-scale adjustment on the smelting furnace front image to obtain an adjusted smelting furnace front image at each scale; For each scale, deep features are extracted from the smelting furnace front image adjusted at the scale to obtain deep features at the scale.
5. The image quality assessment method according to claim 4, wherein: The deep feature extraction is performed on the smelting furnace front image adjusted at the scale to obtain the deep features at the scale, including: By formula: Calculate the deep features at the i-th scale in, It represents the initial features of the smelting furnace image adjusted at the i-th scale after convolution and maximum pooling operations. represents the secondary convolution feature, represents the cubic convolution feature, represents the cubic convolution feature, Both represent convolution, Represents a downsampling operation.
6. The image quality assessment method according to claim 1, wherein: The fusing of the smoke feature, the strong light feature and the deep feature to obtain a fused feature includes: splicing the smoke feature and the strong light feature to obtain an apparent feature; The surface features and the deep features are fused to obtain fused features.
7. The image quality assessment method according to claim 6, wherein: The step of combining the smoke feature and the strong light feature to obtain an apparent feature includes: By formula: Calculate the apparent feature F h ; Among them, concat represents the splicing operation, T represents the smoke feature, S represents the strong light feature, R 2×H×W Indicates the dimension of the apparent feature, H represents the height of the image in front of the smelting furnace, and W represents the width of the image in front of the smelting furnace; The fusing of the surface features and the deep features to obtain fused features includes: By formula: Calculate the fusion features at the i-th scale Among them, reshape represents the size adjustment operation, represents the deep features at the i-th scale, R 2050×(H / 32)×(W / 32) Dimensions that represent apparent features.
8. The image quality assessment method according to claim 1, wherein: The splicing of all coding features to obtain splicing features includes: All coding features are spliced together to obtain the initial spliced features; Adding an embedding sequence before the initial splicing feature to obtain an embedded feature, performing multi-layer encoding on the embedded feature to obtain a final encoded feature; Using a feedforward network to perform linear transformation and nonlinear activation on the final encoding features to obtain feedforward features; Perform a residual connection on the feedforward feature and the encoding feature to obtain a splicing feature.
9. The image quality assessment method according to claim 8, wherein: The image quality assessment result is an image quality score; The image quality assessment is performed based on the stitching features to obtain an image quality assessment result of the smelting furnace front image, including: By formula: score=Linear2(GELU(Linear1(v))) Calculate the image quality score; Among them, Linear1 and Linear2 represent linear mappings, GELU represents the GELU activation function, and v represents the embedded sequence in the splicing feature.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the image quality assessment method for smelting furnace operations according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Low-visibility image defogging model method
CN113962878A
Display screen picture quality detection method
CN114119591A
Multi-scale lightweight smoke image segmentation method and device
CN116503726A
Remote sensing image space spectrum fusion method based on de-noising diffusion probability model and Transform
CN119205523A
Ironmaking product quality evaluation method and device based on multi-source data
CN119624970A