A method for image quality assessment of a smelting furnace front operation and related equipment

By utilizing image quality assessment methods at the smelting furnace front work site, and employing feature extraction of smoke and strong light as well as multi-scale deep feature fusion, the problem of low accuracy in image quality assessment at the smelting furnace front work site has been solved, achieving higher assessment accuracy.

CN120689320BActive Publication Date: 2026-07-21CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2025-06-17
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing no-reference image quality assessment algorithms suffer from low accuracy in image data from smelting furnace operations due to complex and variable visual noise interference such as smoke and strong light.

Method used

By acquiring images in front of the smelting furnace, features of smoke and dust and strong light are extracted. Combined with multi-scale deep feature extraction, feature fusion and pre-self-attention operation are performed. Finally, the images are encoded and stitched together to obtain the image quality assessment results.

Benefits of technology

This improved the information richness and accuracy of image quality assessment, effectively enhancing the accuracy of image quality assessment in front of the smelting furnace.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689320B_ABST
    Figure CN120689320B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image quality evaluation, and provides an image quality evaluation method for smelting furnace front operation and related equipment, which comprises the following steps: smoke dust feature extraction is performed on a smelting furnace front image to obtain a smoke dust feature, and strong light feature extraction is performed on the smelting furnace front image to obtain a strong light feature; multi-scale deep feature extraction is performed on the smelting furnace front image to obtain deep features at multiple scales; the smoke dust feature, the strong light feature and the deep features are fused to obtain fused features, and pre-self-attention operation is performed on the fused features to obtain final fused features; each final fused feature is encoded to obtain an encoding feature corresponding to each final fused feature, and all the encoding features are spliced to obtain a spliced feature; image quality evaluation is performed based on the spliced feature to obtain an image quality evaluation result of the smelting furnace front image. The method can improve the accuracy of image quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image quality assessment technology, and in particular to an image quality assessment method and related equipment for smelting furnace front operations. Background Technology

[0002] With the development of intelligent manufacturing, industrial robots are gradually replacing manual operations in heavy industries such as metallurgy, and automation and intelligence have become the future development trend of the metallurgical industry. This transformation improves production efficiency on the one hand, and allows workers to stay away from dusty and high-temperature environments on the other.

[0003] The intelligent operation of robots in front of smelting furnaces is inseparable from machine vision. Industrial robots need to acquire the positional information of targets and the status information of objects such as smelting furnaces from vision systems. However, the variable visual interference noise in the metallurgical industry, such as dust and lighting, can affect the performance of machine vision. Assessing the level of visual noise interference is the first step in subsequent processing. Therefore, how to accurately assess the image quality of the work site images acquired by industrial camera vision systems is an important issue for the intelligentization of metallurgical operation robots.

[0004] In the field of image quality assessment, early research mainly involved subjective evaluation methods, where humans, as observers, subjectively evaluated the quality of images. With the development of computer vision technology, objective evaluation methods using algorithms to assess image quality have become a research focus. Based on whether or not a reference image is relied upon, objective image quality assessment methods can be divided into three categories: Full Reference (FR), Reduced Reference (RR), and No Reference (NR). Since obtaining reference images in reality is difficult, no-reference image quality assessment algorithms are more widely used in practice. However, existing no-reference image quality assessment algorithms are mostly focused on the quality assessment of everyday or field images. Image data from smelting furnace operation sites contain complex and variable visual noise interference such as smoke, dust, and strong light, resulting in low accuracy in image quality assessment of smelting furnace images. Summary of the Invention

[0005] This application provides an image quality assessment method and related equipment for smelting furnace front operations, which can solve the problem of low accuracy in image quality assessment of smelting furnace front images.

[0006] In a first aspect, embodiments of this application provide an image quality assessment method for smelting furnace front operations, the image quality assessment method comprising:

[0007] The image in front of the smelting furnace is acquired, and the smoke and dust features of the image are extracted to obtain the smoke and dust features of the image in front of the smelting furnace. The strong light features of the image in front of the smelting furnace are also extracted to obtain the strong light features of the image in front of the smelting furnace.

[0008] Multi-scale deep feature extraction was performed on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales.

[0009] For each deep feature, the smoke feature, strong light feature and deep feature are fused to obtain the fused feature, and the fused feature is subjected to a pre-self-attention operation to obtain the final fused feature;

[0010] Encode each final fused feature to obtain the encoded feature corresponding to each final fused feature, and concatenate all the encoded features to obtain the concatenated feature;

[0011] Image quality assessment is performed based on stitching features to obtain the image quality assessment results of the image in front of the smelting furnace.

[0012] Optionally, the smoke and dust features are extracted from the image in front of the smelting furnace to obtain the smoke and dust features, including:

[0013] Through the formula:

[0014]

[0015] Calculate the smoke and dust features T(x) of pixel x in the image in front of the smelting furnace;

[0016] Where ω represents the residual smoke coefficient, α represents the balance weight, θ represents the linear mapping parameter of the color attenuation prior, A represents the global atmospheric light coefficient, D(x) represents the pixel value of pixel x in the dark channel image of the image in front of the smelting furnace, x∈X, X represents the number of all pixels in the image in front of the smelting furnace, and C(x) represents the depth model of the image in front of the smelting furnace.

[0017]

[0018] C(x)=I V (x)-I S (x)

[0019] Where Ω(x) represents a local window centered at pixel x, R represents the red channel, G represents the green channel, B represents the blue channel, and I represents the green channel. V (x) and I s (x) represent the pixel values ​​of pixel x in the V and S channels after converting the image in front of the smelting furnace to HSV space, respectively. c (y) represents the pixel value of pixel y in channel c in the image in front of the smelting furnace;

[0020] By integrating the smoke and dust features of all pixels, the smoke and dust features T of the image in front of the smelting furnace are obtained.

[0021] Optionally, the strong light features are extracted from the image in front of the smelting furnace to obtain strong light features, including:

[0022] Through the formula:

[0023]

[0024] Calculate the strong light feature S(x) of pixel x in the image in front of the smelting furnace;

[0025] Where, m s (x) represents the mirror coefficient, m b (x) represents the diffuse reflectance coefficient:

[0026]

[0027] Where R(x), G(x), and B(x) represent the pixel values ​​of pixel x in the R, G, and B channels of the image in front of the smelting furnace, respectively, and r s g represents the chromaticity of the light source in the red channel. s The colorimetric value of the light source in the green channel, b s The chromaticity of the light source in the blue channel, r b g represents the diffuse chromaticity of an object in the red channel. b b represents the diffuse chromaticity of an object in the green channel. b The diffuse colorimetric hue of an object representing the blue channel;

[0028] By integrating the strong light features of all pixels, the strong light features S of the image in front of the smelting furnace are obtained.

[0029] Optionally, multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain the deep features of the image at multiple scales, including:

[0030] Multi-scale adjustments were made to the images in front of the smelting furnace to obtain adjusted images of the smelting furnace at each scale.

[0031] For each scale, deep features are extracted from the scale-adjusted images in front of the smelting furnace to obtain the deep features at each scale.

[0032] Optionally, deep feature extraction is performed on the scale-adjusted image in front of the smelting furnace to obtain the scale-adjusted deep features, including:

[0033] Through the formula:

[0034]

[0035]

[0036] Calculate deep features at the i-th scale

[0037] in, This represents the initial features obtained by performing convolution and max pooling operations on the adjusted image in front of the smelting furnace at the i-th scale. Represents the features of quadratic convolution. Represents cubic convolution features. Represents cubic convolution features. Both represent convolution, f ds This indicates a downsampling operation.

[0038] Optionally, smoke features, strong light features, and deep features are fused to obtain fused features, including:

[0039] The appearance features are obtained by splicing together the characteristics of smoke and dust and the characteristics of strong light.

[0040] By fusing apparent features and deep features, a fused feature is obtained.

[0041] Optionally, the smoke and dust features and the strong light features are stitched together to obtain the apparent features, including:

[0042] Through the formula:

[0043]

[0044] Calculate apparent features F h ;

[0045] Where concat represents the stitching operation, T represents the smoke and dust feature, S represents the strong light feature, and R represents the strong light feature. 2×H×W The dimensions represent the apparent features, where H represents the height of the image in front of the smelting furnace and W represents the width of the image in front of the smelting furnace.

[0046] By fusing apparent features and deep features, we obtain fused features, including:

[0047] Through the formula:

[0048]

[0049] Calculate the fusion features at the i-th scale

[0050] Here, reshape represents a size adjustment operation. R represents the deep features at the i-th scale. 2050 ×(H / 32)×(W / 32) Dimensions representing apparent features.

[0051] Optionally, all encoded features are concatenated to obtain concatenated features, including:

[0052] All encoded features are concatenated to obtain the initial concatenated features;

[0053] An embedding sequence is added before the initial concatenation features to obtain embedded features. The embedded features are then encoded in multiple layers to obtain the final encoded features.

[0054] The feedforward features are obtained by applying linear transformation and nonlinear activation to the final encoded features using a feedforward network.

[0055] Residual concatenation is performed on the feedforward features and the encoded features to obtain the concatenated features.

[0056] Optionally, the image quality assessment result is an image quality score;

[0057] Image quality assessment based on stitching features yields the following results for the image in front of the smelting furnace:

[0058] Through the formula:

[0059] score=Linear2(GELU(Linear1(v)))

[0060] Calculate the image quality score;

[0061] Where Linear1 and Linear2 represent linear mappings, GELU represents the GELU activation function, and v represents the embedded sequence in the spliced ​​features.

[0062] Secondly, embodiments of this application provide an image quality assessment device for smelting furnace front operations, comprising:

[0063] The feature extraction module is used to acquire images in front of the smelting furnace, extract the smoke and dust features from the images in front of the smelting furnace to obtain the smoke and dust features of the images in front of the smelting furnace, and extract the strong light features from the images in front of the smelting furnace to obtain the strong light features of the images in front of the smelting furnace.

[0064] The multi-scale deep feature extraction module is used to extract multi-scale deep features from the image in front of the smelting furnace, and obtain the deep features of the image in front of the smelting furnace at multiple scales.

[0065] The fusion module is used to fuse smoke features, strong light features and deep features for each deep feature separately to obtain fused features, and to perform pre-self-attention operation on the fused features to obtain the final fused features;

[0066] The encoding module is used to encode each final fused feature to obtain the encoded feature corresponding to each final fused feature, and to concatenate all the encoded features to obtain the concatenated feature;

[0067] The image quality assessment module is used to perform image quality assessment based on stitching features, and obtain the image quality assessment results of the image in front of the smelting furnace.

[0068] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the image quality assessment method for smelting furnace front operations described above.

[0069] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image quality assessment method for smelting furnace front operations described above.

[0070] The above-mentioned solution in this application has the following beneficial effects:

[0071] In some embodiments of this application, by acquiring an image in front of a smelting furnace, the smoke and dust features of the image are extracted to obtain the smoke and dust features of the image in front of the smelting furnace, and the strong light features of the image in front of the smelting furnace are extracted to obtain the strong light features of the image in front of the smelting furnace. Then, multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales. Then, for each deep feature, the smoke and dust features, strong light features and deep features are fused to obtain fused features. Then, a pre-self-attention operation is performed on the fused features to obtain the final fused features. Then, each final fused feature is encoded to obtain the corresponding encoded features. Then, all encoded features are concatenated to obtain concatenated features. Finally, image quality assessment is performed based on the concatenated features to obtain the image quality assessment result of the image in front of the smelting furnace. Among them, extracting smoke and light features from images in front of the smelting furnace can capture visual noise features such as smoke and light. Multi-scale feature extraction of images in front of the smelting furnace enables the expression of feature information of the image at multiple scales. Based on smoke and light features, strong light features, and deep features at multiple scales, image quality assessment of images in front of the smelting furnace is performed, which improves the information richness of image quality assessment and thus effectively improves the accuracy of image quality assessment.

[0072] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0073] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0074] Figure 1 A flowchart illustrating an image quality assessment method for smelting furnace front operations provided in an embodiment of this application;

[0075] Figure 2 This is a schematic diagram of smoke and dust characteristics provided in an embodiment of this application;

[0076] Figure 3 This is a schematic diagram of the strong light characteristics provided in an embodiment of this application;

[0077] Figure 4 This is a schematic diagram illustrating the image quality assessment results of a simulation environment image provided in an embodiment of this application.

[0078] Figure 5 This is a schematic diagram illustrating the image quality assessment results of a real-world image provided in an embodiment of this application.

[0079] Figure 6 A schematic diagram of the structure of an image quality assessment device for smelting furnace front operations provided in an embodiment of this application;

[0080] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0081] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0082] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0083] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0084] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0085] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0086] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0087] To address the issue of low accuracy in image quality assessment of existing smelting furnace front images, this application provides an image quality assessment method for smelting furnace front operations. This method involves acquiring an image of the smelting furnace front, extracting smoke and dust features from the image, extracting strong light features from the image, and then performing multi-scale deep feature extraction to obtain deep features at multiple scales. For each deep feature, the smoke and dust features, strong light features, and deep features are then fused to obtain fused features. A pre-self-attention operation is performed on the fused features to obtain the final fused features. Each final fused feature is then encoded to obtain corresponding encoded features. All encoded features are then concatenated to obtain concatenated features. Finally, image quality assessment is performed based on the concatenated features to obtain the image quality assessment result for the smelting furnace front image. Among them, extracting smoke and light features from images in front of the smelting furnace can capture visual noise features such as smoke and light. Multi-scale feature extraction of images in front of the smelting furnace enables the expression of feature information of the image at multiple scales. Based on smoke and light features, strong light features, and deep features at multiple scales, image quality assessment of images in front of the smelting furnace is performed, which improves the information richness of image quality assessment and thus effectively improves the accuracy of image quality assessment.

[0088] The image quality assessment method for smelting furnace front operations provided in this application will be illustrated by example below.

[0089] like Figure 1 As shown, the image quality assessment method for smelting furnace front operations provided in this application includes the following steps:

[0090] Step 11: Obtain the image in front of the smelting furnace, extract the smoke and dust features from the image in front of the smelting furnace to obtain the smoke and dust features of the image in front of the smelting furnace, and extract the strong light features from the image in front of the smelting furnace to obtain the strong light features of the image in front of the smelting furnace.

[0091] The above images of the smelting furnace are from the smelting furnace operation site.

[0092] In some embodiments of this application, a camera mounted on a metallurgical robot operating in front of a smelting furnace can be used to acquire images of the smelting furnace operation site. The steps of extracting smoke and dust features from the images of the smelting furnace to obtain smoke and dust features, and extracting strong light features from the images of the smelting furnace to obtain strong light features, include:

[0093] The first step is to extract the smoke and dust features from the image in front of the smelting furnace to obtain the smoke and dust features of the image in front of the smelting furnace.

[0094] First, through the formula:

[0095]

[0096] Calculate the smoke and dust features T(x) of pixel x in the image in front of the smelting furnace.

[0097] Where ω represents the smoke residue coefficient, α represents the balance weight, θ represents the linear mapping parameter of the color attenuation prior, A represents the preset global atmospheric light coefficient, used to describe the overall bias of the smoke ambient light on the image brightness, D(x) represents the pixel value of pixel x in the dark channel image of the image in front of the smelting furnace, x∈X, X represents the number of all pixels in the image in front of the smelting furnace, and C(x) represents the depth model of the image in front of the smelting furnace.

[0098]

[0099] C(x)=I V (x)-I S (x)

[0100] Where Ω(x) represents a local window centered at pixel x, R represents the red channel, G represents the green channel, B represents the blue channel, and I represents the green channel. V (x) and I s (x) represent the pixel values ​​of pixel x in the V and S channels after converting the image in front of the smelting furnace to HSV space, respectively. c (y) represents the pixel value of pixel y in channel c in the image in front of the smelting furnace.

[0101] Then, the smoke and dust features of all pixels are integrated to obtain the smoke and dust features T of the image in front of the smelting furnace.

[0102] Specifically, the smoke and dust features of all pixels are integrated into one image according to the position of the pixels in the image in front of the smelting furnace, thus obtaining the smoke and dust features of the image in front of the smelting furnace.

[0103] The second step is to extract the strong light features from the image in front of the smelting furnace to obtain the strong light features of the image in front of the smelting furnace.

[0104] First, through the formula:

[0105]

[0106] Calculate the strong light feature S(x) of pixel x in the image in front of the smelting furnace.

[0107] Where, m s (x) represents the mirror coefficient, m b (x) represents the diffuse reflectance coefficient:

[0108]

[0109] Where R(x), G(x), and B(x) represent the pixel values ​​of pixel x in the R, G, and B channels of the image in front of the smelting furnace, respectively, and r s g represents the chromaticity of the light source in the red channel. s The colorimetric value of the light source in the green channel, b s The chromaticity of the light source in the blue channel, r b g represents the diffuse chromaticity of an object in the red channel. b b represents the diffuse chromaticity of an object in the green channel. b This represents the diffuse color of an object in the blue channel.

[0110] Then, the strong light features of all pixels are integrated to obtain the strong light features S of the image in front of the smelting furnace.

[0111] Specifically, the strong light features of all pixels are integrated into one image according to the position of the pixels in the image in front of the smelting furnace, thus obtaining the strong light features of the image in front of the smelting furnace.

[0112] The following example illustrates this step.

[0113] The original image (i.e., the image in front of the smelting furnace) and the smoke and dust inspection image (i.e., the smoke and dust characteristics) are as follows: Figure 2 As shown, Figure 2 'a' represents the original image. Figure 2 b is the smoke and dust inspection image. The original image and the strong light inspection image (i.e., strong light features) are as follows: Figure 3 As shown, Figure 3 'a' represents the original image. Figure 3 b is the image from the strong light test.

[0114] Step 12: Perform multi-scale deep feature extraction on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales.

[0115] In some embodiments of this application, the steps of performing multi-scale deep feature extraction on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales include:

[0116] The first step is to perform multi-scale adjustments on the image in front of the smelting furnace to obtain the adjusted image in front of the smelting furnace at each scale.

[0117] The above scales refer to the image resolution. Adjusting the resolution of the image in front of the smelting furnace to each scale yields the adjusted image in front of the smelting furnace at each scale.

[0118] It should be noted that the scale of the original image in front of the smelting furnace can also be used as a scale, in which case the image in front of the smelting furnace can be directly used as the adjusted image in front of the smelting furnace at that scale.

[0119] For example, there are three scales: the original scale, the first scale, and the second scale. The original image of the furnace is... H represents the image height, and W represents the image width. The adjusted images of the two smelting furnaces are as follows: H1 represents the image height at the first scale, W1 represents the image width at the first scale, H2 represents the image height at the second scale, and W2 represents the image width at the second scale, in order to obtain... For example, let's set the width of the adjusted image to s1, the scaling factor to s1 / W, and the calculated H1 to be H*(s1 / W). This method can maintain the original aspect ratio. Simultaneously, the original image in front of the smelting furnace is used as the adjusted image, resulting in adjusted images of the smelting furnace at three scales:

[0120] The second step is to extract deep features from the scale-adjusted images in front of the smelting furnace for each scale, thereby obtaining the deep features at each scale.

[0121] Specifically, through the formula:

[0122]

[0123] Calculate deep features at the i-th scale

[0124] in, This represents the initial features obtained by performing convolution and max pooling operations on the adjusted image in front of the smelting furnace at the i-th scale. Represents the features of quadratic convolution. Represents cubic convolution features. Represents cubic convolution features. Both represent convolution, f ds This indicates a downsampling operation with a step size of 2.

[0125] It should be noted that the number of channels and the convolution kernel are different between the above convolutions. The superscript indicates the number of channels in the convolution, and the subscript indicates the convolution kernel. It is a 1×1 convolution with 64 channels.

[0126] For example, with For example, the image can be input into a normalization layer, then processed through a 7×7 convolution (64 channels, stride of 2), and finally processed through a max pooling layer to obtain the initial features. The dimension of this initial feature is The final dimension of the deep features is

[0127] Step 13: For each deep feature, fuse the smoke feature, strong light feature and deep feature to obtain the fused feature, and perform a pre-self-attention operation on the fused feature to obtain the final fused feature.

[0128] In some embodiments of this application, the steps of fusing smoke features, strong light features, and deep features to obtain fused features, and performing a pre-self-attention operation on the fused features to obtain the final fused features include:

[0129] The first step is to stitch together the smoke and dust features and the strong light features to obtain the apparent features.

[0130] Specifically, through the formula:

[0131]

[0132] Calculate apparent features F h .

[0133] Where concat represents the stitching operation, T represents the smoke and dust feature, S represents the strong light feature, and R represents the strong light feature. 2×H×W The dimensions represent the apparent features, where H represents the height of the image in front of the smelting furnace and W represents the width of the image in front of the smelting furnace.

[0134] The second step is to fuse the apparent features and deep features to obtain the fused features.

[0135] Specifically, through the formula:

[0136]

[0137] Calculate the fusion features at the i-th scale

[0138] Here, reshape represents a size adjustment operation. R represents the deep features at the i-th scale. 2050 ×(H / 32)×(W / 32) Dimensions representing apparent features.

[0139] The third step is to perform a pre-self-attention operation on the fused features to obtain the final fused features.

[0140] Specifically, the fused features are flattened to obtain a flattened feature map. Then, attention is calculated on the flattened feature map to obtain a weighted feature map. The weighted feature map and the flattened feature map are then residually connected to obtain a residually connected feature map. The residually connected feature map is then recalibrated to obtain a recalibrated feature sequence. Finally, it is restored to the original feature map spatial dimension to obtain the final fused features.

[0141] For example, the input feature map (i.e., the fused features) is first flattened to... For example, it can be abbreviated as C represents the number of feature map channels, H F and W F These represent the height and width of the feature map, respectively.

[0142] The flattened features are Generate query Q using three fully connected layers F Key K F Sum V F :

[0143] Q F =W q ·F flat

[0144] K F =W k ·F flat

[0145] V F =W v ·F flat

[0146] in, This is a learnable parameter matrix used to map the input feature map to the attention matrix space; thus, we obtain...

[0147] Calculate attention weights:

[0148]

[0149] in, Softmax is a normalization function that makes the sum of the weights equal to 1.

[0150] Using attention weights on V F Perform a weighted summation:

[0151] Y F =A·V F

[0152] in, This represents the weighted feature map.

[0153] In order to preserve the information of the original feature map, the obtained Y F With feature map F flat Perform residual connections:

[0154] Z = Y F +F flat

[0155] in, This is the feature map after residual connection.

[0156] To enhance the model's response to important feature channels, the obtained feature maps are recalibrated using the SE channel attention module. The specific implementation process is as follows:

[0157] For Z in sequence dimension H F W F Perform global average pooling to obtain channel statistics:

[0158]

[0159] Here, Z[i,:] represents each row vector of Z.

[0160] Then, s is transformed through two fully connected layers to obtain the channel weights:

[0161] w = σ(W2·ReLU(W1s))

[0162] Where ReLU is the ReLU activation function, σ is the Sigmoid activation function, and W1 and W2 represent the parameters of the two fully connected layers, respectively.

[0163] Then, each channel in Z is recalibrated using the aforementioned channel weights:

[0164]

[0165] in, This represents the recalibrated feature sequence.

[0166] Finally, Restore the original feature map spatial dimensions from the sequence format:

[0167]

[0168] in, The reshape operation is a feature map adjustment operation used to map the sequence dimensions to the original feature map space dimensions.

[0169] Step 14: Encode each final fused feature to obtain the encoded feature corresponding to each final fused feature, and concatenate all the encoded features to obtain the concatenated feature.

[0170] In some embodiments of this application, the steps of encoding each final fusion feature to obtain the encoded feature corresponding to each final fusion feature, and concatenating all the encoded features to obtain the concatenated feature include:

[0171] The first step is to encode each final fused feature to obtain the encoded feature corresponding to each final fused feature.

[0172] Specifically, scale embedding is first added to the final fused features, and then sinusoidal position encoding is added to obtain the encoded features.

[0173] Exemplary, learnable scale embedding The size of S is H S ×W S Consistent with the size of the embedded feature map, the scale embedding method is pixel-wise addition, adding scale embedding to the feature corresponding to each pixel in the final fused feature. The expression for the above sinusoidal position encoding is:

[0174]

[0175] Where PE is the obtained positional encoding, pos is the position index, i is the dimension index, and d is the total dimension of the embedding vector. The sinusoidal positional encoding is added pixel-by-pixel, adding a sinusoidal positional encoding to the feature corresponding to each pixel in the final fused feature.

[0176] The second step is to concatenate all the encoded features to obtain the initial concatenated features.

[0177] Specifically, each encoded feature is flattened, and then all the flattened encoded features are concatenated to obtain the initial concatenated features.

[0178] For example, the features at each scale are flattened to... For example, the features after flattening are: The features flattened at multiple scales are concatenated along the sequence dimension, i.e., the HW dimension, to obtain the initial concatenated features. N total This represents the total length of the initial splicing feature.

[0179] The third step is to add an embedding sequence before the initial splicing features to obtain the embedded features, and then perform multi-layer encoding on the embedded features to obtain the final encoded features.

[0180] The aforementioned embedding sequence is used for quality score prediction. The embedding sequence is learnable and reflects the feature information within the feature containing the embedding sequence. When processing features such as those containing the embedding sequence through fusion or concatenation, the information contained in the embedding sequence is learned as the feature changes. It is typically a pre-defined token sequence, such as the token sequence cls token in the Transformer architecture. For example, the embedding sequence added for concatenating features is a vector. It is added at the very beginning of the spliced ​​features, and during training, it interacts with other sequences in the spliced ​​features through attention to learn the summary of the overall features.

[0181] For example, multiple stacked coding layers can be used to compute the embedded features to obtain the final encoded features. The multiple coding layers are connected sequentially, and the operation of each coding layer is as follows:

[0182] The input data's Q, K, and V are obtained through linear transformation, and multi-head self-attention is then calculated.

[0183]

[0184] Where, d k For each head in multi-head attention, there is a dimension.

[0185] The above calculation results, after residual connection and layer normalization, are used as the output of the current coding layer. The output after stacking multiple layers is:

[0186] The fourth step involves using a feedforward network to perform linear transformations and nonlinear activations on the final encoded features to obtain feedforward features.

[0187] For example, nonlinear activation can be achieved using the GELU function, and linear transformation can be achieved using fully connected layers.

[0188] The fifth step is to perform residual connection on the feedforward features and the encoded features to obtain the spliced ​​features.

[0189] Step 15: Based on the stitching features, perform image quality assessment to obtain the image quality assessment results of the image in front of the smelting furnace.

[0190] The above image quality assessment results are image quality scores. The higher the image quality score, the better the quality of the image in front of the smelting furnace; the lower the image quality score, the worse the quality of the image in front of the smelting furnace.

[0191] Specifically, through the formula:

[0192] score=Linear2(GELU(Linear1(v)))

[0193] Calculate the image quality score.

[0194] Where Linear1 and Linear2 represent linear mappings, GELU represents the GELU activation function, and v represents the embedded sequence in the spliced ​​features.

[0195] It should be noted that, through the feature processing in step 14, the embedded sequence added therein will be updated and learned with each step of feature processing, and finally capture and integrate the feature information in the spliced ​​features, and provide the feature information of the spliced ​​features for image quality score prediction.

[0196] It is worth mentioning that extracting the smoke and light features from the image in front of the smelting furnace can capture visual noise features such as smoke and light. Multi-scale feature extraction of the image in front of the smelting furnace enables the expression of feature information of the image at multiple scales. Based on the smoke and light features, the deep features at multiple scales, and the deep features of the image in front of the smelting furnace, the image quality is evaluated, which improves the information richness of the image quality evaluation and thus effectively improves the accuracy of the image quality evaluation.

[0197] The method of this application will be illustrated below with a specific example.

[0198] Comparative experiments were conducted using several existing methods and the method of this application. Spearman Rank Correlation Coefficient (SRCC) and Pearson Linear Correlation Coefficient (PLCC) were selected as performance indicators. The experimental results are shown in Table 1.

[0199] Table 1

[0200] BRISQUE 0.665 0.681 ILNIQE 0.507 0.523 HOSA 0.671 0.694 WaDIQaM 0.797 0.805 PQR 0.880 0.884 DBCNN 0.875 0.884 MetaIQA 0.850 0.887 BIQA 0.906 0.917 MUSIQ 0.905 0.919 The method of this application 0.920 0.939

[0201] Among them, BRISQUE is a Blind / Referenceless Image Spatial Quality Evaluator, ILNIQE is an ImageLocalizedNaturalness Image Quality Evaluator, HOSA is a Higher Order Statistics-based Image Quality Assessment, WaDIQaM is a Wavelet Domain Image Quality Assessment Method, PQR is a Perceptual Quality Rating Method, DBCNN is a Deep Bilateral Convolutional Neural Network, MetaIQA is a MetaImage Quality Assessment Method, BIQA is a Blind Image Quality Assessment Method, and MUSIQ is a Multiscale Image Quality Assessment Method.

[0202] The results of the ablation experiment conducted on the method of this application are shown in Table 2.

[0203] Table 2

[0204]

[0205] Therefore, it can be seen that the accuracy of image quality assessment using the method of this application is higher than that of other existing methods.

[0206] The results of quality assessment of images in a simulated environment using the method of this application are as follows: Figure 4 As shown, Figure 4 When 'a' represents the Gaussian noise of the image (Gauss_noise_sigma = 0), the image quality score Iqa = 0.6216. Figure 4 When b represents the Gaussian noise of the image (Gauss_noise_sigma = 5), the image quality score Iqa = 0.5750. Figure 4 When c represents the Gaussian noise of the image (Gauss_noise_sigma = 10), the image quality score Iqa = 0.4728. Figure 4 When d represents the Gaussian noise of the image (Gauss_noise_sigma = 15), the image quality score Iqa = 0.4203.

[0207] The results of quality assessment of images in a real-world working environment using the method described in this application are as follows: Figure 5 As shown, Figure 5 'a' represents the image quality score of the first image taken under real-world conditions: Iqa = 0.3148. Figure 5 'a' represents the image quality score of the second image taken under real-world conditions: Iqa = 0.3174. Figure 5 a represents the image quality score of the third image under real-world conditions, Iqa = 0.4197.

[0208] Therefore, it can be seen that the method of this application can accurately assess the image quality of images in both simulated and real-world operating environments.

[0209] The image quality assessment device for smelting furnace front operations provided in this application will be described exemplarily below.

[0210] like Figure 6 As shown, this application embodiment provides an image quality assessment device for smelting furnace front operations. The image quality assessment device 600 for smelting furnace front operations includes:

[0211] The feature extraction module 601 is used to acquire the image in front of the smelting furnace, extract the smoke and dust features from the image in front of the smelting furnace to obtain the smoke and dust features of the image in front of the smelting furnace, and extract the strong light features from the image in front of the smelting furnace to obtain the strong light features of the image in front of the smelting furnace.

[0212] The multi-scale deep feature extraction module 602 is used to extract multi-scale deep features from the image in front of the smelting furnace, and obtain the deep features of the image in front of the smelting furnace at multiple scales.

[0213] The fusion module 603 is used to fuse the smoke feature, the strong light feature and the deep feature for each deep feature to obtain the fused feature, and to perform a pre-self-attention operation on the fused feature to obtain the final fused feature;

[0214] The encoding module 604 is used to encode each final fused feature to obtain the encoded feature corresponding to each final fused feature, and to concatenate all the encoded features to obtain the concatenated feature;

[0215] The image quality assessment module 605 is used to perform image quality assessment based on stitching features to obtain the image quality assessment results of the image in front of the smelting furnace.

[0216] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0217] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0218] like Figure 7 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, which, when executing the computer program D102, implements the steps in any of the above method embodiments.

[0219] Specifically, when the processor D100 executes the computer program D102, it acquires an image in front of the smelting furnace, extracts smoke and dust features from the image, obtains smoke and dust features, extracts strong light features from the image, obtains strong light features, then performs multi-scale deep feature extraction on the image, obtains deep features at multiple scales, then fuses the smoke and dust features, strong light features, and deep features for each deep feature, obtains fused features, performs pre-self-attention operation on the fused features, obtains final fused features, then encodes each final fused feature, obtains corresponding encoded features, and concatenates all encoded features to obtain concatenated features. Finally, it performs image quality assessment based on the concatenated features to obtain the image quality assessment result of the image in front of the smelting furnace. Among them, extracting smoke and light features from images in front of the smelting furnace can capture visual noise features such as smoke and light. Multi-scale feature extraction of images in front of the smelting furnace enables the expression of feature information of the image at multiple scales. Based on smoke and light features, strong light features, and deep features at multiple scales, image quality assessment of images in front of the smelting furnace is performed, which improves the information richness of image quality assessment and thus effectively improves the accuracy of image quality assessment.

[0220] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0221] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0222] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0223] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0224] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an image quality assessment method device / terminal device for operations in front of a smelting furnace, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0225] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0226] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0227] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for image quality assessment of operations in front of a smelting furnace, characterized in that, include: A smelting furnace front image is acquired, and the smoke and dust features of the smelting furnace front image are extracted to obtain the smoke and dust features of the smelting furnace front image. The strong light features of the smelting furnace front image are also extracted to obtain the strong light features of the smelting furnace front image. Multi-scale deep feature extraction is performed on the image in front of the smelting furnace to obtain the deep features of the image in front of the smelting furnace at multiple scales; For each deep feature, the smoke feature, the strong light feature, and the deep feature are fused to obtain a fused feature, and a pre-self-attention operation is performed on the fused feature to obtain the final fused feature; Encode each final fused feature to obtain the encoded feature corresponding to each final fused feature, and concatenate all the encoded features to obtain the concatenated feature; Image quality assessment is performed based on the stitching features to obtain the image quality assessment result of the image in front of the smelting furnace. The smoke and dust features, the strong light features, and the deep features are fused to obtain fused features, including: The smoke and dust features and the strong light features are spliced ​​together to obtain the appearance features; The apparent features and the deep features are fused to obtain the fused features; The concatenation of all encoded features to obtain concatenated features includes: All encoded features are concatenated to obtain the initial concatenated features; An embedding sequence is added before the initial splicing feature to obtain the embedding feature. The embedding feature is then encoded in multiple layers to obtain the final encoded feature. The final encoded features are obtained by performing linear transformation and nonlinear activation on a feedforward network. The feedforward features and the encoded features are residually concatenated to obtain the spliced ​​features.

2. The image quality assessment method according to claim 1, characterized in that, The step of extracting smoke and dust features from the image in front of the smelting furnace to obtain smoke and dust features includes: Through the formula: Calculate the pixels in the image in front of the smelting furnace Smoke characteristics ; in, Indicates the residual smoke and dust coefficient. Indicates the balancing weights. The linear mapping parameters represent the color attenuation prior. This represents the preset global atmospheric light coefficient. The pixels in the dark channel image representing the image in front of the smelting furnace pixel values, , This represents the number of all pixels in the image in front of the smelting furnace. The depth-of-field model representing the image in front of the smelting furnace is as follows: in, Represented in pixels A local window centered on the center. Indicates the red channel. Indicates a green channel. Indicates the blue channel. These represent the pixels in the V and S channels of the image before the smelting furnace after being converted to HSV space. pixel values, Represents the pixels in the image in front of the smelting furnace In the passage The pixel value below; By integrating the smoke and dust features of all pixels, the smoke and dust features of the image in front of the smelting furnace are obtained. .

3. The image quality assessment method according to claim 1, characterized in that, The step of extracting strong light features from the image in front of the smelting furnace to obtain strong light features includes: Through the formula: Calculate the pixels in the image in front of the smelting furnace strong light characteristics ; in, Indicates the mirror coefficient. Indicates the diffuse reflectance coefficient: in, , and These represent the pixels in the image in front of the smelting furnace. In the pixel values ​​of the R, G, and B channels, This represents the chromaticity of the light source in the red channel. Indicates the chromaticity of the light source in the green channel. This represents the chromaticity of the light source in the blue channel. Represents the diffuse chromaticity of an object in the red channel. The diffuse chromaticity of an object representing the green channel. The diffuse colorimetric hue of an object representing the blue channel; The strong light features of all pixels are integrated to obtain the strong light features of the image in front of the smelting furnace. .

4. The image quality assessment method according to claim 1, characterized in that, The step of performing multi-scale deep feature extraction on the image in front of the smelting furnace to obtain the deep features of the image at multiple scales includes: The image in front of the smelting furnace is adjusted at multiple scales to obtain the adjusted image in front of the smelting furnace at each scale. For each scale, deep feature extraction is performed on the adjusted furnace front image at the scale to obtain the deep features at the scale.

5. The image quality assessment method according to claim 4, characterized in that, The process of extracting deep features from the scale-adjusted image in front of the smelting furnace to obtain deep features at the scale includes: Through the formula: Calculate the first Deep features at various scales ; in, Indicates the first The initial features are obtained by performing convolution and max pooling operations on the smelting furnace front image adjusted at each scale. Represents the features of quadratic convolution. Represents cubic convolution features. Represents cubic convolution features. , , , , , , , , , Both represent convolution. This indicates a downsampling operation.

6. The image quality assessment method according to claim 1, characterized in that, The process of stitching together the smoke and dust features and the intense light features to obtain the apparent features includes: Through the formula: Calculate apparent features ; in, This indicates a splicing operation. Indicates the characteristics of smoke and dust. Indicates the characteristics of strong light. Dimensions representing apparent features This indicates the height of the image in front of the smelting furnace. Indicates the width of the image in front of the smelting furnace; The process of fusing the apparent features and the deep features to obtain the fused features includes: Through the formula: Calculate the first Fusion features at various scales ; in, This indicates a size adjustment operation. Indicates the first Deep features at various scales Dimensions representing apparent features.

7. The image quality assessment method according to claim 1, characterized in that, The image quality assessment result is the image quality score; The image quality assessment based on the stitching features, to obtain the image quality assessment result of the image in front of the smelting furnace, includes: Through the formula: Calculate image quality score ; in, , Represents a linear mapping. This represents the GELU activation function. This represents the embedded sequence in the splicing features.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image quality assessment method for smelting furnace front operations as described in any one of claims 1 to 7.