Non-ferrous metal rolling process quality control method based on machine vision

By combining a multi-angle structured light group and a high-frequency linear scan camera, efficient identification and real-time control of defects in the rolling process of non-ferrous metals are achieved, solving the image processing problems in high reflectivity and complex environments, and improving detection accuracy and stability.

CN121883485APending Publication Date: 2026-04-17LIANYUNGANG DEYAO MASCH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIANYUNGANG DEYAO MASCH TECH CO LTD
Filing Date
2026-03-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In the rolling process of non-ferrous metals, the high reflectivity and complex environmental noise interference make it difficult for traditional image processing methods to accurately identify surface defects, especially fine scratches, roll marks and dot defects.

Method used

A multi-angle structured light group is used in conjunction with a high-frequency linear scan camera. By controlling the alternating exposure of the structured light through synchronous trigger pulses, multi-source image sequences are acquired and image decomposition is performed. Combined with adaptive masking and multi-scale fusion, a feature pyramid network and deformable convolution are used to generate candidate detection boxes for defects, and the process status is judged in real time and control commands are fed back.

Benefits of technology

It effectively suppresses high reflectivity interference, overcomes oil mist noise, improves the accuracy and stability of defect identification, and ensures real-time high-precision quality control of the calendering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883485A_ABST
    Figure CN121883485A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a non-ferrous metal calendaring process quality control method based on machine vision, which comprises the following steps: adopting multi-angle structured light to cooperate with a line-scan digital camera for alternate exposure, and acquiring an aligned multi-source image sequence through synchronous trigger pulse and pixel offset compensation; deconstructing the sequence into diffuse reflection and specular reflection components by using an image decomposition algorithm, and performing multi-scale fusion in combination with an adaptive mask to reconstruct a defect texture base map; inputting the base map into the feature pyramid network, and shielding oil mist and water vapor interference through adjacent frame residual analysis; defect candidate frames are adaptively generated by using deformable convolution, and overlapping is eliminated based on confidence coefficient smooth conversion. According to the method, high-frequency phase distortion information is extracted from a light reflection principle, and a time-space cooperative residual analysis and adaptive geometric perception mechanism is introduced, so that high reflection of nonferrous metals and noise interference of a complex production environment are effectively inhibited, and closed-loop control of a process state is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a machine vision-based method for quality control of non-ferrous metal rolling processes. Background Technology

[0002] The rolling of non-ferrous metals (such as aluminum foil and copper strip) is a core process in the metal processing industry. With technological advancements, utilizing machine vision technology to replace manual surface defect detection has become a mainstream trend for automating quality control. However, in actual high-speed non-ferrous metal rolling production lines, the highly dynamic environment and the inherent physical properties of the materials pose significant challenges to image processing. Non-ferrous metal surfaces typically have extremely high specular reflectivity; under illumination, they are prone to severe "halo" or "overexposure" phenomena, causing defect features in the image to be masked by environmental noise. This high reflectivity makes it difficult for traditional grayscale thresholding or simple edge detection operators to suppress false defects while maintaining high sensitivity. Simultaneously, the rolling mill environment is often accompanied by high-speed flowing oil mist, water vapor, and mechanical vibration. These factors lead to motion blur, extremely low contrast, and non-uniform occlusion in the acquired images, severely interfering with the accurate identification of minute scratches, roll marks, and point defects by machine vision systems. Therefore, designing a method to effectively suppress high reflectivity interference and overcome oil mist noise for quality control in the non-ferrous metal rolling environment is a key technical problem that urgently needs to be solved in this field. To address this, a machine vision-based method for quality control in non-ferrous metal rolling processes is proposed. Summary of the Invention

[0003] The purpose of this invention is to provide a machine vision-based method for quality control of non-ferrous metal rolling processes, in order to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: Machine vision-based quality control methods for non-ferrous metal rolling processes include: A multi-angle structured light group combined with a high-frequency linear array scanning camera is used to image the surface of non-ferrous metal strips in high-speed rolling. By controlling the alternating exposure of structured light at different incident angles through synchronous trigger pulses, a multi-source image sequence containing geometric morphology and surface reflectivity characteristics is obtained. The multi-source temporal image sequence is decomposed into diffuse reflection component and specular reflection component using an image decomposition algorithm. High-frequency phase distortion information in the specular reflection component is extracted, and the high-frequency phase distortion information and the diffuse reflection component are fused at multiple scales using adaptive mask processing to reconstruct a single-frame defect texture base map. The defect texture base map of consecutive frames is input into the feature pyramid network. The residual analysis mechanism between the feature maps of adjacent frames is used to identify and shield the background interference of oil mist and water vapor, extract the defect features, and input the defect features into deformable convolution to generate defect candidate detection boxes. The candidate detection boxes are subjected to overlap elimination processing based on confidence smoothing transformation to accurately locate the defect position and extract the edge contour. The current rolling process status is judged in real time, and control commands are generated and fed back based on the judgment results.

[0005] Preferably, the specific process of alternating exposure of structured light at different incident angles by controlling the synchronous trigger pulse includes: The pulse signal from the encoder of the calendering production line is acquired in real time. The row trigger cycle of the linear scan camera is calculated based on the real-time running speed of the strip and used as the master control frequency for global synchronization. The trigger pulses generated by the master control frequency are sequentially distributed to different output channels, with each channel corresponding to a set of structured light sources with a specific incident angle. When each row trigger pulse arrives, the pulse controller drives the light source of one of the channels to perform instantaneous strobe, while simultaneously triggering the linear scan camera. Combining the physical distribution spacing of each set of structured light sources in the calendering direction and the real-time calendering speed, pixel offset compensation is performed on the row images acquired under different lighting conditions. The row images acquired from different channels are recombined into a multi-source image sequence with spatially aligned positions.

[0006] Preferably, the image decomposition algorithm is based on the multi-source image sequence and uses a sub-pixel registration algorithm to perform spatial alignment to eliminate physical displacement deviation; compare the brightness distribution of aligned pixels under different incident angles, and use sorted median filtering to extract the background signal common to each channel as the initial estimate of the diffuse reflection component; The original image sequence is differentially processed with the initial estimate to obtain a residual image. A guided filter operator is used to retain the reflective features in the residual image, guided by the original image. Structural tensor analysis is performed on the reflective features. Gradient vector correction and logical superposition are performed on the bright spots under different channels in combination with the incident geometry angle of the structured light. The gradient deviation field reflecting the micro-geometric deformation of the surface is extracted as the high-frequency phase distortion information. The multi-source image sequence is deconstructed into diffuse reflection components and specular reflection components.

[0007] Preferably, the specific process of multi-scale fusion of the high-frequency phase distortion information and the diffuse reflection component includes constructing a dual-stream convolutional neural network to simultaneously extract multi-scale feature maps of the diffuse reflection component and the high-frequency phase distortion information at different spatial resolutions; performing saliency analysis on the feature maps of the high-frequency phase distortion information using an attention mechanism to generate an adaptive mask; weighting the phase distortion feature maps of the corresponding scales pixel by pixel using the adaptive mask; performing multi-level cross-channel stitching of the weighted phase distortion features and the feature maps of the diffuse reflection component; performing nonlinear feature integration using convolutional layers; and reconstructing a single-frame defect texture base map by upsampling level by level through a decoding network.

[0008] Preferably, the dual-stream convolutional neural network has a heterogeneously configured background texture branch and geometric feature branch. The background texture branch uses large-scale convolutional kernels or dilated convolutions to extract globally consistent texture features of the diffuse reflection component, while the geometric feature branch uses high-resolution residual structures to capture local geometric gradient features of the high-frequency phase distortion information. During the multi-scale feature extraction process of the two branches, a lateral interactive connection module is introduced between feature layers of the same resolution to achieve displacement alignment correction of feature spatial positions, and a hierarchical feature map containing different spatial resolutions is output synchronously.

[0009] Preferably, the feature pyramid network is a multi-scale representation structure composed of multiple feature layers with different resolutions; The defect texture base map of consecutive frames is input into each feature level of the feature pyramid network. The feature maps of the same level in adjacent frames are spatially shifted and aligned using real-time rolling speed, and the residual of the aligned feature map is calculated. Based on the spatial morphological features of the residual, the feature responses of oil mist and water vapor regions are identified and shielded. The shielded features of each level are upsampled and fused from top to bottom and supplemented with lateral information to extract defect features.

[0010] Preferably, the specific process of generating defect candidate detection boxes using deformable convolution is as follows: the extracted defect features are input into a deformable convolutional layer; the offset prediction branch is used to learn the offset vector of the local geometric deformation of the defect, and a sampling offset matrix that adaptively matches the topological shape of the defect is generated; the sampling position of the convolutional kernel on the feature map is adjusted according to the sampling offset matrix, so that the receptive field of the convolutional kernel is deformed to cover the defect region with irregular edges, and the deformed and enhanced defect feature map is output; multi-scale regression calculation is performed on the deformed and enhanced defect feature map to simultaneously predict the category confidence score and bounding box coordinate offset value of the defect target, and the initial defect candidate detection boxes are generated by filtering according to the confidence threshold.

[0011] Preferably, the overlap elimination process based on confidence smoothing transformation involves sorting all the candidate detection boxes from high to low according to their initial confidence, and selecting the detection box with the highest current ranking as the reference box. The overlap between the reference box and the other candidate detection boxes is calculated using a continuous decay function, and the confidence of the other candidate detection boxes is updated in real time with reduced weight based on the overlap. The confidence ranking of each candidate detection box is updated through multiple rounds of iteration. The center position and coverage area of ​​the defect are determined according to the distribution law of the ranked confidence. The edge contour of the defect is extracted based on the determined area, and the current rolling process status is output in real time.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention changes the traditional single-source imaging mode by alternating exposure and synchronous triggering mechanism of multi-angle structured light groups. It uses image decomposition algorithm to deconstruct the original sequence into diffuse reflection and specular reflection components with clear physical meaning. This processing method can extract high-frequency phase distortion information reflecting the micro-deformation of the surface from the principle of light reflection and fuse it with diffuse reflection texture at multiple scales to reconstruct a defect texture base map containing rich geometric details, providing more distinguishable underlying data input.

[0013] 2. This invention introduces a residual analysis mechanism for adjacent frames based on real-time velocity compensation into the feature pyramid network, achieving a dimensional leap from spatial positioning to time series. By calculating the differences in the aligned feature maps, it can utilize the different characteristics of dynamic backgrounds such as oil mist and water vapor and metal surface defects in the temporal evolution to identify and shield environmental interference at the feature level, thereby making the extracted defect features purer and ensuring the logical rigor of the detection process in complex physical environments.

[0014] 3. This invention achieves adaptive matching of the detection operator to the topological shape of the defect by employing deformable convolution and overlap elimination processing based on confidence-smooth transformation. It uses offset prediction branches to adjust the sampling position of the convolution kernel, enabling the receptive field to dynamically deform with the edge contour of the irregular defect. Combined with a smooth confidence-weighted update mechanism, it can more accurately lock the defect center and extract the edge contour, providing high-precision geometric representation data for real-time closed-loop discrimination of the rolling process status. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating a machine vision-based quality control method for non-ferrous metal rolling processes. Figure 2 This is a schematic diagram of the closed-loop control logic of the present invention, from the acquisition of the original signal to the final generation of control commands; Figure 3 This is a diagram of the dual-stream network branch architecture of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: Please see Figure 1 and Figure 2 This invention provides a machine vision-based method for quality control in non-ferrous metal rolling processes, the technical solution of which is as follows: Machine vision-based quality control methods for non-ferrous metal rolling processes include: A multi-angle structured light group combined with a high-frequency linear array scanning camera is used to image the surface of non-ferrous metal strips in high-speed rolling. By controlling the alternating exposure of structured light at different incident angles through synchronous trigger pulses, a multi-source image sequence containing geometric morphology and surface reflectivity characteristics is obtained. The multi-source temporal image sequence is decomposed into diffuse reflection component and specular reflection component using an image decomposition algorithm. High-frequency phase distortion information in the specular reflection component is extracted, and the high-frequency phase distortion information and the diffuse reflection component are fused at multiple scales using adaptive mask processing to reconstruct a single-frame defect texture base map. The defect texture base map of consecutive frames is input into the feature pyramid network. The residual analysis mechanism between the feature maps of adjacent frames is used to identify and shield the background interference of oil mist and water vapor, extract the defect features, and input the defect features into deformable convolution to generate defect candidate detection boxes. The candidate detection boxes are subjected to overlap elimination processing based on confidence smoothing transformation to accurately locate the defect position and extract the edge contour. The current rolling process status is judged in real time, and control commands are generated and fed back based on the judgment results.

[0018] The specific process of imaging the surface of non-ferrous metal strip in high-speed rolling using a multi-angle structured light group in conjunction with a high-frequency linear scan camera is as follows: at the inspection station of the rolling production line, the high-frequency linear scan camera is set up directly above the strip in the direction of strip movement, and its scan line spans the entire width of the strip; multiple sets of structured light sources are arranged alternately around the scan line in the forward and backward direction of strip movement, and these light sources are pointed at the scan line at different tilt angles; The multi-angle structured light group comprises three to five line light sources with different incident angles. The selection of the incident angles must be based on the specular reflection characteristics of the material: for materials with extremely high specular reflectivity, such as aluminum foil, the angle configuration should cover tangential incident angles (20° to 40°), oblique incident angles (40° to 60°), and near-positive incident angles (70° to 85°), thereby capturing micro-roughness, geometric deformation, and overall brightness information respectively; the instantaneous flicker pulse width of the light source is strictly limited to within 80% of the single-line exposure time of the camera; The specific process of alternating exposure of structured light at different incident angles by synchronously triggering pulses includes: The pulse signal from the encoder of the calendering production line is acquired in real time. The row trigger cycle of the linear scan camera is calculated based on the real-time running speed of the strip and used as the master control frequency for global synchronization. The preset cycle counting cycle distributes the trigger pulses generated by the master control frequency to different output channels in sequence. Each channel corresponds to a set of structured light sources with a specific incident angle. When each row trigger pulse arrives, the pulse controller drives the light source of one of the channels to flash instantaneously, while triggering the linear scan camera. Combining the physical distribution spacing of each set of structured light sources in the calendering direction and the real-time calendering speed, pixel offset compensation is performed on the row images acquired under different lighting conditions. The row images acquired by different channels are recombined into a multi-source image sequence with spatially aligned positions. A rotary encoder is installed on the shaft of the drive roller or on the measuring roller that is in direct contact with the strip. When the strip moves at high speed with the calender, the rotary encoder synchronously generates pulse signals with a fixed phase difference. The central controller collects these pulse signals in real time and calculates the current precise running speed of the strip based on the preset number of pulses per revolution. In order to ensure that the image is not stretched or compressed in the longitudinal direction, the line triggering cycle of the linear scan camera is dynamically calculated based on the real-time speed of the strip and the preset longitudinal resolution. The line triggering cycle serves as the master control frequency for global synchronization, ensuring that each line acquisition action of the camera precisely corresponds to a fixed displacement interval on the surface of the strip. A pre-set counting cycle corresponding to the number of structured light source groups is used. In this embodiment, when three groups of structured lights at different angles are installed, the counting cycle is set to three. After receiving the trigger pulses generated by the main control frequency, the pulse controller sequentially distributes the pulses to different output channels through its internal logic distributor. Each channel is independently connected to a group of structured light sources with a specific incident angle at the hardware level, thereby establishing a one-to-one correspondence between the trigger pulse sequence and the physical light source group. At the instant each line trigger pulse arrives, the pulse controller generates a microsecond-level driving current to excite the light source corresponding to the current counting channel to flash instantaneously. Since the non-ferrous metal surface is extremely sensitive to light, the width of the flash pulse is strictly controlled within the camera's photosensitive time to obtain instantaneous brightness and freeze motion shadows. At the same time as the light source flashes, the same trigger pulse is synchronously sent to the external trigger interface of the line scan camera, enabling the camera to complete the exposure and readout of one line image. In this way, the line scan camera alternately captures the surface reflection information generated by structured light illumination from different angles within a continuous line cycle. Considering the physical distribution spacing of multiple structured light sources in the rolling direction, the center lines of the strip surface illuminated by light sources at different angles do not coincide in space. Based on the pre-calibrated physical distance between each group of light sources and the real-time rolling speed of the strip at the current moment, the relative time offset between image rows at different angles is calculated. Based on the relative time offset, the row image streams acquired by different channels are logically indexed and rearranged. By shifting a specific number of pixel data rows in the memory buffer, dynamic compensation for the physical position difference is achieved. After pixel offset compensation, the image rows of each channel are fully aligned in spatial coordinates. The multi-line images that belong to the same physical scanning position but are acquired at different incident angles are merged and reassembled into a multi-source image sequence with spatial alignment.

[0019] The pulse-synchronized time-division exposure mechanism effectively solves the alignment problem of multi-angle imaging in high-speed rolling; the main control frequency is dynamically bound to the strip speed, eliminating image stretching caused by speed fluctuations and ensuring the stability of detection accuracy; multi-channel instantaneous stroboscopic technology is used to capture features of the same position under different lighting conditions in a very short time, significantly improving the ability to perceive the microstructure of non-ferrous metals; and pixel-level alignment is achieved through a physical spacing compensation algorithm, eliminating displacement deviation caused by asynchronous acquisition.

[0020] The image decomposition algorithm is based on the multi-source image sequence and uses a sub-pixel registration algorithm to perform spatial alignment to eliminate physical displacement deviation; compare the brightness distribution of aligned pixels under different incident angles, and use sorted median filtering to extract the background signal common to each channel as the initial estimate of the diffuse reflection component; The original image sequence is differentially processed with the initial estimate to obtain a residual image. A guided filter operator is used to retain the reflective features in the residual image, guided by the original image. Structural tensor analysis is performed on the reflective features. Gradient vector correction and logical superposition are performed on the bright spots under different channels in combination with the incident geometry angle of the structured light. The gradient deviation field reflecting the micro-geometric deformation of the surface is extracted as the high-frequency phase distortion information. The multi-source image sequence is deconstructed into diffuse reflection components and specular reflection components.

[0021] Because mechanical vibration or speed fine-tuning is inevitable during high-speed calendering of the strip, physical displacement deviations occur in images acquired from different channels. This embodiment employs a sub-pixel registration algorithm. First, the integer pixel displacement between images is determined by calculating the cross-correlation function. Then, within a 3×3 pixel region around the integer pixel extreme point, the continuous extreme center of the correlation energy function is found using a quadratic polynomial fitting method. By solving the zero-point derivative of the fitted surface, the offset vector accurate to 0.05 pixels is calculated. Finally, a trilinear interpolation algorithm is used to resample and reconstruct the non-reference channel image, eliminating physical displacement deviations in the spatial dimension and enabling complete point-to-point overlap of pixels at different angles. For each physical coordinate point, the grayscale value of that point in all acquisition channels is extracted. Due to the strong directionality of specular reflection on metal surfaces, bright spots are usually generated only in one or two specific angle channels, while the background brightness remains stable in other channels. This embodiment uses a median filtering algorithm to sort the grayscale values ​​of pixels in each channel in ascending order of numerical value, and selects the grayscale value in the middle of the sorted sequence as the signal output of that point. Utilizing the physical characteristic that specular reflection points are outliers in grayscale distribution, transient bright interference in each channel is eliminated, and a common background signal that reflects the true color of the strip is extracted as the initial estimate of the diffuse reflection component. The physical logic of the median filtering is that, when processing multi-source image sequences, for each pixel coordinate that is spatially perfectly aligned, the grayscale values ​​of that pixel in all light source channels are extracted to form a vector. These grayscale values ​​are sorted in ascending order of numerical value, and the value in the middle of the sorted sequence is selected as the diffuse reflection estimate of that pixel.

[0022] The original multi-source image sequence is subtracted pixel-by-pixel from the aforementioned initial estimate of diffuse reflection to obtain a residual image containing surface geometric deformation information. To eliminate high-frequency noise during the imaging process and preserve defect edges, a guided filtering operator is used. Using the original image as a reference, the mean and variance of the reference image are calculated within a set local window (a 3×3 pixel area in this embodiment). The residual image is then smoothed by combining the statistical characteristics within this window. The texture structure of the reference image is used to constrain the output of the residual image, ensuring that the reflective features in the residual image do not have blurred contour edges while filtering out noise, thus fully preserving the morphological information of micro-defects. The first-order gradient partial derivatives of the reflective feature image in the horizontal and vertical directions are calculated using the central difference operator. Based on the gradient values ​​in these two directions, a 2×2 symmetric positive definite matrix, i.e., the structure tensor, is constructed for each pixel. On this basis, combined with the preset physical incident geometry angle of the structured light group, gradient vector correction is performed on the bright spots captured by different channels. Through geometric mapping, the brightness gradients generated by light and shadow in different directions are restored to a unified surface normal vector deviation. By logically superimposing the gradient contributions of each channel, the gradient deviation field reflecting the microscopic geometric deformation of the surface is calculated. The extracted gradient deviation field is defined as high-frequency phase distortion information. Through the above decomposition process, the multi-source image sequence is successfully decomposed into diffuse reflection components and specular reflection components. Among them, the diffuse reflection component eliminates reflection interference and reflects the basic texture of the material; the specular reflection component (i.e., the high-frequency phase distortion information) highly sensitively captures the defect features of the rolled surface caused by minute deformation.

[0023] By coupling subpixel registration with sorted median filtering, extremely high-precision spatial alignment was achieved under high-speed vibration environment, and reflective interference was accurately removed, significantly improving the robustness of background extraction. Guided filtering was used to achieve edge-preserving denoising, ensuring high-fidelity retention of micro-defect features. Through structural tensor and gradient vector correction, unstable light and shadow fluctuations were transformed into uniform surface normal vector deviations, achieving in-depth deconstruction of geometric deformation.

[0024] The high-frequency phase distortion information and the diffuse reflection component are fused at multiple scales using adaptive masking to reconstruct a single-frame defect texture base map. The specific process of multi-scale fusion of the high-frequency phase distortion information and the diffuse reflection component includes constructing a dual-stream convolutional neural network to simultaneously extract multi-scale feature maps of the diffuse reflection component and the high-frequency phase distortion information at different spatial resolutions; performing saliency analysis on the feature maps of the high-frequency phase distortion information using an attention mechanism to generate an adaptive mask; weighting the phase distortion feature maps at corresponding scales pixel-by-pixel using the adaptive mask; performing multi-level cross-channel stitching of the weighted phase distortion features and the feature maps of the diffuse reflection component; using convolutional layers for nonlinear feature integration; and using a decoding network to progressively upsample and restore the spatial resolution to reconstruct a single-frame defect texture base map.

[0025] Specifically, the preprocessed diffuse reflection component image and high-frequency phase distortion information map are input into a two-stream convolutional neural network. An attention mechanism is introduced for saliency analysis. For each layer's extracted phase distortion feature map, the contribution weight of each feature channel is calculated through global pooling and fully connected layers. At the same time, a spatial attention module is used to identify regions with abnormal response intensity in the feature map. These regions usually correspond to physical deformations such as scratches, roll marks, or pits on the strip surface. Through this analysis, it is possible to automatically identify which regions' phase information is more valuable for defect determination, thereby generating an adaptive mask with the same size as the feature map. After generating the adaptive mask, it is weighted pixel-by-pixel with the phase distortion feature map of the corresponding scale. In this pixel-by-pixel weighting process, the adaptive mask is a two-dimensional weight matrix generated by an attention mechanism, where the value of each pixel ranges from 0 to 1, representing the confidence that the location belongs to a defect feature. During implementation, this weight matrix is ​​aligned with a high-frequency phase distortion feature map of the same size, and the value at each spatial coordinate point in the feature map is directly multiplied by the weight value at the corresponding coordinate point in the mask. Through this point-by-point multiplication operation, regions in the mask close to 1 retain and amplify the high-frequency details in the phase distortion, while regions in the mask close to 0 (such as normal background textures or flat areas) have their feature values ​​attenuated to near zero, thereby accurately filtering out non-defect-related physical fluctuations from the original signal. The weighted phase distortion feature map is stitched together with the diffuse reflection component feature map of the same layer resolution in a multi-level cross-channel manner. This stitching method forces the luminance information reflecting the surface material properties and the geometric information reflecting the surface micro-deformation to be fused in the channel dimension, forming a fused feature vector with rich dimensions. The multi-level cross-channel stitching ensures that the weighted phase distortion feature map and diffuse reflection component feature map in each level have completely consistent spatial resolution (i.e., equal height and width pixel count). Tensor stitching is used to merge the two sets of features along the channel dimension (depth dimension) of the feature maps. To achieve multi-level fusion, the stitching process is repeated at different spatial scales of the network. During downsampling, corresponding diffuse reflection and phase distortion features are extracted at multiple scale branches at 1 / 2, 1 / 4, and 1 / 8 of the original resolution for the aforementioned channel stitching. The stitching result at each level is passed through a 3×3 convolutional layer responsible for nonlinear dimensionality reduction and feature integration of the stitched multi-dimensional channels, transforming the original features into higher-order defect semantic representations. These fused feature blocks at different scales serve as input to the decoding network, participating in the global resolution restoration process through a skip connection mechanism, ultimately reconstructing a single-frame defect texture base map. By constructing a dual-stream convolutional neural network and combining it with adaptive mask processing, this invention achieves cross-channel deep fusion of material brightness attributes and micro-geometric deformation information. It utilizes a pixel-by-pixel weighting mechanism to accurately amplify weak defect signals and suppress background noise fluctuations at the feature level. Combined with multi-scale stitching and integration technology, it ensures full-scale feature capture from micro-texture to macro-morphology, significantly improving the signal-to-noise ratio and contrast of the defect texture base map, thereby effectively overcoming the interference of complex environments such as oil mist and water vapor at the calendering site.

[0026] See Figure 3 The dual-stream convolutional neural network consists of a heterogeneously configured background texture branch and a geometric feature branch. The background texture branch uses dilated convolution to extract globally consistent texture features of the diffuse reflection component, while the geometric feature branch uses a high-resolution residual structure to capture local geometric gradient features of the high-frequency phase distortion information. During the multi-scale feature extraction process of the two branches, a lateral interactive connection module is introduced between feature layers of the same resolution to achieve displacement alignment correction of the feature space position and synchronously output a hierarchical feature map containing different spatial resolutions. In the background texture branch, a multi-layer dilated convolutional structure is employed to obtain a larger receptive field without sacrificing spatial resolution. Specifically, different dilation rates are set in each convolutional layer (in this embodiment, they cycle in a 1, 2, 4 ratio), allowing the convolutional kernel to cover a wider pixel area in leaps during computation. In this way, the background texture branch can bypass subtle random disturbances on the metal surface and extract globally consistent texture features that reflect the overall material properties of the strip surface. In the geometric feature branch, to ensure that minute defect signals (such as shallow scratches and tiny pits) in high-frequency phase distortion are not filtered out by the downsampling operation, a high-resolution residual structure is employed. By establishing skip connections with identity mappings between adjacent convolutional layers, high-frequency geometric gradient features can be directly transmitted across layers, avoiding signal attenuation in deep networks. During feature extraction, this branch maintains a high feature map resolution, utilizing small 3×3 convolutional kernels to capture phase change gradients in local regions at high frequencies, ensuring that every geometric detail reflecting surface micro-deformation is accurately converted into a feature vector. To address the feature space mismatch issue that may arise from different convolution logics during parallel processing of heterogeneous branches, this embodiment introduces a lateral interaction connection module between corresponding layers (i.e., feature layers with completely identical resolution) of the two branches. The implementation process includes: when the background texture branch and the geometric feature branch output feature maps of the same scale, the spatial positions of the two sets of feature maps are compared using a cross-correlation calculation module; based on the calculated displacement deviation, sub-pixel level spatial transformation correction is performed on the feature maps in the lateral connection channel. This lateral interaction mechanism ensures that the physical positions represented by the features from the two branches are completely overlapped and aligned at the pixel level before entering the subsequent fusion step. Through the collaborative work of the heterogeneous branches and lateral interactive connections described above, the two-stream convolutional neural network can simultaneously output hierarchical feature maps containing different spatial resolutions. During downsampling, each branch simultaneously outputs two sets of spatially aligned feature sequences at 1 / 2, 1 / 4, and 1 / 8 of the original resolution: one set is a texture feature map representing a globally consistent background, and the other set is a geometric gradient feature map representing local fine deformations. These two sets of hierarchical feature maps will serve as input to the attention mechanism, used for subsequent generation of adaptive masks and multi-scale feature fusion to reconstruct a high-quality single-frame defect texture base map. The heterogeneous dual-stream convolutional neural network architecture described in this embodiment achieves differentiated and accurate extraction of global background texture and local micro-geometric features through heterogeneous branch design. It utilizes dilated convolution and residual structure to balance large-scale noise suppression and preservation of small defect signals. The introduction of a lateral interactive connection module eliminates the spatial displacement deviation caused by heterogeneous algorithms through sub-pixel level correction, ensuring seamless integration of multi-source features at the pixel level.

[0027] The defect texture base map of consecutive frames is input into the feature pyramid network. The residual analysis mechanism between the feature maps of adjacent frames is used to identify and shield the background interference of oil mist and water vapor, and extract the defect features. The feature pyramid network is a multi-scale representation structure composed of multiple feature layers with different resolutions. The defect texture base map of consecutive frames is input into each feature level of the feature pyramid network. The feature maps of the same level in adjacent frames are spatially shifted and aligned using real-time rolling speed, and the residual of the aligned feature map is calculated. Based on the spatial morphological features of the residual, the feature responses of oil mist and water vapor regions are identified and shielded. The shielded features of each level are upsampled and fused from top to bottom and supplemented with lateral information to extract defect features.

[0028] The feature pyramid network is a multi-scale representation structure composed of multiple feature layers with different resolutions. The reconstructed single-frame defect texture base maps of consecutive frames (including the current frame and the adjacent frames of the previous time step) are input into the network. Through a bottom-up path, multiple convolutional kernels with a stride of 2 are used to downsample the image, generating a series of feature layer levels with spatial resolution decreasing by a factor of two. During processing, the speed data of the encoder on the rolling production line is retrieved in real time to calculate the displacement of the strip in the rolling direction during the time interval between two adjacent frames. For each feature level with the same resolution in the feature pyramid, this displacement is used to perform spatial displacement alignment compensation on the feature maps of adjacent frames. Specifically, the feature map of the previous frame is translated in the spatial dimension according to the calculated pixel offset value, so that the feature maps of the two frames are precisely superimposed on the same physical coordinate point of the metal strip. Element-by-element subtraction is performed on the aligned feature map of the current frame and the feature map of the previous frame to obtain the feature map residual reflecting the dynamic changes in the time domain. For the calculated feature map residuals, a spatial morphology feature discrimination mechanism is used to identify oil mist and water vapor regions. Since oil mist and water vapor typically exhibit a floating, diffused, and extremely blurred clump-like distribution at the calendering site, they appear in the residual map as low-frequency features with large connected regions but extremely slow gray-level gradient changes. In contrast, real defects on the strip surface exhibit high-frequency features with sharp contours, high contrast, and specific geometric orientations (such as distribution along the calendering direction). In practice, by calculating the ratio of the local gradient variance to the perimeter area of ​​each connected region in the residual map, a dynamic threshold is set to identify interference regions belonging to oil mist and water vapor. This dynamic threshold is set as the sum of the local statistical mean of the residual feature map and the deviation gain based on calendering speed compensation. After identifying the interference region, a corresponding spatial interference mask is generated to mask the feature response values ​​in the current frame feature map that are located within the interference region (i.e., set the response values ​​to 0). Feature fusion is performed using the top-down path of the feature pyramid. The high-level feature map containing high-level semantic information is upsampled by 2 times to align its spatial size with the adjacent low-level feature map, and then horizontally stitched with the masked feature map of the same layer after 1×1 convolution. This deeply combines the global discrimination information of the high-level layer with the local detail information of the low-level layer, while ensuring that the interference signal does not propagate to the lower level. Through multi-level top-down fusion and lateral connections, the feature pyramid network outputs pure features after shielding environmental interference at each resolution level. These features not only eliminate false responses caused by oil mist and water vapor, but also enhance the salience of small defects in complex backgrounds through cross-scale feature completion. The features fused from each level are summarized to extract the final defect feature vector. By employing spatiotemporal collaborative residual analysis and utilizing real-time velocity compensation to accurately remove oil mist and water vapor interference that does not move with the strip, the problem of excessive false responses in complex production environments is solved. Combining dynamic threshold adaptive shielding and feature pyramid multi-scale fusion, diffuse background noise is efficiently removed while the detailed features of high-frequency minute defects are enhanced by supplementing lateral information.

[0029] The defect features are then input into a deformable convolutional layer to generate candidate defect detection boxes. Specifically, the extracted defect features are input into the deformable convolutional layer, and the offset prediction branch is used to learn the offset vector of the local geometric deformation of the defect, generating a sampling offset matrix that adaptively matches the defect topology. The sampling position of the convolutional kernel on the feature map is adjusted according to the sampling offset matrix, causing the receptive field of the convolutional kernel to deform and cover the defect region with irregular edges, outputting a deformed and enhanced defect feature map. Multi-scale regression calculations are performed on the deformed and enhanced defect feature map to simultaneously predict the category confidence score and bounding box coordinate offset value of the defect target, and initial candidate defect detection boxes are generated based on the confidence threshold. The defective feature map is input to a deformable convolutional module, which includes a dedicated offset prediction branch. This branch consists of a set of 3×3 convolutional layers. During execution, the offset prediction branch performs a local context scan on each pixel in the feature map, calculates the trend of geometric deformation around that point, and outputs a sampling offset matrix. The spatial dimensions (rows and columns) of this sampling offset matrix completely correspond to the height and width of the input defective feature map, meaning that each pixel position on the feature map has unique row and column coordinates in the matrix. At each intersection sampling point of the matrix, 18 values ​​are stored. These 18 values ​​are divided into 9 pairs of coordinate offset vectors, each corresponding to one of the 9 original sampling points in the 3×3 convolutional kernel. Each pair of vectors precisely indicates the pixel distance that the sampling point needs to move in the horizontal and vertical directions. After obtaining the offset matrix, an irregular sampling operation is performed. Specifically, when convolving the feature map, the sampling coordinates of the convolution kernel are no longer fixed. Based on the values ​​in the offset matrix, the offset target coordinates are found in the neighborhood of the current pixel. Since the offset coordinates usually do not fall on integer pixels, weighted interpolation is performed using the feature values ​​of the four surrounding pixels to obtain sub-pixel level feature responses. This sampling method allows the receptive field of the convolution kernel to be stretched, rotated, or bent according to the offset vector, so that its coverage area completely matches the irregular edges of defects such as thin scratches and irregular patches, thereby outputting a deformed and enhanced defect feature map. The deformed and enhanced defect feature map is fed into the detection head for multi-scale regression calculation. The detection head uses two parallel 1×1 convolutional branches: the classification branch is responsible for feature activation of each sampling region and calculating its confidence score belonging to a specific defect category, with a value range between 0 and 1; the regression branch is responsible for coordinate correction of the reference anchor frame, calculating four offset correction values ​​for the center point coordinates, width, and height of the bounding box. The coordinate correction process involves obtaining the original horizontal and vertical coordinates, original width, and original height of the preset reference anchor frame. In the translation correction stage, the predicted horizontal and vertical offset ratios of the center point are multiplied by the width and height of the original anchor frame, respectively, to calculate the absolute pixel displacement, which is then accumulated to the original center coordinates to obtain the corrected center point position. In the scale correction stage, the predicted logarithmic scaling ratio of width and height is used as the exponent, and the result of the natural constant power operation with a base of 2.718 is used as the scaling factor and multiplied by the original width and height to complete the nonlinear scale adjustment. Initial candidate boxes are generated by comparing with a confidence threshold. During implementation, predictions with a confidence score below 0.8 are directly discarded, retaining only highly reliable predictions. For each selected target, the four correction values ​​output by its regression branch are applied to the corresponding baseline anchor box. The pixel coordinate range of the defect in the original image is calculated using a coordinate restoration algorithm, ultimately generating a set of initial candidate detection boxes that can tightly wrap around the defect edges. This embodiment achieves adaptive matching between the receptive field and the defect topology shape through deformable convolution, solving the problem that fixed convolution kernels cannot accurately cover irregular edges; combined with sub-pixel interpolation sampling and nonlinear coordinate correction logic, this method effectively eliminates the geometric deviation between the preset anchor frame and the real target.

[0030] The candidate detection boxes are subjected to overlap elimination processing based on confidence level smoothing transformation to accurately locate the defect position and extract the edge contour, determine the current rolling process status in real time, generate control commands based on the determination results and provide feedback; the overlap elimination processing based on confidence level smoothing transformation is to sort all the candidate detection boxes from high to low according to the initial confidence level, and select the detection box with the highest current ranking as the reference box. The overlap between the reference box and the other candidate detection boxes is calculated using a continuous decay function, and the confidence of the other candidate detection boxes is updated in real time with reduced weight based on the overlap. The confidence ranking of each candidate detection box is updated through multiple rounds of iteration. The center position and coverage area of ​​the defect are determined according to the distribution law of the ranked confidence. The edge contour of the defect is extracted based on the determined area, and the current rolling process status is output in real time. Multiple initial candidate detection boxes are aggregated and stored in a set to be processed. For each candidate detection box in the set, its corresponding initial confidence score is extracted, and all detection boxes are globally sorted in descending order of confidence score. The detection box with the highest confidence score in the current sorting result is selected as the reference box. This reference box is regarded as the most definitive defect target in the current region, and it is removed from the set to be processed and stored in the final determined target sequence. After selecting a reference box, the overlap between the reference box and all remaining candidate detection boxes in the set to be processed is calculated. Specifically, the intersection-union algorithm is used to calculate the ratio of the overlap area between the reference box and each of the other detection boxes to the total area of ​​their union. A continuous decay function is used to update the confidence of the remaining detection boxes in the set to be processed in real time by reducing their weight. The continuous decay function is preferably a Gaussian decay function. The logic is that when the overlap between the remaining detection boxes and the reference box is greater, the calculated decay factor is smaller (approaching 0). By multiplying the decay factor by the current confidence of the detection box, the confidence is smoothly reduced, rather than being directly removed. After one round of weight reduction, the bounding boxes in the set to be processed are re-sorted from high to low according to the updated confidence level, and the top-ranked bounding box is selected again as the new reference box. The process of calculating overlap and confidence decay is repeated. Through this multi-round iterative update, the confidence of redundant bounding boxes that highly overlap with the true defect center will smoothly decay to extremely low values ​​as the number of iterations increases. This smooth conversion mechanism avoids the erroneous deletion problem that is prone to occur when the traditional hard thresholding method is used to process adjacent or overlapping defects, ensuring that the most accurate representation box is retained for each independent defect target. The iteration stops when the confidence scores of all remaining detection boxes in the set to be processed are below 0.1. At this point, based on the distribution pattern of each detection box in the final target sequence, the geometric center position and final coverage area of ​​the defect target are locked. Based on the locked area, the edge contour of the defect is accurately extracted on the reconstructed defect texture map, and the geometric features including area, length and aspect ratio are calculated. Geometric quantification analysis is performed on the extracted defect edge contours to calculate the defect's characteristic parameters, including length, width, area, and aspect ratio. These parameters are then input into a pre-defined defect classification decision tree. The classification decision tree has three pre-defined quality standards: a defect less than 5 mm in length and less than 10 square millimeters in area is classified as a Level 1 minor defect; a defect between 5 and 20 mm in length or with continuously distributed minor flaws is classified as a Level 2 warning defect; and a defect exceeding 20 mm in length or exceeding 50 square millimeters in area is classified as a Level 3 severe defect. Through comparison, the system can determine in real time whether the current rolling process is in normal operation, performance drift, or abnormal shutdown. Based on the determined process status, the logic controller generates control commands in real time according to preset mapping rules. If a level two warning defect is identified, indicating a deviation in lubrication or cooling conditions during the current calendering process, the logic controller generates a command to increase the spray pressure, increasing the cooling oil flow by 10% to 20% by adjusting the proportional valve opening to eliminate thermal scratches caused by temperature rise or uneven friction. If a level three severe defect is identified, indicating a possible misalignment of the calender roll gap or foreign object intrusion, the logic controller immediately generates a command to adjust the calender roll gap, driving the servo to perform micron-level gap compensation, and simultaneously triggering an audible and visual warning signal to alert the operator. The generated control commands are fed back to the rolling execution system in real time via an industrial fieldbus. The execution system completes the action response within one control cycle after receiving the command (set to 20 milliseconds in this embodiment), realizing fully automatic online closed-loop control of the quality of non-ferrous metal rolling process.

[0031] The Gaussian attenuation smoothing conversion mechanism solves the problem of missed detection of adjacent defects, ensuring the completeness of the positioning. Combined with the three-level hierarchical decision-making logic, it realizes a fully automatic closed loop from precise feature analysis to micron-level compensation of process parameters, effectively transforming post-processing scrap into online correction, and significantly improving the yield of non-ferrous metal strip and the intelligent control accuracy of the production line.

[0032] This invention addresses key technical bottlenecks in high-speed calendering processes, such as low defect contrast, diverse defect morphologies, and severe interference from oil mist and water vapor. Through multi-source image sequence deconstruction and feature pyramid residual analysis, it achieves precise removal of weak defect signals and effective shielding of environmental noise in extremely complex environments. Utilizing deformable convolution and confidence smoothing mechanisms, it overcomes the challenges of locating irregular topological defects and easily missing overlapping targets. The final closed-loop control architecture of sampling-analysis-discrimination-feedback ensures stable operation and intelligent management of the calendering process.

[0033] Example 2: This embodiment applies a machine vision-based quality control method for non-ferrous metal rolling processes to an aluminum foil finishing mill with a width of 1500mm and a production speed of up to 1200m / min. The inspection station is equipped with a 16K high-frequency linear scanning camera with a pixel clock of 800MHz and a line scanning frequency set to 100kHz to ensure a longitudinal resolution of 0.2mm / line at high speed. The multi-angle structured light group consists of three sets of high-power red linear lasers with a wavelength of 635nm. The incident angles are precisely preset to 30° (tangential light), 45° (oblique light), and 75° (near-normal light) to cover all-around feature capture, from microscopic scratches to macroscopic flatness. A rotary encoder is installed on the unit's steering roller, generating 5000 pulses per revolution. The pulse controller receives the encoder signal and calculates the current main control frequency as 83.3kHz. A three-channel time-division switching logic is used: the first pulse cycle triggers a 30° light source, the second triggers a 45° light source, and the third triggers a 75° light source, repeating cyclically. Considering the light source spacing is 50mm, a buffer logic index is used to perform 250 lines of physical displacement compensation offset on the three-line image sequence to ensure that the overlap error of the synthesized RGB format multi-source image sequence in spatial coordinates is controlled within 0.02 pixels. The preset subpixel registration operator is invoked, and quadratic polynomial fitting is used to align the three-channel images. For the highly reflective points on the aluminum foil surface, sorted median filtering is used to extract the background signal, effectively eliminating approximately 15% of specular outlier noise. When extracting high-frequency phase distortion information, the regularization parameter of the guided filter is set to... This ensures that while filtering out thermal noise from the linear scan camera, the depth is fully preserved. The subtle gradient field at the edge of the roller print.

[0034] The dual-stream network employs a heterogeneous design: the background texture branch consists of four layers of dilated convolutions with a cyclic dilation rate of [1, 2, 4, 8] to obtain a receptive field exceeding 256 pixels and extract global uniformity features; the geometric feature branch uses a hierarchical structure composed of three sets of ResNet-18 residual blocks to maintain the feature map resolution at least 1 / 4 of the original image. The lateral interaction connection module performs feature correction after each residual block using a 1×1 convolution. The attention mechanism module performs saliency weighting on the phase feature map, and the activation threshold for mask generation is set to 0.65 to highlight deformed regions.

[0035] The feature pyramid network contains four feature levels; the frequency of the exhaust fan at the mill exit is read in real time, and the threshold coefficient of the residual analysis is dynamically adjusted. For feature maps of adjacent frames (with a time interval of 1ms), the residual is calculated after spatial alignment using velocity vectors; by setting a discrimination rule that the perimeter-to-area ratio of connected components is greater than 2.5 and the gradient variance is less than 0.15, dynamic oil mist responses with an occlusion area ratio of 20% are automatically identified and blocked. The detection head integrates a deformable convolutional layer, whose offset prediction branch outputs an 18-channel offset matrix (corresponding to 3×3 sampling points). The convolutional kernel adaptively deforms according to the offset vector, enabling its receptive field to perfectly cover longitudinal, slender scratches exceeding 10mm in length. The regression branch presets four reference anchor boxes with different aspect ratios for common defects in aluminum foil; the classification branch sets the category confidence threshold to 0.8, and the number of initially generated candidate boxes is controlled to within 50 per frame to reduce subsequent computational load. Overlap elimination employs a Gaussian decay function with a standard deviation parameter set to 0.5. After multiple iterations, when the overlap exceeds 0.45, the confidence level of redundant boxes rapidly decreases exponentially until it falls below 0.1, at which point they are eliminated. The final determined defect edge contour accuracy reaches 0.1 mm. Based on the hierarchical decision tree, if continuously distributed spots (level 2 defects) are identified, the logic controller will send a signal to the actuator within 20 ms to increase the rolling oil spray pressure by 0.15 MPa. If a perforation (level 3 defect) is identified, a deceleration command is immediately triggered (decelerating to 200 m / min) and the defect coordinates are recorded.

[0036] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for quality control of non-ferrous metal calendering process based on machine vision, characterized in that, include: A multi-angle structured light group combined with a high-frequency linear array scanning camera is used to image the surface of non-ferrous metal strips in high-speed rolling. By controlling the alternating exposure of structured light at different incident angles through synchronous trigger pulses, a multi-source image sequence containing geometric morphology and surface reflectivity characteristics is obtained. The multi-source image sequence is decomposed into diffuse reflection component and specular reflection component using an image decomposition algorithm. High-frequency phase distortion information in the specular reflection component is extracted. The high-frequency phase distortion information and the diffuse reflection component are fused at multiple scales through adaptive mask processing to reconstruct a single-frame defect texture base map. The defect texture base map of consecutive frames is input into the feature pyramid network. The residual analysis mechanism between the feature maps of adjacent frames is used to identify and shield the background interference of oil mist and water vapor, extract the defect features, and input the defect features into deformable convolution to generate defect candidate detection boxes. The candidate detection boxes are subjected to overlap elimination processing based on confidence smoothing transformation to locate the defect position and extract the edge contour. The current rolling process status is judged in real time, and control commands are generated and fed back based on the judgment results.

2. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The specific process of alternating exposure of structured light at different incident angles by synchronously triggering pulses includes: The pulse signal from the encoder of the calendering production line is acquired in real time. The row trigger cycle of the linear scan camera is calculated based on the real-time running speed of the strip and used as the master control frequency for global synchronization. The trigger pulses generated by the master control frequency are sequentially distributed to different output channels, with each channel corresponding to a set of structured light sources with a specific incident angle. When each row trigger pulse arrives, the pulse controller drives the light source of one of the channels to perform instantaneous strobe, while simultaneously triggering the linear scan camera. Combining the physical distribution spacing of each set of structured light sources in the calendering direction and the real-time calendering speed, pixel offset compensation is performed on the row images acquired under different lighting conditions. The row images acquired from different channels are recombined into a multi-source image sequence with spatially aligned positions.

3. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The image decomposition algorithm is based on the multi-source image sequence and uses a sub-pixel registration algorithm to perform spatial alignment to eliminate physical displacement deviation; compare the brightness distribution of aligned pixels under different incident angles, and use sorted median filtering to extract the background signal common to each channel as the initial estimate of the diffuse reflection component; The original image sequence is differentially processed with the initial estimate to obtain the residual image. Then, guided filtering is used to retain the reflective features in the residual image, guided by the original image. Structural tensor analysis is performed on the reflective features. Gradient vector correction and logical superposition are performed on the bright spots under different channels in combination with the incident geometry angle of the structured light. The gradient deviation field reflecting the micro-geometric deformation of the surface is extracted as the high-frequency phase distortion information. The multi-source image sequence is deconstructed into diffuse reflection component and specular reflection component.

4. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The specific process of multi-scale fusion of the high-frequency phase distortion information and the diffuse reflection component includes constructing a two-stream convolutional neural network, simultaneously extracting multi-scale feature maps of the diffuse reflection component and the high-frequency phase distortion information at different spatial resolutions; and using an attention mechanism to perform saliency analysis on the feature map of the high-frequency phase distortion information to generate an adaptive mask. The adaptive mask is used to weight the phase distortion feature map of the corresponding scale pixel by pixel. The weighted high-frequency phase distortion information is then stitched together with the feature map of the diffuse reflection component in a multi-level cross-channel manner. Convolutional layers are used to perform nonlinear integration of features, and the spatial resolution is restored by upsampling at each level through a decoding network to reconstruct a single-frame defect texture base map.

5. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 4, characterized in that, The dual-stream convolutional neural network has a heterogeneously configured background texture branch and geometric feature branch. The background texture branch uses dilated convolution to extract the globally consistent texture features of the diffuse reflection component, while the geometric feature branch uses residual structure to capture the local geometric gradient features of the high-frequency phase distortion information. In the multi-scale feature extraction process of the two branches, the displacement alignment correction of the feature space position is realized by introducing a lateral interactive connection module between feature layers of the same resolution, and the hierarchical feature map containing different spatial resolutions is output synchronously.

6. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The feature pyramid network is a multi-scale representation structure composed of multiple feature layers with different resolutions. The defect texture base map of consecutive frames is input into each feature level of the feature pyramid network. The feature maps of the same level in adjacent frames are spatially shifted and aligned using real-time rolling speed, and the residual of the aligned feature map is calculated. Based on the spatial morphological features of the residual, the feature responses of oil mist and water vapor regions are identified and shielded. The shielded features of each level are upsampled and fused from top to bottom and supplemented with lateral information to extract defect features.

7. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The specific process of generating defect candidate detection boxes using deformable convolution is as follows: the extracted defect features are input into the deformable convolutional layer, the offset prediction branch is used to learn the offset vector of the local geometric deformation of the defect, and a sampling offset matrix is ​​generated that adaptively matches the topological shape of the defect; the sampling position of the convolution kernel on the feature map is adjusted according to the sampling offset matrix, so that the receptive field of the convolution kernel is deformed to cover the defect region with irregular edges, and the deformed and enhanced defect feature map is output. Multi-scale regression calculation is performed on the deformed and enhanced defect feature map to simultaneously predict the category confidence score and bounding box coordinate offset of the defect target, and initial defect candidate detection boxes are generated based on the confidence threshold.

8. The machine vision-based quality control method for non-ferrous metal rolling processes according to claim 1, characterized in that, The overlap elimination process based on confidence smoothing transformation involves sorting all candidate detection boxes from high to low according to their initial confidence, and selecting the detection box with the highest current confidence as the reference box. The overlap between the reference box and the other candidate detection boxes is calculated using a continuous decay function, and the confidence of the other candidate detection boxes is updated in real time with reduced weight based on the overlap. The confidence ranking of each candidate detection box is updated through multiple rounds of iteration. The center position and coverage area of ​​the defect are determined according to the distribution law of the ranked confidence. The edge contour of the defect is extracted based on the determined area, and the current rolling process status is output in real time.