Method and system for positioning a soft zone of a mandrel based on machine vision
By employing a machine vision-based mandrel soft area positioning method, and utilizing a pre-calibrated set of reference parameters and parallax correction technology, high-precision mandrel soft area positioning was achieved. This solved the problems of large positioning errors and insufficient reliability in existing technologies, and improved the production efficiency and quality of cold-rolled strip steel coiling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PRIMETALS TECH (CHINA) LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the positioning methods of mechanical limit and encoder feedback for mandrel soft zone positioning have large positioning errors and cannot adapt to fluctuations in production conditions. Visual inspection solutions are limited by installation conditions and are prone to visual deception, resulting in insufficient strip coiling accuracy and reliability, which affects production efficiency and quality.
A machine vision-based soft area localization method for mandrels is adopted. By acquiring a pre-calibrated set of reference parameters, continuously acquiring images of the mandrel surface, performing target detection and quality screening, and combining physical coordinate mapping and parallax correction, high-precision soft area localization is achieved.
It significantly improves the accuracy and reliability of mandrel soft zone positioning, reduces deviation and tower-shaped defects during strip coiling, and improves production efficiency and product quality.
Smart Images

Figure CN122435019A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial machine vision technology, and in particular relates to a machine vision-based method and system for locating the soft area of a mandrel. Background Technology
[0002] With the iterative upgrading of automation technology in the metallurgical rolling industry, continuous and intelligent production modes are gradually becoming popular in the strip steel processing field, and coiling automation control technology is widely used. This type of technology relies on the coordinated operation of electrical control system and mechanical transmission structure to complete strip steel coiling operations. It has the characteristics of strong production continuity and high industrial adaptability, thus forming a traditional lead-positioning operation method with mechanical limit and electrical control sensing as the core, which has become the mainstream implementation method for soft area positioning of coiling machine core shaft at this stage.
[0003] Currently, the industry mainly uses an indirect positioning method combining mechanical limiters with encoder feedback to complete the positioning of the mandrel soft area. For example, the reference range of the mandrel soft area is calibrated by a mechanical limiter structure, and then the encoder is used to collect parameters such as the travel stroke and rotation speed of the transmission mechanism. Combined with the electronic control logic, the relative position of the strip head is calculated to complete the alignment operation between the strip head and the soft area, meeting the basic strip coiling production requirements. At the same time, some simplified vision inspection solutions are applied to positioning scenarios, relying on basic image capture methods to collect material images and achieve a rough determination of the strip head position.
[0004] However, mechanical positioning structures are susceptible to transmission gaps and component wear due to long-term equipment vibration and material friction, resulting in positioning reference deviations. Furthermore, they cannot adapt to fluctuations in operating conditions during production, leading to significant positioning errors. Encoder-based indirect measurement relies on transmission parameters to deduce position data, resulting in data lag and difficulty in achieving real-time, accurate correction. In addition, conventional visual inspection solutions are limited by industrial installation conditions, with tilted shooting angles. Combined with the inherent height difference between the strip and the mandrel, this easily leads to visual deception, interfering with position determination accuracy. They also cannot automatically correct subsequent production parameters based on positioning deviations, exhibiting poor consistency and stability. This can easily cause the strip head to deviate from the soft zone, resulting in strip surface damage, reduced finished product quality, and even production line downtime, significantly limiting the improvement of strip coiling production efficiency and quality. Summary of the Invention
[0005] Therefore, it is necessary to provide a machine vision-based mandrel soft area positioning method and system to address the above-mentioned technical problems. This method can improve the stability of target detection under dynamic working conditions, so as to meet the requirements of soft area positioning accuracy and reliability in the cold-rolled strip coiling process.
[0006] In a first aspect, this application provides a machine vision-based method for locating the soft area of a mandrel, including:
[0007] Obtain the pre-calibrated set of mandrel soft-area positioning reference parameters; the set of mandrel soft-area positioning reference parameters includes physical coordinate mapping parameters, head-lifting parallax correction parameters, and soft-area reference positioning parameters;
[0008] In response to the arrival signal of the strip head, multiple frames of mandrel surface images are continuously acquired within the mandrel stationary window to obtain the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the start of the mandrel rotation and winding.
[0009] Target detection is performed on each frame of the target image in the image sequence to be detected, and the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge are obtained.
[0010] Based on the physical coordinate mapping parameters, the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge are converted into physical coordinates, and the original positioning deviation and detection quality parameters corresponding to each frame of the target image are calculated.
[0011] Target images with detection quality parameters not lower than the quality threshold are selected as valid frames. The original positioning deviations of the valid frames are weighted and fused according to the detection quality parameters to obtain the weighted average original deviation.
[0012] Based on the obtained current strip thickness parameters and strip head float parallax correction parameters, the weighted average original deviation is corrected for height difference parallax to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
[0013] In one embodiment, obtaining a pre-calibrated set of mandrel soft-area positioning reference parameters includes:
[0014] Acquire multi-pose calibration images of the calibration board on the soft area plane of the mandrel, perform corner detection on the multi-pose calibration images, and solve the camera intrinsic parameters based on the corner pixel coordinates and the known physical coordinates corresponding to the corner pixel coordinates obtained by corner detection, so as to obtain the camera intrinsic parameter matrix and distortion coefficient vector.
[0015] The calibration plate is flatly attached to the cylindrical surface where the soft area of the mandrel is located to acquire a planar calibration image. The pixel coordinates of multiple corner control points and the corresponding physical coordinates of the corner control points are extracted from the planar calibration image. Based on the correspondence between the pixel coordinates and physical coordinates of the corner control points, the homography matrix from the image plane to the physical plane of the soft area is solved by direct linear transformation. The scale conversion coefficient is calculated based on the ratio of the physical distance between the corner control points to the corresponding pixel distance. The homography matrix and the scale conversion coefficient are used as physical coordinate mapping parameters.
[0016] The original deviation value sequence is obtained, and linear fitting is performed based on the original deviation value sequence and the corresponding strip thickness parameters to obtain the strip head floating parallax correction parameter between the strip head floating height and the projection systemic offset; wherein, the original deviation value sequence is measured by the vision system when the physical deviation of the strip head is manually confirmed to be zero under multiple strip thickness specifications.
[0017] Under calibration, the left and right edge pixel coordinates of the soft area of the mandrel are identified, the pixel coordinates of the center line of the soft area are calculated based on the left and right edge pixel coordinates, and the pixel coordinates of the center line of the soft area are mapped to the physical coordinate system through the homography matrix to obtain the soft area reference positioning parameters.
[0018] In one embodiment, target detection is performed on each frame of the target image in the image sequence to be detected, obtaining the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, including:
[0019] Each frame of the target image is converted to a color space, converting the RGB color space to the HSV color space to obtain an HSV image containing hue, saturation, and lightness components.
[0020] Thresholding is performed on the brightness components in the HSV image by setting a preset brightness threshold. Pixels with brightness values greater than the brightness threshold are identified as soft area candidate pixels, and a binary image of the soft area candidate region is obtained.
[0021] Morphological opening operations are performed on the binary image to eliminate scattered noise and retain the main soft region, resulting in a denoised soft region binary image.
[0022] Connected component labeling is performed on the denoised soft region binary image, and the pixel area of each connected component is extracted. The connected component with the largest pixel area is determined as the soft region.
[0023] The soft region is fitted with a minimum bounding rectangle, and the pixel coordinates of the left and right edges of the minimum bounding rectangle are extracted. The average value of the pixel coordinates of the left and right edges is taken to obtain the pixel coordinates of the center line of the soft region.
[0024] In one embodiment, target detection is performed on each frame of the target image in the image sequence to be detected, obtaining the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, including:
[0025] Each frame of the target image is input into the improved YOLO-seg instance segmentation model for forward inference to obtain a pixel-level binary mask with the head; the improved YOLO-seg instance segmentation model is trained on an optimized training dataset constructed by high-precision polygon annotation of the edge region at the junction of the head and the core.
[0026] Contour extraction is performed on the pixel-level binary mask to obtain complete contour data with the header;
[0027] Based on the complete contour data, the position of the foremost edge of the belt head facing the soft area of the mandrel is identified, and the pixel coordinates of the belt head edge are extracted from the foremost edge position.
[0028] Furthermore, the method also includes:
[0029] Multi-condition mandrel surface images under various working conditions were collected as the original image dataset. These images cover different strip thicknesses, different lighting conditions, and the state of the mandrel at different lateral positions of the strip head.
[0030] Each image in the original image dataset is annotated with polygons based on the edge encryption at the intersection, and the complete outline with the head is annotated to obtain a high-precision polygon annotation mask;
[0031] We use a nano-level YOLO-seg network as the basic architecture, set the number of output categories to 1 to predict the leading region, and divide the original image dataset after high-precision polygon labeling masking into training set and validation set.
[0032] The YOLO-seg network is trained on the training set. During training, a weighted sum of the binary cross-entropy loss function and the Dice loss function is used as the segmentation loss. Training is stopped when the validation set loss no longer decreases for 5 consecutive epochs, resulting in an improved YOLO-seg segmentation model. The improved YOLO-seg segmentation model is then exported as a TensorRTFP16 inference engine file for deployment on the GPU memory of the vision system server.
[0033] In one embodiment, the weighted average original deviation is corrected for height difference parallax based on the obtained current strip thickness parameters and strip head float parallax correction parameters to obtain the mandrel soft zone positioning deviation, including:
[0034] The following formula is used to correct the weighted average original deviation by subtracting the height difference parallax, based on the current strip thickness parameters, strip head float parallax correction parameters, and physical coordinate mapping parameters, to obtain the mandrel soft area positioning deviation:
[0035] ;
[0036] in, This is due to the positioning deviation of the mandrel soft area. The weighted average of the original deviations, Parallax correction parameters for levitation. The current strip thickness parameters are as follows: These are the scale transformation coefficients in the physical coordinate mapping parameters.
[0037] Secondly, this application also provides a machine vision-based mandrel soft area localization system for use with the machine vision-based mandrel soft area localization method provided in the first aspect, the system comprising:
[0038] The reference parameter calibration module is used to obtain a pre-calibrated set of reference parameters for mandrel soft-area positioning; the set of reference parameters for mandrel soft-area positioning includes physical coordinate mapping parameters, parallax correction parameters for head-mounted floating, and soft-area reference positioning parameters;
[0039] The image acquisition module is used to continuously acquire multiple frames of mandrel surface images within the mandrel stationary window in response to the arrival signal of the strip head, so as to obtain the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the start of the mandrel rotation and winding.
[0040] The target detection module is used to perform target detection on each frame of the image sequence to be detected, and obtain the pixel coordinates of the soft area center line and the pixel coordinates of the leading edge.
[0041] The coordinate transformation module is used to convert the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge into physical coordinates based on the physical coordinate mapping parameters, and to calculate the original positioning deviation and detection quality parameters corresponding to each frame of the target image;
[0042] The image fusion module is used to filter target images whose detection quality parameters are not lower than the quality threshold as valid frames, and to perform weighted fusion on the original positioning deviation of the valid frames according to the detection quality parameters to obtain the weighted average original deviation.
[0043] The parallax correction module is used to perform height difference parallax correction on the weighted average original deviation based on the obtained current strip thickness parameters and strip head float parallax correction parameters, so as to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
[0044] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the machine vision-based mandrel soft area positioning method as provided in the first aspect.
[0045] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the machine vision-based mandrel soft area localization method as provided in the first aspect.
[0046] The aforementioned machine vision-based mandrel soft-area localization method and system establishes a unified and accurate measurement benchmark for the entire localization process by acquiring a set of benchmark parameters for mandrel soft-area localization, including physical coordinate mapping parameters, strip head levitation parallax correction parameters, and soft-area benchmark localization parameters. Based on this benchmark, in response to the arrival signal of the strip head, multiple frames of mandrel surface images are continuously acquired within a static window from when the strip head adheres to the mandrel until the mandrel begins to rotate, forming a sequence of images to be detected. This fundamentally avoids the errors caused by motion compensation in traditional dynamic detection schemes. Subsequently, target detection is performed on each frame of the target image to obtain the pixel coordinates of the soft-area centerline and the strip head edge. Based on the physical coordinate mapping parameters, these coordinates are converted into physical coordinates, and the original localization deviation and surface deviation for each frame are calculated. The system employs detection quality parameters to assess the reliability of single-frame detection. Reliable valid frames are then selected based on these parameters, and the original positioning deviations of these valid frames are weighted and fused according to these parameters. This effectively suppresses the impact of random interference on the detection results. Finally, the weighted average original deviation is corrected for height difference parallax by combining the real-time acquired strip thickness parameters and strip head float parallax correction parameters. This eliminates systematic measurement errors caused by strip head float, ultimately yielding accurate mandrel soft-zone positioning deviations that instruct the PLC to adjust the strip's lateral position. This significantly improves the accuracy and reliability of mandrel soft-zone positioning, effectively reducing deviation and tower-shaped defects during cold-rolled strip coiling, and enhancing production efficiency and product quality. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 A schematic flowchart of the machine vision-based mandrel soft area localization method provided in an embodiment of the present invention;
[0049] Figure 2 This is a schematic diagram of the structure of the mandrel soft area positioning system based on machine vision provided in an embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0051] First, a brief introduction to the terms used in the embodiments of this application will be given.
[0052] In the coiling machine of a cold-rolled strip steel production line, the mandrel is the core component used for coiling the strip steel. It is typically a cylindrical expansion mechanism. During coiling, the strip head (strip tip) must accurately conform to a specific area on the mandrel surface. The mandrel surface usually has a soft zone to provide cushioning support for the strip head at the beginning of coiling, preventing edge damage or plastic deformation caused by rigid contact. The strip tip refers to the leading edge of the strip, which is the first part to contact the mandrel during threading. In the coiling process, the spatial position of the strip tip determines the accuracy of the coiling starting point. Because the strip tip may float during transport (creating a height difference with the mandrel surface), visual inspection must consider measurement deviations caused by perspective projection. If the strip tip fails to fall within the soft zone, the coiling tension will concentrate at the strip edge, easily causing quality defects such as indentations and edge cracks. The axial positioning accuracy of the mandrel directly determines the coiling quality and subsequent rolling stability.
[0053] The YOLO-seg instance segmentation model is an extended version of the YOLO series of object detection algorithms, capable of simultaneously performing object detection and pixel-level mask generation in a single forward propagation. Compared to traditional two-stage segmentation methods, YOLO-seg offers advantages such as faster inference speed and end-to-end optimization. The "Nano" level refers to a lightweight variant with the smallest number of model parameters, suitable for edge deployment scenarios.
[0054] Based on the above definitions, the implementation environment of the machine vision-based mandrel soft area positioning method provided in this application embodiment is described. Indicatively, this implementation environment may include: an industrial camera, a laser rangefinder, a position sensor, an edge computing server, a programmable logic controller (PLC), and a human-machine interface terminal. The industrial camera is used to acquire multiple frames of images of the strip head and the mandrel surface within a stationary window of the mandrel; the laser rangefinder, as an optional enhancement component, is used to acquire the height difference between the strip head and the mandrel surface in real time to support dynamic correction; the position sensor is used to detect when the strip head reaches the detection position and simultaneously trigger vision acquisition and primary control logic; the edge computing server is used to perform image preprocessing, YOLO-seg instance segmentation inference, homography matrix coordinate mapping, and deviation correction calculations, and is equipped with an artificial intelligence chip to accelerate model inference using the TensorRTFP16 engine; the PLC controller, as the core of the primary automation system, is used to receive the deviation value output by the vision system and execute position compensation closed-loop control in the next winding cycle; the human-machine interface terminal is used to provide a parameter configuration interface, calibration status monitoring, and visualization of measurement results. The sensors include, but are not limited to, industrial cameras, laser rangefinders, and position sensors; the processors include, but are not limited to, edge computing servers, PLC controllers, central processing units, multi-core processors, or artificial intelligence chips, etc., without any specific limitations.
[0055] Based on the above explanations of terms and implementation environments, the application scenarios of the embodiments of this application will be described. The machine vision-based mandrel soft area localization method provided in the embodiments of this application can be applied to scenarios including but not limited to the following:
[0056] In the coiling process of a cold-rolled strip steel production line, such as in a continuous annealing or pickling rolling unit, the strip steel, after rolling, needs to be coiled into a coil by a coiler. The landing position of the strip head at the beginning of the coiling directly determines the end face quality and tapered accuracy of the entire coil. This solution uses machine vision to detect the relative positional relationship between the strip head and the soft zone of the mandrel in real time, providing precise positioning deviation feedback to the primary automation system. This ensures the consistency of coiling quality in a flexible production environment with multiple specifications and varieties, avoiding coiling outside the soft zone due to strip head skew and the resulting strip edge damage.
[0057] In direct coiling after hot rolling or rewinding inspection units after cold rolling, the production line operates at a tight pace with frequent roll changes. The position of the soft zone on the mandrel may drift slightly due to equipment maintenance or wear of the expansion blocks. This solution supports rapid self-inspection and benchmark calibration after each shift's start-up by establishing physical coordinate mapping parameters and soft zone reference positioning parameters during the calibration phase. It can confirm whether the spatial reference of the soft zone centerline has shifted without repeated manual measurements, thereby significantly shortening the debugging time after roll changes and improving the production line's operating rate.
[0058] When the coiler operates continuously at high speeds, the real-time positioning of the strip head directly affects production efficiency. This invention completes image acquisition and calculation within the brief static window after the strip head reaches the mandrel surface and triggers the position sensor. Upon outputting the positioning deviation, the PLC immediately controls the mandrel rotation and simultaneously adjusts the lateral position of the strip, achieving seamless integration from detection to execution. Simultaneously, the system records the positioning deviation of each strip coil. If a continuous deviation trend occurs (e.g., due to reference drift caused by mechanical vibration), the operator can promptly remind the operator to recalibrate or adjust the soft-zone reference positioning parameters based on the deviation statistics. This closed-loop feedback mechanism ensures that positioning accuracy does not decrease with long-term production line operation, effectively guaranteeing the consistency of batch coiling and the quality of the finished product.
[0059] This is merely an illustrative example; the machine vision-based mandrel soft area localization method provided in this application embodiment can also be applied to other application scenarios. It is only an example and does not limit the specific application scenarios.
[0060] In one exemplary embodiment, such as Figure 1As shown, a machine vision-based mandrel soft area localization method is provided. This embodiment illustrates the application of this method to a terminal in the aforementioned implementation environment. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps 101 to 106:
[0061] Step 101: Obtain the pre-calibrated mandrel soft-area positioning reference parameter set; the mandrel soft-area positioning reference parameter set includes physical coordinate mapping parameters, head-lifting parallax correction parameters, and soft-area reference positioning parameters.
[0062] Specifically, the set of reference parameters for mandrel soft zone positioning is generated once through a standardized calibration process after the initial system installation or each equipment maintenance and stored in non-volatile memory. It can be directly accessed during system operation without repeated calculations. For example, the physical coordinate mapping parameters can be implemented using a planar homography transformation algorithm to establish a one-to-one correspondence between the image pixel coordinate system and the mandrel surface physical coordinate system; the strip lifting parallax correction parameters are obtained by calibrating strips of various thicknesses to quantify the influence of different lifting heights on lateral positioning accuracy; and the soft zone reference positioning parameters are the standard physical coordinates of the mandrel soft zone centerline, serving as the absolute reference for all positioning measurements.
[0063] Step 102: In response to the acquisition of the strip head arrival signal, multiple frames of mandrel surface images are continuously acquired within the mandrel stationary window to obtain the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the start of the mandrel rotation and winding.
[0064] Specifically, when the strip head is conveyed from the roller conveyor to the mandrel surface and adheres to it, a position sensor installed at the end of the roller conveyor detects the arrival signal of the strip head. This signal simultaneously triggers the vision acquisition system and the programmable logic controller (PLC). The PLC immediately latches the control signal of the mandrel drive motor, keeping the mandrel completely stationary. During this stationary period, the vision acquisition system continuously acquires multiple frames of images at fixed frame intervals. For example, the duration of this stationary window can be set from 250ms to 500ms depending on the on-site process cycle time to fully cover the time required for image acquisition and processing, avoiding the cumulative errors and mechanical vibration interference caused by motion compensation algorithms in traditional dynamic detection schemes.
[0065] Step 103: Perform target detection on each frame of the target image in the image sequence to be detected, and obtain the pixel coordinates of the soft area center line and the pixel coordinates of the leading edge.
[0066] Specifically, the method first performs preprocessing operations on each frame of the target image, including distortion correction and Gaussian filtering denoising based on pre-calibrated camera intrinsic parameters, to eliminate image noise caused by lens distortion and electromagnetic interference in the industrial environment. For example, soft area detection can be achieved by converting the image to the HSV color space and extracting the luminance channel for threshold segmentation. After removing noise through morphological operations, the left and right edges of the largest connected component are extracted, and the centerline coordinates are calculated. Head detection can use a lightweight instance segmentation model for inference, outputting a pixel-level binary mask of the head, and then obtaining the precise edge coordinates at the junction of the head and the spindle through an edge extraction algorithm. Furthermore, since the geometry and position of the soft area remain stable during production, subsequent frames after successful detection of the first frame can reuse the soft area detection results, only re-detecting when the detected soft area position offset exceeds a preset threshold, thereby improving processing efficiency.
[0067] Step 104: Based on the physical coordinate mapping parameters, convert the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge into physical coordinates, and calculate the original positioning deviation and detection quality parameters corresponding to each frame of the target image.
[0068] Specifically, this method utilizes physical coordinate mapping parameters to convert two-dimensional pixel coordinates on the image plane into three-dimensional physical coordinates on the mandrel surface, achieving a transformation from pixel-level detection to millimeter-level positioning. For example, the original positioning deviation is obtained by calculating the axial difference between the physical coordinates of the head edge and the physical coordinates of the soft area centerline, directly reflecting the relative positional relationship between the head and the soft area reference in a single frame image. Furthermore, the detection quality parameters are obtained by comprehensively evaluating the reliability of the soft area detection results and the head detection results. The reliability of soft area detection can be quantified by the matching degree between the detected soft area area and the standard area, and the deviation between the soft area centerline position and the reference position. The reliability of head detection can be quantified by the target confidence output by the instance segmentation model and the average gradient value of the head edge. The combination of these two methods can effectively distinguish between high-quality detection results and low-quality detection results affected by interference.
[0069] Step 105: Select target images whose detection quality parameters are not lower than the quality threshold as valid frames, and perform weighted fusion on the original positioning deviation of the valid frames according to the detection quality parameters to obtain the weighted average original deviation.
[0070] Specifically, this method sets a detection quality threshold to remove low-quality image frames affected by factors such as on-site oil stains, oxide scale, and sudden changes in lighting, retaining only valid frames with reliable detection quality for subsequent calculations. For example, the weighted fusion process can assign different weights based on the detection quality parameters of each valid frame, with higher-quality frames receiving greater weights, thus maximizing the contribution of high-quality data to the final result. Furthermore, this method can introduce temporal correlation weights, assigning higher weights to frames in the middle of the image sequence, as the leading state of these middle frames is the most stable, further reducing the impact of random errors and demonstrating higher robustness compared to the traditional simple arithmetic average method.
[0071] Step 106: Based on the obtained current strip thickness parameters and strip head float parallax correction parameters, the weighted average original deviation is corrected for height difference parallax to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
[0072] Specifically, in actual production, the strip head cannot be perfectly flat against the mandrel surface, resulting in a certain amount of float. This height difference causes parallax in camera imaging, leading to a systematic deviation between the detected strip head position and the actual position. For example, this method uses a programmable logic controller (PLC) to acquire the current strip thickness parameters in real time. Based on pre-calibrated strip head float parallax correction parameters, it calculates the parallax correction amount corresponding to the current thickness through linear interpolation. Subtracting this correction amount from the weighted average original deviation yields the final positioning deviation that eliminates the parallax effect. Furthermore, this method can also use a laser rangefinder mounted directly above the mandrel to collect the actual float amount of the strip head in real time, replacing the strip thickness parameters for correction, to further improve positioning accuracy under extreme conditions. The final mandrel soft zone positioning deviation is sent to the PLC. The PLC adjusts the guide device and roller speed based on this deviation value, ensuring the strip head accurately moves to the soft zone centerline position. After centering, the mandrel is then controlled to begin rotation and winding.
[0073] In summary, the machine vision-based mandrel soft-area localization method provided in this application establishes a unified and accurate measurement benchmark for the entire localization process by acquiring a set of benchmark parameters for mandrel soft-area localization, including physical coordinate mapping parameters, strip head levitation parallax correction parameters, and soft-area benchmark localization parameters. Based on this benchmark, in response to the arrival signal of the strip head, multiple frames of mandrel surface images are continuously acquired within a static window from when the strip head is close to the mandrel until the mandrel begins to rotate, forming a sequence of images to be detected. This fundamentally avoids the errors caused by motion compensation in traditional dynamic detection schemes. Subsequently, target detection is performed on each frame of the target image to obtain the pixel coordinates of the soft-area centerline and the strip head edge. Based on the physical coordinate mapping parameters, these coordinates are converted into physical coordinates, and the original localization of each frame is calculated. The system employs deviation and detection quality parameters to characterize the reliability of single-frame detection. Reliable valid frames are then selected based on these parameters, and the original positioning deviations of these valid frames are weighted and fused according to the detection quality parameters. This effectively suppresses the impact of random interference on the detection results. Finally, the weighted average original deviation is corrected for height difference parallax by combining the real-time acquired current strip thickness parameters and strip head float parallax correction parameters. This eliminates systematic measurement errors caused by strip head float, ultimately yielding accurate mandrel soft area positioning deviation, which is used to instruct the PLC to perform strip lateral position adjustment. This significantly improves the accuracy and reliability of mandrel soft area positioning, effectively reduces deviation and tower-shaped defects during cold-rolled strip coiling, and improves production efficiency and product quality.
[0074] In one embodiment, obtaining a pre-calibrated set of mandrel soft-area positioning reference parameters includes:
[0075] Acquire multi-pose calibration images of the calibration board on the soft area plane of the mandrel, perform corner detection on the multi-pose calibration images, and solve the camera intrinsic parameters based on the corner pixel coordinates and the known physical coordinates corresponding to the corner pixel coordinates, to obtain the camera intrinsic parameter matrix and distortion coefficient vector.
[0076] Specifically, the calibration board can be a checkerboard calibration board, with the grid side lengths being known standard physical dimensions. Multi-pose calibration images are acquired by placing the calibration board on the plane of the mandrel's soft area and changing its tilt angle, rotation angle, and position; the number of images acquired is sufficient to cover the entire field of view of the camera. Corner detection can employ the Harris corner detection algorithm or the Shi-Tomasi corner detection algorithm to extract the sub-pixel level coordinates of all checkerboard corners in each calibration image. Combined with the known physical coordinates corresponding to each corner, the camera intrinsic parameter matrix and distortion coefficient vector are obtained using the Zhang Zhengyou calibration method. The camera intrinsic parameter matrix includes the camera's focal length and principal point coordinates, while the distortion coefficient vector includes radial and tangential distortion coefficients, used to correct image distortion caused by lens distortion. For example, corner detection uses the Shi-Tomasi corner detection algorithm to extract the sub-pixel level coordinates of all checkerboard corners in each calibration image, and combines this with the known physical coordinates of each corner to obtain the camera intrinsic parameter matrix and distortion coefficient vector using the Zhang Zhengyou calibration method. This method can correct distortion in the original image using the following formula to obtain a distortion-free image:
[0077]
[0078] in, and These are the normalized coordinates of the pixels in the original image. and These are the normalized coordinates after distortion correction. The square of the distance from the pixel to the center of the image. , , The radial distortion coefficient is... , These are the tangential distortion coefficients, and together they form the distortion coefficient vector. Further, the camera intrinsic parameter matrix is used to convert normalized coordinates to pixel coordinates, and its expression is:
[0079]
[0080] in, and These are the camera's focal lengths along the x and y axes, respectively, in pixels. and The pixel coordinates of the principal point of the image are given. This step is performed only once after the system is initially installed. The resulting camera intrinsic parameter matrix and distortion coefficient vector are stored in non-volatile memory and can be directly called during subsequent system operation without repeated calibration.
[0081] The calibration plate is flatly attached to the cylindrical surface where the soft area of the mandrel is located to acquire a planar calibration image. The pixel coordinates of multiple corner control points and their corresponding physical coordinates are extracted from the planar calibration image. Based on the correspondence between the pixel coordinates and physical coordinates of the corner control points, the homography matrix from the image plane to the physical plane of the soft area is solved by direct linear transformation. The scale conversion coefficient is calculated based on the ratio of the physical distance between the corner control points to the corresponding pixel distance. The homography matrix and the scale conversion coefficient are used as physical coordinate mapping parameters.
[0082] Specifically, since the target area where the soft area of the mandrel is located is relatively small compared to the overall size of the mandrel, the curvature of its cylindrical surface can be ignored, and therefore it can be approximated as a plane. A planar homography transformation is used to establish the mapping relationship between pixel coordinates and physical coordinates. For example, at least six checkerboard corner points distributed throughout the entire target area of the soft area are selected as corner control points to ensure the accuracy of the homography matrix across the entire measurement area. A direct linear transformation is used to establish an overdetermined system of equations and solve it using the least squares method to obtain a 3×3 homography matrix. The mapping relationship from pixel coordinates to physical coordinates is established using the following formula:
[0083]
[0084] in, It is a 3×3 homography matrix. and These are the pixel coordinates of the pixels in the image. , , These are intermediate calculation variables. The final physical coordinates are obtained by normalizing these intermediate calculation variables using the following formula:
[0085]
[0086] in, and pixel coordinates The corresponding physical coordinates of the mandrel surface, in millimeters, are used in this normalization step to effectively correct nonlinear image distortion caused by camera perspective distortion, achieving accurate conversion from pixel coordinates to physical coordinates. Furthermore, the scale conversion coefficient is a local linear approximation of the homography matrix near the measurement area, used for rapid conversion in subsequent deviation calculations. This method can extract the scale conversion coefficient from the homography matrix using the following formula:
[0087]
[0088] in, This is a scale conversion factor, expressed in millimeters per pixel. , , , , Homography matrix The corresponding elements in the model. In this embodiment, the homography matrix provides a complete nonlinear perspective mapping, suitable for accurate coordinate transformation across the entire field of view, while the scale conversion coefficient provides a simplified linear conversion relationship, suitable for rapid deviation calculation within the measurement area. The combination of the two can balance computational accuracy and efficiency.
[0089] The original deviation value sequence is obtained, and linear fitting is performed based on the original deviation value sequence and the corresponding strip thickness parameters to obtain the strip head floating parallax correction parameter between the strip head floating height and the projection system offset; wherein, the original deviation value sequence is measured by the vision system when the physical deviation of the strip head is manually confirmed to be zero under the strip threading conditions of multiple strips with different thicknesses.
[0090] Specifically, for each thickness specification of strip steel, a mechanical centering fixture is used to precisely fix the strip head at the centerline position of the mandrel's soft zone, ensuring that the physical deviation between the strip head and the centerline of the soft zone is zero. At this time, the vision system continuously acquires multiple frames of images and calculates the corresponding original positioning deviation, taking the average value as the original deviation value for that thickness specification. The original deviation value sequence contains original deviation values corresponding to various strip steel thickness specifications. Since the strip head lifting height is positively correlated with the strip steel thickness, a linear relationship between the strip head lifting height and the systematic offset of the projection system can be established through linear fitting. For example, this method can linearly fit the original deviation value sequence to the corresponding strip steel thickness parameters using the following formula:
[0091]
[0092] in, The height at which the leader floats is The corresponding systematic offset of the projection system, in millimeters. For parallax correction factor for buoyancy, The buoyancy height is expressed in millimeters. The parallax correction factor for buoyancy can be calculated using the following formula:
[0093]
[0094] in, The quantity of strip steel thickness specifications to be included in the calibration. This represents the average original deviation value corresponding to the j-th thickness specification of strip steel, in millimeters. Let be the thickness of the strip of the j-th thickness specification, in millimeters. Furthermore, during system operation, the original positioning deviation can be corrected for height difference parallax using the following formula:
[0095]
[0096] in, The corrected positioning deviation is expressed in millimeters. This represents the original positioning deviation, in millimeters. This refers to the thickness of the currently produced strip steel, in millimeters. Using the scale conversion coefficient, this method can effectively eliminate systematic measurement errors caused by the floating of strip heads of strips with different thicknesses, and improve the consistency of positioning accuracy for strips of different specifications. For example, this calibration process can be completed during the system commissioning phase, or it can be updated periodically during production to adapt to changes in equipment status.
[0097] Under calibration, the left and right edge pixel coordinates of the soft area of the mandrel are identified, the center line pixel coordinates of the soft area are calculated based on the left and right edge pixel coordinates, and the center line pixel coordinates of the soft area are mapped to the physical coordinate system through homography matrix to obtain the soft area reference positioning parameters.
[0098] Specifically, the calibration state refers to a state where the mandrel surface is free of strip steel, the mandrel remains stationary, and the lighting conditions are consistent with normal production conditions. In this state, an image of the mandrel surface is acquired, and the soft region of the mandrel is identified using an image segmentation algorithm. For example, soft region edge recognition is achieved by converting the image to the HSV color space and extracting the luminance channel for threshold segmentation. After removing noise through morphological opening operations, the left and right edge pixel coordinates of the largest connected component are extracted. The pixel coordinates of the soft region centerline are calculated using the following formula:
[0099]
[0100] in, The pixel coordinates of the center line of the soft area. and These are the pixel coordinates of the left and right edges of the soft area, respectively. Using the homography matrix obtained earlier, these pixel coordinates are converted into physical coordinates, which are the soft area reference positioning parameters. Specifically, this method can complete the mapping from pixel coordinates to physical coordinates to obtain the soft area reference positioning parameters using the following formula:
[0101]
[0102]
[0103] in, It is a planar homography matrix. The radial pixel coordinates of the soft area centerline. , , For intermediate calculation variables, The axial physical coordinates of the soft area centerline are the soft area reference positioning parameters. These parameters serve as the absolute reference for the entire positioning system; all measurements of the lead-end position are relative to this reference. The initial positioning deviation calculated in subsequent steps is the difference between the physical coordinates of the lead-end edge and these soft area reference positioning parameters. Furthermore, the system can perform a quick self-check after each shift's startup, acquiring a frame of mandrel surface image and calculating the current physical coordinates of the soft area centerline. These coordinates are then compared with the soft area reference positioning parameters. If the deviation exceeds a preset threshold, recalibration is prompted to ensure the long-term accuracy and stability of the system.
[0104] In one embodiment, target detection is performed on each frame of the target image in the image sequence to be detected, obtaining the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, including:
[0105] Each frame of the target image undergoes color space conversion, transforming the RGB color space into the HSV color space to obtain an HSV image containing hue, saturation, and lightness components.
[0106] Specifically, industrial lighting conditions are complex and variable, with interference such as light intensity fluctuations, localized reflections, and shadows. Since all three components of the RGB color space are related to light intensity, direct segmentation within the RGB space leads to unstable detection results. For example, the HSV color space separates the color and brightness information of an image. The brightness component only reflects the brightness level of a pixel and is unaffected by color changes, effectively resisting interference from variations in light intensity. Therefore, soft-area detection is performed in the HSV color space. Furthermore, this color space conversion process is achieved through a linear transformation, converting the RGB value of each pixel into its corresponding H, S, and V component values. The converted HSV image has the same resolution as the original RGB image.
[0107] By performing threshold segmentation on the luminance components in the HSV image using a preset luminance threshold, pixels with luminance values greater than the preset threshold are identified as soft area candidate pixels, thus obtaining a binary image of the soft area candidate region.
[0108] Specifically, the soft area of the mandrel is a high-contrast marked area pre-set on the surface of the mandrel. Its visual characteristic is that it is significantly brighter than the surrounding mandrel surface. Therefore, the soft area can be separated from the background area by brightness thresholding. For example, the preset brightness threshold is determined according to the actual lighting conditions during the system calibration stage. It is usually set to a value that can clearly distinguish the soft area from the background, while retaining a certain margin to accommodate small changes in lighting. Further, the thresholding process compares the brightness value of each pixel in the brightness component image with the preset threshold. Pixels with brightness values greater than the threshold are assigned a value of 255 (white), and pixels with brightness values less than or equal to the threshold are assigned a value of 0 (black), thereby obtaining a binary image of the candidate soft area region with clear black and white distinction.
[0109] Morphological opening operations are performed on the binary image to eliminate scattered noise and retain the main soft region, resulting in a denoised soft region binary image.
[0110] Specifically, the binary image after thresholding often contains isolated bright noise points caused by surface reflections, oil stains, oxide scale, or electromagnetic interference from the steel strip. These noise points can affect the accuracy of subsequent connected component analysis. For example, the morphological opening operation consists of two operations: erosion followed by dilation. The erosion operation eliminates small bright areas, while the dilation operation restores the original shape of the remaining large areas. Therefore, the opening operation can effectively remove small scattered noise points without changing the shape and position of the main soft region. Furthermore, the morphological opening operation uses a 3×3 rectangular structuring element, the size of which matches the size of common noise points, achieving a good balance between denoising effect and soft region edge preservation.
[0111] Connected component labeling is performed on the denoised soft region binary image, the pixel area of each connected component is extracted, and the connected component with the largest pixel area is determined as the soft region.
[0112] Specifically, connected component labeling connects adjacent white pixels in a binary image into a single entity, forming multiple independent connected components, each representing a potential target region. For example, the core soft region is the largest high-brightness area in the image, while other smaller connected components are typically residual noise points or localized reflections. Therefore, by comparing the pixel areas of each connected component and selecting the largest soft region, most interference can be effectively eliminated. Furthermore, connected component labeling can simultaneously extract features such as the bounding rectangle and center coordinates of each connected component, providing foundational data for subsequent edge extraction and centerline calculation.
[0113] The soft region is fitted with a minimum bounding rectangle, and the pixel coordinates of the left and right edges of the minimum bounding rectangle are extracted. The average value of the pixel coordinates of the left and right edges is taken to obtain the pixel coordinates of the center line of the soft region.
[0114] Specifically, the soft area of the mandrel is typically processed into a regular rectangular shape. Minimum bounding rectangle fitting can accurately describe the overall position and boundary of the soft area, avoiding errors caused by irregular edges. For example, the minimum bounding rectangle is the smallest rectangle that completely encloses all pixels of the soft area, with its left and right edges corresponding to the left and right boundaries of the soft area, respectively. Further, the center line of the soft area is its geometric center, obtained by calculating the average of the pixel coordinates of the left and right edges of the minimum bounding rectangle. This center line is the reference position of the soft area on the image plane. Since the geometry and position of the soft area remain stable during production, the pixel coordinates of the center line can be reused in subsequent frames after successful detection in the first frame. The entire soft area detection process is repeated only when the detected position offset exceeds a preset threshold, thereby improving overall processing efficiency.
[0115] In one embodiment, target detection is performed on each frame of the target image in the image sequence to be detected, obtaining the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, including:
[0116] Each frame of the target image is input into the improved YOLO-seg instance segmentation model for forward inference to obtain a pixel-level binary mask of the header; the improved YOLO-seg instance segmentation model is trained on an optimized training dataset constructed by high-precision polygon annotation of the edge region at the junction of the header and the core.
[0117] Specifically, traditional edge detection and threshold segmentation methods struggle to adapt to complex working conditions such as oil stains, oxide scale, scratches, and fluctuating ambient light on strip steel surfaces, easily leading to edge breakage or false detections. In contrast, instance segmentation models can learn the visual features of the strip head end-to-end, outputting accurate pixel-level segmentation results with stronger robustness and generalization ability. For example, the training dataset is collected from the production site, covering strip steel of different thicknesses, surface conditions, and image samples under different lighting conditions, with no fewer than 500 samples. Unlike conventional rectangular annotation, this technical solution uses polygonal annotation, focusing on sub-pixel precision annotation of the boundary edge where the strip head contacts the mandrel. This ensures the model can accurately learn the features of this key area, thereby improving the detection accuracy of the strip head edge. After training, the model is exported to the TensorRT FP16 acceleration engine, achieving a single-frame inference time of no more than 2 milliseconds on an industrial-grade GPU to meet production cycle requirements. During forward inference, the model outputs a probability map of the leading region. The probability map is converted into a binary image by a preset confidence threshold. Pixels with a value of 1 represent the leading region and pixels with a value of 0 represent the background region. This binary image is the pixel-level binary mask of the leading region.
[0118] Contour extraction is performed on the pixel-level binary mask to obtain complete contour data with the header.
[0119] Specifically, the purpose of contour extraction is to extract the set of boundary pixels of the leading region from the binary mask, providing basic data for subsequent edge localization. For example, this method can use the Suzuki85 contour tracking algorithm for contour extraction. This algorithm can traverse all pixels in the binary image, identify continuous closed boundaries, and store contour data according to hierarchical relationships. Since there may be small isolated regions in the binary mask caused by model misdetection or image noise, after contour extraction, it is necessary to perform area filtering on all extracted contours. The pixel area enclosed by each contour is calculated, and contours with areas smaller than a preset area threshold are identified as noise contours and removed. Only the contour with the largest area is retained as the complete leading contour data. This filtering step effectively eliminates most interference, ensuring the accuracy of subsequent processing.
[0120] Based on the complete contour data, the position of the foremost edge of the belt head facing the soft area of the spindle is identified, and the pixel coordinates of the belt head edge are extracted from the foremost edge position.
[0121] Specifically, the conveying direction of the strip steel on the production line is fixed, so the orientation of the strip head in the image is known. The foremost edge facing the soft area is the position where the strip head first contacts the mandrel, and it is also the key reference for positioning. For example, for strip steel conveyed from right to left, the foremost edge of the strip head facing the soft area corresponds to the set of all pixels with the smallest x-coordinate in the contour. Since there may be tiny burrs or deformations at the edge of the strip head, directly taking a single pixel as the edge would introduce a large error. Therefore, a random sampling consensus algorithm is used to fit a straight line to all pixels at the foremost edge, resulting in a straight line equation that represents the overall edge trend of the strip head. This straight line equation is the mathematical expression of the strip head edge. The pixel coordinates of any point on the strip head edge can be calculated using this equation. Taking the midpoint coordinates of this straight line as the final pixel coordinates of the strip head edge effectively reduces the impact of local edge deformation on positioning accuracy.
[0122] In another embodiment, the method further includes:
[0123] Multi-condition mandrel surface images were collected as the original image dataset, which covered different strip thicknesses, different lighting conditions, and the mandrel state at different lateral positions of the strip head.
[0124] Specifically, since the generalization ability of deep learning models directly depends on the diversity and representativeness of the training dataset, and the production conditions in industrial sites are complex and varied, models trained on data collected under a single condition are prone to performance degradation in practical applications. Furthermore, strip steel of different thicknesses has different surface reflectivity and edge morphology; thin strip steel has sharper edges and is prone to warping, while thick strip steel has blunter edges and greater lift. Different lighting conditions include natural light during the day, artificial lighting at night, and dynamic light and shadow changes during equipment operation. Different lateral positions of the strip head include three typical states: left-leaning, centering, and right-leaning. The relative positional relationship between the strip head and the mandrel is different in each state. In this embodiment, the method uses the same industrial camera and lens as the actual deployment during the dataset acquisition process, maintaining the same shooting distance, angle, and exposure parameters to ensure that the distribution of training data and actual inference data is consistent, avoiding domain offset problems.
[0125] Each image in the original image dataset is annotated with polygons based on the edge encryption at the intersection, and the complete outline of the annotation is obtained by annotating the head, resulting in a high-precision polygon annotation mask.
[0126] Specifically, conventional rectangular annotation methods include a large amount of background area within the annotation box, causing the model to learn a large number of irrelevant features. Polygon annotation, on the other hand, can accurately delineate the target's outline, improving the model's segmentation accuracy. This technical solution employs a polygon annotation method based on edge densification at the junction. In the junction area where the head and core meet, a annotation point is set every 1 to 2 pixels, while in other edge areas of the head, a annotation point is set every 5 to 10 pixels. This annotation method can significantly enhance the annotation accuracy of key localization areas without significantly increasing the annotation workload, enabling the model to more accurately learn the subtle features of the junction edge between the head and core, thereby improving the final localization accuracy. Furthermore, after annotation, all annotated data needs to be cross-validated to remove incorrectly labeled or incomplete samples, ensuring the quality of the dataset.
[0127] We use a nano-level YOLO-seg network as the basic architecture, set the number of output categories to 1 to predict the leading region, and divide the original image dataset after high-precision polygon labeling masking into training set and validation set.
[0128] Specifically, the nano-level YOLO-seg network is a lightweight instance segmentation network characterized by its small parameter count and fast inference speed, making it ideal for real-time detection needs in industrial settings. For example, compared to larger-scale YOLO-seg models, the nano-level model has only about one-tenth the number of parameters, yet achieves comparable accuracy in single-class object detection tasks. Since this approach only requires detecting the head-bearing object class, the number of output classes is set to 1, simplifying the network structure and further improving inference speed and accuracy. Furthermore, this method can randomly divide the labeled dataset into training and validation sets in an 8:2 ratio. The training set is used to update model parameters, while the validation set is used to evaluate the model's generalization ability and monitor performance during training.
[0129] The YOLO-seg network is trained on the training set. During training, a weighted sum of the binary cross-entropy loss function and the Dice loss function is used as the segmentation loss. Training is stopped when the validation set loss no longer decreases for 5 consecutive epochs, resulting in an improved YOLO-seg segmentation model. The improved YOLO-seg segmentation model is then exported as a TensorRTFP16 inference engine file for deployment on the GPU memory of the vision system server.
[0130] Specifically, a single loss function often struggles to simultaneously satisfy the requirements of segmentation accuracy and region integrity. The binary cross-entropy loss function calculates the loss independently for each pixel, improving pixel-level classification accuracy, while the Dice loss function calculates the loss for the entire segmented region, improving region integrity, especially suitable for cases of foreground-background imbalance. This method can calculate the weighted segmentation loss using the following formula:
[0131]
[0132] in, For the total partition loss, For binary cross-entropy loss, For Dice's loss, and Let be the weight coefficient, and satisfy... For example, this technical solution will... Set to 0.5. Setting it to 0.5 allows the model to balance pixel-level accuracy and region integrity. The binary cross-entropy loss is calculated using the following formula:
[0133]
[0134] in, and These are the height and width of the image, respectively, in pixels. For pixels The true label is 1, indicating it belongs to the header area, and 0, indicating it belongs to the background area. Predict pixels for the model The probability of belonging to the leading region. The Dice loss is calculated using the following formula:
[0135]
[0136] The numerator is the area of the intersection of the predicted and ground truth regions, and the denominator is the sum of the areas of the predicted and ground truth regions. Furthermore, an early stopping strategy can be employed during training: when the validation set loss no longer decreases for five consecutive epochs, the model is considered converged, and training is stopped. This effectively prevents overfitting and saves training time. Even further, data augmentation techniques are used during training, including random flipping, random cropping, and random brightness and contrast adjustments, to further improve the model's generalization ability.
[0137] Preferably, in this embodiment, the method exports the improved YOLO-seg segmentation model as a TensorRT FP16 inference engine file for deployment in the GPU memory of the vision system server. Specifically, the original PyTorch or ONNX format models have slow inference speeds on GPUs, which cannot meet the real-time requirements of industrial applications. TensorRT, developed by NVIDIA, is a high-performance deep learning inference optimizer that can perform a series of optimizations on the model, such as layer fusion, weight quantization, and kernel optimization, significantly improving inference speed.
[0138] In one embodiment, the weighted average original deviation is corrected for height difference parallax based on the obtained current strip thickness parameters and strip head float parallax correction parameters to obtain the mandrel soft zone positioning deviation, including:
[0139] The following formula is used to correct the weighted average original deviation by subtracting the height difference parallax, based on the current strip thickness parameters, strip head float parallax correction parameters, and physical coordinate mapping parameters, to obtain the mandrel soft area positioning deviation:
[0140] ;
[0141] in, This is due to the positioning deviation of the mandrel soft area. The weighted average of the original deviations, Parallax correction parameters for levitation. The current strip thickness parameters are as follows: These are the scale transformation coefficients in the physical coordinate mapping parameters.
[0142] Specifically, in actual production, due to its own rigidity and the mechanical impact during conveying, the strip head cannot be completely flat against the mandrel surface, resulting in a certain amount of float. This height difference causes parallax in the camera image, causing a systematic offset between the detected lateral position of the strip head and its actual position. This offset is approximately linearly related to the float height of the strip head. However, under normal threading conditions, the float height of the strip head has a stable positive correlation with the strip thickness. Therefore, the original positioning deviation can be corrected by combining a pre-calibrated strip head float parallax correction parameter with the current strip thickness parameter. Where, The final mandrel soft area positioning deviation, in millimeters, is the accurate positioning result after multi-frame weighted fusion and height difference parallax correction; The weighted average original deviation obtained from the aforementioned steps is expressed in millimeters and is the weighted average of the original positioning deviations of multiple valid frames. The parallax correction parameter for the levitation head obtained from the aforementioned calibration steps is the slope of the linear relationship between the levitation head height and the systematic offset of the projection, in millimeters per millimeter. The thickness parameter of the current production strip is obtained in real time from the PLC, in millimeters; This is the scale conversion coefficient in the physical coordinate mapping parameters obtained from the aforementioned calibration steps, measured in millimeters per pixel. It is used to convert pixel-level offsets into physical size offsets. It should be noted that the physical meaning of this formula is: the lateral projection offset caused by strip head floating is equal to the strip head floating parallax correction coefficient multiplied by the current strip thickness, and then multiplied by the scale conversion coefficient. Since this offset causes the detected original deviation to be greater than the actual deviation, it needs to be subtracted from the weighted average original deviation to obtain the true mandrel soft area positioning deviation. This correction method does not require additional hardware; it can be implemented using only the existing strip thickness parameters and pre-calibrated system parameters from the production process. Without increasing costs, it effectively eliminates the systematic measurement error caused by strip head floating and significantly improves the consistency of positioning accuracy for strips of different thicknesses. After correction, the obtained mandrel soft area positioning deviation is sent to the PLC as the direct control basis for the PLC to perform lateral position adjustment of the strip. The PLC adjusts the position of the guide device according to this deviation value, so that the strip head accurately moves to the center line position of the mandrel soft area. After centering is completed, the PLC controls the mandrel to start rotating and winding.
[0143] In summary, this technical solution addresses the industry pain points of low positioning accuracy and poor stability of the soft area of the mandrel in the cold-rolled strip coiling process. It constructs a complete positioning system from five core technical dimensions: measurement timing, benchmark establishment, target detection, data fusion, and error correction. The solution establishes a unified measurement benchmark before calibration, including physical coordinate mapping parameters, strip head floating parallax correction parameters, and soft area benchmark positioning parameters, providing accurate coordinate transformation and error correction basis for all subsequent detection steps. Secondly, it utilizes the process gap between the strip head and the start of coiling to form a static window for the mandrel, continuously acquiring multiple frames of images within this window, fundamentally avoiding the cumulative errors and mechanical vibration interference caused by motion compensation algorithms in traditional dynamic detection schemes. Subsequently, it employs a threshold segmentation method based on the HSV color space and an improved YOLO-seg instance segmentation model for soft area and strip head detection, respectively. The instance segmentation model is trained on a dataset with encrypted annotations of boundary edges, enabling it to output data in complex industrial environments. The system obtains precise pixel-level edge information; then, based on physical coordinate mapping parameters, it converts the detected pixel coordinates into physical coordinates, calculates the original positioning deviation of each frame and the detection quality parameters characterizing the detection reliability; then, it filters out valid frames through the detection quality parameters, and performs weighted fusion on the original positioning deviation of the valid frames according to the detection quality parameters, effectively suppressing random interference such as changes in ambient lighting and surface defects of the strip steel; finally, it combines the current strip steel thickness parameters sent by the PLC in real time and the pre-calibrated strip head floating parallax correction parameters to perform height difference parallax correction on the weighted average original deviation, eliminating the systematic measurement error caused by strip head floating, and finally obtaining accurate mandrel soft area positioning deviation, which is used to guide the PLC to perform strip steel lateral position adjustment.
[0144] This technical solution completely solves the inherent error problem of dynamic detection by measuring the mandrel static window. Combined with multi-frame weighted fusion technology, it significantly improves the stability of the detection results. Precise parallax correction of the lead-up ensures the consistency of positioning accuracy for strip steel of different thicknesses. The use of a lightweight instance segmentation model ensures accuracy while meeting the real-time requirements of industrial sites. The complete three-level calibration system ensures the accuracy and stability of the system during long-term operation. It can effectively reduce deviation and tower-shaped defects during the coiling process of cold-rolled strip steel, and improve production efficiency and product quality.
[0145] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0146] Based on the same inventive concept, this application also provides a machine vision-based mandrel soft-area localization system 10 for implementing the machine vision-based mandrel soft-area localization method described above. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more machine vision-based mandrel soft-area localization system 10 embodiments provided below can be found in the limitations of the machine vision-based mandrel soft-area localization method described above, and will not be repeated here.
[0147] In one exemplary embodiment, such as Figure 2 As shown, a machine vision-based mandrel soft area localization system is provided, comprising:
[0148] The reference parameter calibration module 11 is used to obtain a pre-calibrated set of reference parameters for mandrel soft-area positioning; the set of reference parameters for mandrel soft-area positioning includes physical coordinate mapping parameters, parallax correction parameters for head-mounted floating, and soft-area reference positioning parameters;
[0149] Image acquisition module 12 is used to continuously acquire multiple frames of mandrel surface images within the mandrel stationary window in response to the acquisition of the strip head arrival signal, so as to obtain the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the mandrel starting to rotate and wind up;
[0150] The target detection module 13 is used to perform target detection on each frame of the image sequence to be detected, and to obtain the pixel coordinates of the soft area center line and the pixel coordinates of the leading edge.
[0151] The coordinate transformation module 14 is used to convert the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge into physical coordinates based on the physical coordinate mapping parameters, and to calculate the original positioning deviation and detection quality parameters corresponding to each frame of the target image.
[0152] The image fusion module 15 is used to filter target images whose detection quality parameters are not lower than the quality threshold as valid frames, and to perform weighted fusion on the original positioning deviation of the valid frames according to the detection quality parameters to obtain the weighted average original deviation.
[0153] The parallax correction module 16 is used to perform height difference parallax correction on the weighted average original deviation based on the obtained current strip thickness parameters and strip head float parallax correction parameters to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
[0154] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the machine vision-based mandrel soft area localization method as described above.
[0155] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0156] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0157] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A machine vision-based method for locating soft areas of a mandrel, characterized in that, The method includes: Obtain a pre-calibrated set of mandrel soft-area positioning reference parameters; the set of mandrel soft-area positioning reference parameters includes physical coordinate mapping parameters, head-lift parallax correction parameters, and soft-area reference positioning parameters; In response to the arrival signal of the strip head, multiple frames of mandrel surface images are continuously acquired within the mandrel stationary window to obtain the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the start of the mandrel rotation and winding. Target detection is performed on each frame of the target image in the image sequence to be detected, and the pixel coordinates of the soft area center line and the pixel coordinates of the leading edge are obtained; Based on the physical coordinate mapping parameters, the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge are converted into physical coordinates, and the original positioning deviation and detection quality parameters corresponding to each frame of the target image are calculated. The target images whose detection quality parameters are not lower than the quality threshold are selected as valid frames. The original positioning deviations of the valid frames are weighted and fused according to the detection quality parameters to obtain the weighted average original deviation. Based on the obtained current strip thickness parameters and the strip head float parallax correction parameters, the weighted average original deviation is corrected for height difference parallax to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
2. The method according to claim 1, characterized in that, The process of obtaining the pre-calibrated set of mandrel soft-zone positioning reference parameters includes: Acquire multi-pose calibration images of the calibration board on the soft area plane of the mandrel, perform corner detection on the multi-pose calibration images, and solve the camera intrinsic parameters based on the corner pixel coordinates obtained by the corner detection and the known physical coordinates corresponding to the corner pixel coordinates to obtain the camera intrinsic parameter matrix and distortion coefficient vector. The calibration plate is flatly attached to the cylindrical surface where the soft area of the mandrel is located to acquire a planar calibration image. The pixel coordinates of multiple corner control points and the physical coordinates corresponding to the corner control points are extracted from the planar calibration image. Based on the correspondence between the pixel coordinates and physical coordinates of the corner control points, the homography matrix from the image plane to the physical plane of the soft area is solved by direct linear transformation. The scale conversion coefficient is calculated based on the ratio of the physical distance between the corner control points to the corresponding pixel distance. The homography matrix and the scale conversion coefficient are used as the physical coordinate mapping parameters. Obtain the original deviation value sequence, and perform linear fitting based on the original deviation value sequence and the corresponding strip thickness parameters to obtain the strip head floating parallax correction parameter between the strip head floating height and the projection systemic offset; wherein, the original deviation value sequence is measured by the vision system when the physical deviation of the strip head is manually confirmed to be zero under multiple strip thickness specifications. Under calibration, the left and right edge pixel coordinates of the soft area of the mandrel are identified, the center line pixel coordinates of the soft area are calculated based on the left and right edge pixel coordinates, and the center line pixel coordinates of the soft area are mapped to the physical coordinate system through the homography matrix to obtain the soft area reference positioning parameters.
3. The method according to claim 1, characterized in that, The step of performing target detection on each frame of the target image in the image sequence to be detected, to obtain the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, includes: For each frame of the target image, a color space conversion is performed, converting the RGB color space to the HSV color space to obtain an HSV image containing hue components, saturation components, and lightness components; The brightness component in the HSV image is segmented by a preset brightness threshold, and pixels with brightness values greater than the brightness threshold are identified as soft area candidate pixels, thus obtaining a binary image of the soft area candidate region. The binary image is processed by morphological opening to eliminate scattered noise and retain the main soft region, resulting in a denoised soft region binary image. Connected component labeling is performed on the denoised soft region binary image, and the pixel area of each connected component is extracted. The connected component with the largest pixel area is determined as the soft region. The soft region is fitted with a minimum bounding rectangle, and the pixel coordinates of the left and right edges of the minimum bounding rectangle are extracted. The average value of the left and right edge pixel coordinates is taken to obtain the pixel coordinates of the center line of the soft region.
4. The method according to claim 1, characterized in that, The step of performing target detection on each frame of the target image in the image sequence to be detected, and obtaining the pixel coordinates of the soft region centerline and the pixel coordinates of the leading edge, includes: Each frame of the target image is input into the improved YOLO-seg instance segmentation model for forward inference to obtain a pixel-level binary mask with the head; wherein, the improved YOLO-seg instance segmentation model is trained on an optimized training dataset constructed by performing high-precision polygon annotation on the edge region at the junction of the head and the core. Contour extraction is performed on the pixel-level binary mask to obtain complete contour data with header; Based on the complete contour data, the position of the foremost edge of the belt head facing the soft area of the mandrel is identified, and the pixel coordinates of the belt head edge are extracted from the foremost edge position.
5. The method according to claim 4, characterized in that, The method further includes: Multi-condition mandrel surface images under various working conditions are collected as the original image dataset. These multi-condition mandrel surface images cover different strip thicknesses, different lighting conditions, and the mandrel state at different lateral positions of the strip head. Each image in the original image dataset is annotated with polygonal labels based on the edge encryption at the intersection, and the complete outline with the head is annotated to obtain a high-precision polygonal label mask; A nano-level YOLO-seg network is used as the basic architecture, and the number of output categories is set to 1 to predict the leading region. The original image dataset after being annotated by the high-precision polygon label mask is divided into a training set and a validation set. The YOLO-seg network is trained based on the training set. During training, a weighted sum of the binary cross-entropy loss function and the Dice loss function is used as the segmentation loss. Training stops when the validation set loss no longer decreases for 5 consecutive epochs, resulting in the improved YOLO-seg segmentation model. The improved YOLO-seg segmentation model is then exported as a TensorRT FP16 inference engine file for deployment on the GPU memory of the vision system server.
6. The method according to claim 1, characterized in that, The weighted average original deviation is corrected for height difference parallax based on the obtained current strip thickness parameters and the strip head float parallax correction parameters to obtain the mandrel soft area positioning deviation, including: The mandrel soft zone positioning deviation is obtained by subtracting the height difference parallax from the weighted average original deviation using the following formula, based on the current strip thickness parameter, the strip head floating parallax correction parameter, and the physical coordinate mapping parameter: ; in, The positioning deviation of the soft area of the mandrel. The weighted average original deviation, The parallax correction parameters for the levitation head are as follows: The current strip thickness parameter, is the scale transformation coefficient in the physical coordinate mapping parameters.
7. A machine vision-based mandrel soft area positioning system for implementing the method according to any one of claims 1 to 6, characterized in that, The system includes: The reference parameter calibration module is used to obtain a pre-calibrated set of reference parameters for mandrel soft area positioning; the set of reference parameters for mandrel soft area positioning includes physical coordinate mapping parameters, parallax correction parameters for head-mounted floating, and soft area reference positioning parameters; The image acquisition module is used to continuously acquire multiple frames of mandrel surface images within a mandrel stationary window in response to the acquisition of the strip head arrival signal, thereby obtaining the image sequence to be detected; wherein, the mandrel stationary window is the time interval between the strip head touching the mandrel surface and the start of mandrel rotation and winding; The target detection module is used to perform target detection on each frame of the target image in the image sequence to be detected, and to obtain the pixel coordinates of the soft area center line and the pixel coordinates of the leading edge. The coordinate transformation module is used to convert the pixel coordinates of the soft area centerline and the pixel coordinates of the leading edge into physical coordinates based on the physical coordinate mapping parameters, and to calculate the original positioning deviation and detection quality parameters corresponding to each frame of the target image; The image fusion module is used to filter the target images whose detection quality parameters are not lower than the quality threshold as valid frames, and to perform weighted fusion on the original positioning deviation of the valid frames according to the detection quality parameters to obtain a weighted average original deviation. The parallax correction module is used to perform height difference parallax correction on the weighted average original deviation based on the acquired current strip thickness parameters and the strip head floating parallax correction parameters to obtain the mandrel soft area positioning deviation; wherein, the mandrel soft area positioning deviation is used to instruct the PLC to perform strip lateral position adjustment to control the mandrel rotation and winding.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.