A machine vision-based solar support square tube surface defect online detection method and system

CN122510221APending Publication Date: 2026-08-04ANYANG TAIFU NEW ENERGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANYANG TAIFU NEW ENERGY CO LTD
Filing Date
2026-05-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0004]但是,上述及类似技术方案在面对热镀锌方管检测时仍存在以下不足:其依赖常规图像预处理与特征提取方法,难以应对锌层高反光和锌花纹理干扰,易造成误检和漏检;同时,其单视角采集方式无法完整覆盖方管棱边区域,且仅能进行缺陷有无的二分类判断,无法输出缺陷的物理尺寸及跨面连续坐标,难以对接后续自动打磨或分级处置工序

Benefits of technology

一种基于机器视觉的太阳能支架方管表面缺陷在线检测方法及系统,通过基于晶体粒度频谱的物理成像模型与生成对抗网络相结合的合成数据生成方式,实现了对锌花纹理通过引入融合工艺容忍度规则的领域知识图谱进行合规性推理,打通了从缺陷原始感知到产线自动化处置决策的闭环,提升了检测系统的工业实用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510221A_ABST
    Figure CN122510221A_ABST
Patent Text Reader

Abstract

This invention discloses an online detection method and system for surface defects of square tubes used in solar panel supports based on machine vision, belonging to the field of surface defect detection technology. The method includes: S1, acquiring normal zinc flower texture of the square tube, constructing a zinc flower texture spectrum model characterizing the differences in crystal grain size distribution, and generating a simulated background image; establishing a micro-element bidirectional reflection distribution function model of the target defect, embedding the target defect into the background image through physical rendering, generating a synthetic training image set, and pre-training the defect detection network; S2, acquiring multi-view temporal image sequences surrounding the square tube, predicting the characteristic spatial displacement field and non-rigid deformation caused by the square tube's motion, and registering and aligning the multi-view feature images. This invention eliminates the detection blind spot in the edge region of the square tube through multi-view spatiotemporal feature alignment, displacement field prediction decoupled from rigidity, and a spatial attention mechanism for the edge region, achieving clear imaging and accurate detection of the entire surface of the square tube under high-speed motion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surface defect detection technology, specifically to an online detection method and system for surface defects of square tube solar panel brackets based on machine vision. Background Technology

[0002] With the rapid development of the photovoltaic power generation industry, the demand for solar panel mounting systems continues to rise. Square tubes, as a major structural component of solar panel mounting systems, are mostly made of hot-dip galvanized or zinc-aluminum-magnesium alloys, and their surface quality directly affects the weather resistance and service life of the mounting system. During the production of square tubes, online detection of surface defects (such as incomplete plating, zinc nodules, scratches, and pitting) is a core aspect of ensuring product quality. However, square tubes have characteristics such as right-angled square cross-sections, highly reflective zinc coating surfaces, and snowflake-like zinc patterns. Combined with the high-speed continuous operation of production lines, automated detection of surface defects faces numerous technical challenges. Therefore, developing efficient and accurate machine vision-based online inspection methods has become an urgent problem to be solved in this field.

[0003] Patent CN214201214U discloses a machine vision-based seamless steel pipe defect detection system. This system transports steel pipes via a conveyor, captures images of the pipe's outer surface using a camera, and then preprocesses and identifies defects using a computer. This solution achieves a certain degree of automated detection of steel pipe surface defects, improving detection efficiency and success rate.

[0004] However, the above and similar technical solutions still have the following shortcomings when dealing with the inspection of hot-dip galvanized square tubes: they rely on conventional image preprocessing and feature extraction methods, which are difficult to deal with the high reflectivity of the zinc layer and the interference of zinc flower texture, and are prone to false detection and missed detection; at the same time, their single-view acquisition method cannot completely cover the edge area of ​​the square tube, and can only perform binary judgment of the presence or absence of defects, and cannot output the physical size of the defects and the continuous coordinates across the surface, making it difficult to connect with subsequent automatic grinding or graded treatment processes. Summary of the Invention

[0005] The purpose of this invention is to provide an online detection method and system for surface defects of square tube solar panel brackets based on machine vision, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an online detection method for surface defects of square tube solar panel brackets based on machine vision, comprising: S1. Collect normal zinc flower texture of square tube, construct a zinc flower texture spectrum model to characterize the difference in crystal grain size distribution, and generate a simulated background image; establish a micro-facet bidirectional reflection distribution function model of the target defect, embed the target defect into the background image through physical rendering, generate a synthetic training image set, and pre-train the defect detection network. S2. Obtain multi-view temporal image sequences surrounding the square tube, predict the characteristic spatial displacement field and non-rigid deformation caused by the movement of the square tube, and register and align the multi-view feature maps. S3. Input the registered and aligned multi-view feature map into the defect detection network, aggregate the features of the overlapping areas of adjacent view fields to reconstruct the defects in the edge area of ​​the square tube, and output the defect type, physical size and cross-plane continuous coordinates based on the three-dimensional mapping established by the alignment process. S4. Instance-based input of defect type, physical size, and cross-plane continuous coordinates into a pre-built domain knowledge graph containing process tolerance rules, perform compliance reasoning, output defect level, and generate corresponding post-processing control instructions.

[0007] Furthermore, S1 includes: Normal zinc spangle texture samples generated under different galvanizing process parameters were collected. Fourier transform was performed on each sample to extract its amplitude spectrum in the spatial frequency domain. Based on the correspondence between different process parameters and the amplitude spectrum, a zinc spangle texture spectrum model was established. Using a random noise vector and a crystal grain size parameter vector sampled from the spectral model as joint inputs, a simulated zinc flower background image with specified crystal grain size distribution characteristics is generated through a generative adversarial network. Obtain the three-dimensional geometric morphology parameters of the target defect, and construct a micro-element bidirectional reflection distribution function model describing the reflected light intensity distribution of the target defect under different incident angles and viewing angles; Based on the preset arrangement parameters of multi-angle light sources in the actual production line, at least two different incident directions of light are simulated, and the reflection image generated by the micro-facet bidirectional reflection distribution function model under the simulated light is obtained. The reflected image and the background image are physically rendered and fused. During the fusion, the embedding position of the target defect, the pixel-level contour mask and the defect type are recorded simultaneously, and a synthetic training image set with truth labels is automatically generated.

[0008] Furthermore, it also includes: performing unsupervised clustering on the initial zinc flower images, removing abnormal samples in the clustering results that are close to known defect feature clusters, and retaining only samples that characterize the zinc flower crystallization features without surface defects.

[0009] Furthermore, S2 includes: Two adjacent image acquisition devices arranged around a square tube acquire first and second images sequentially with a preset acquisition time difference during the movement of the square tube. The first image and the second image are respectively input into a convolutional coding network with shared weights, and the corresponding first feature map and second feature map are extracted and concatenated to obtain a concatenated feature map. The stitched feature map is input into the displacement field prediction network, and a dense displacement field is output. Each element in the dense displacement field represents the amount of spatial displacement required for the corresponding pixel in the first feature map to align with the second feature map. The dense displacement field is decoupled into a global rigid transformation matrix and a residual displacement field. The global rigid transformation matrix is ​​used to compensate for the overall rigid motion of the square tube within the acquisition time difference. The residual displacement field is used to characterize the non-rigid offset of the square tube edge region caused by perspective distortion and local deformation. The global rigid transformation matrix and the residual displacement field are superimposed and applied to the first feature map. The first feature map is then meshed and sampled through a differentiable spatial transformation layer to generate a transformed first feature map that is registered and aligned with the second feature map in the feature space.

[0010] Furthermore, when predicting the dense displacement field, the displacement field prediction network introduces a preset spatial attention weight map. The spatial attention weight map has a higher response value in the corresponding square tube edge region than in the planar region, so as to guide the displacement prediction network to prioritize optimizing the alignment accuracy of the edge region.

[0011] Furthermore, the defects in the reconstructed square tube edge region in S3 include: Obtain the transformed first feature map and second feature map after S2 registration and alignment; Based on the known overlapping regions of the fields of view of two adjacent image acquisition devices, the edge local feature maps corresponding to the overlapping regions of the fields of view are extracted from the transformed first feature map and the second feature map, respectively. The edge local feature map of the transformed first feature map is concatenated with the edge local feature map of the second feature map, and input into the cross-view aggregation sub-network to output a cross-surface fusion feature map. The cross-surface fusion feature map represents the continuity features of defects between adjacent surfaces within the overlapping area of ​​the field of view. Based on the cross-surface fusion feature map, the defect representation of the square tube edge region is reconstructed, and it is determined whether the defect is a continuous defect that spans adjacent surfaces.

[0012] Furthermore, S3 includes: During the registration and alignment process from the first feature map to the second feature map, S2 obtains the mapping transformation matrix between the two-dimensional pixel coordinates and the three-dimensional surface points of the square tube, which is established synchronously. When the defect detection network detects a defect instance, it obtains the pixel-level segmentation mask corresponding to the defect instance in the transformed first feature map or second feature map. Using the mapping transformation matrix, the pixel-level segmentation mask is mapped to a unified three-dimensional space, and the complete three-dimensional representation of the defect instance is obtained by stitching it together in the three-dimensional space. The physical dimensions of the defect instance are calculated in the three-dimensional space, and the surfaces and edges spanned by the defect instance are determined in combination with the splicing shape. The continuous coordinates of the spanning surfaces and the defect type are output.

[0013] Furthermore, the construction of the domain knowledge graph in S4 includes: defining defect type entity, defect location entity, defect size entity, tolerance threshold entity, defect level entity, and post-processing process entity; the tolerance threshold entity is associated with the defect type entity and the defect location entity; the defect level entity is associated with the defect type entity, defect location entity, defect size entity, and defect tolerance threshold entity; and the post-processing process entity is associated with the defect level entity.

[0014] Furthermore, the compliance reasoning in S4 includes: The defect type, physical size, and cross-plane continuous coordinates output by S3 are instantiated into the defect type entity, defect size entity, and defect location entity in the domain knowledge graph. The type of the square tube surface region where the defect is located is determined based on the cross-plane continuous coordinates, and characterized by the defect location entity; Based on the association between the tolerance threshold entity and the defect type entity and defect location entity, query the tolerance threshold corresponding to the instantiated defect type and defect location; Based on the association relationship of the defect level entities, the defect level of the defect is determined by the instantiated defect type, defect location, defect size and the tolerance threshold; Based on the association between the post-processing entity and the defect level entity, corresponding post-processing control instructions are generated according to the defect level.

[0015] A machine vision-based online detection system for surface defects of square tubes used in solar energy brackets employs any of the aforementioned machine vision-based online detection methods for surface defects of square tubes used in solar energy brackets.

[0016] Compared with the prior art, the beneficial effects of the present invention are: A machine vision-based online detection method and system for surface defects of square tube solar panel brackets is proposed. By combining a physical imaging model based on crystal grain size spectrum with a generative adversarial network to generate synthetic data, compliance reasoning of zinc flower texture is achieved by introducing a domain knowledge graph that incorporates process tolerance rules. This opens up a closed loop from the initial perception of defects to the automated handling decision on the production line, thereby enhancing the industrial practical value of the detection system.

[0017] Meanwhile, by combining a physical imaging model based on crystal grain size spectrum with a generative adversarial network to generate synthetic data, accurate distinction between zinc flower texture and real defects is achieved, significantly reducing false detection and false negative rates. Through three-dimensional mapping based on alignment process and cross-view feature aggregation, the defect type, precise physical size, and cross-plane continuous coordinates can be output at once, providing a complete data foundation for automated processing on the production line. Through multi-view spatiotemporal feature alignment, displacement field prediction with rigid and non-rigid decoupling, and spatial attention mechanism for edge regions, the detection blind spot in the edge region of square tubes is eliminated, achieving clear imaging and accurate detection of the entire surface of square tubes under high-speed movement. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall process of the online surface defect detection method of the present invention; Figure 2 This is a schematic diagram of the zinc flower texture spectrum model and the synthesis training image generation process of the present invention; Figure 3 This is a schematic diagram of the multi-view acquisition and spatiotemporal feature alignment method of the present invention; Figure 4 This is a schematic diagram of the cross-view aggregation and edge defect reconstruction process of this invention. Figure 5 This is a schematic diagram of the compliance reasoning process of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] like Figure 1 As shown, the present invention provides a technical solution: an online detection method for surface defects of square tube solar panel brackets based on machine vision, comprising: S1. Collect normal zinc flower texture of square tube, construct a zinc flower texture spectrum model to characterize the difference in crystal grain size distribution, and generate a simulated background image; establish a micro-facet bidirectional reflection distribution function model of the target defect, embed the target defect into the background image through physical rendering, generate a synthetic training image set, and pre-train the defect detection network.

[0021] like Figure 2 As shown, this invention provides a zinc flower texture spectrum model and a method for generating synthetic training images; Specifically: Step 1: Collect normal zinc flower texture samples generated under different galvanizing process parameters, perform Fourier transform on each sample to extract its amplitude spectrum in the spatial frequency domain, and establish a zinc flower texture spectrum model based on the correspondence between different process parameters and the amplitude spectrum.

[0022] It is important to note that samples of hot-dip galvanized square tubes produced under different galvanizing process parameters should be collected. These different galvanizing process parameters include, but are not limited to, zinc bath temperature (e.g., 440℃, 460℃, 480℃), immersion time (e.g., 30 seconds, 60 seconds, 90 seconds), and cooling rate (e.g., air cooling, natural cooling, water mist cooling). Different process parameters will produce zinc flower textures with different crystal grain sizes. For example, higher zinc bath temperature and slower cooling rate will produce large-grained zinc flowers, while vice versa, fine-grained zinc flowers will be produced.

[0023] Surface images of each sample were captured under diffuse uniform illumination to obtain an initial set of zinc flower images. Unsupervised clustering was then performed on the initial zinc flower images: the images were divided into image blocks of fixed size, texture features were extracted and clustered, and abnormal samples that were close to known defect feature clusters were removed, retaining only normal image blocks without defects as normal zinc flower texture samples.

[0024] Perform a two-dimensional discrete Fourier transform on each sample to obtain a complex matrix F(a,b), where a is the frequency coordinate in the horizontal direction and b is the frequency coordinate in the vertical direction. Take its magnitude to obtain the amplitude spectrum |F(a,b)|.

[0025] To establish the correlation with crystal grain size, the radial average curve A(r) of the amplitude spectrum is extracted as a feature, and the calculation method is as follows:

[0026] Where r is the frequency radius, representing the distance from the frequency point to the center of the spectrum (DC component), r≥0; θ is the polar angle, the angle of counterclockwise rotation from the center of the spectrum as the origin, θ∈[0,2π); N r Let F(a,b) be the number of pixels in the amplitude spectrum that satisfy the frequency radius r within the interval [r-(△r / 2), r+(△r / 2)], where △r is the preset frequency radius sampling step size (e.g., △r = 1 pixel unit); |F(rcosθ, rsinθ)| is the amplitude obtained by interpolating the amplitude spectrum |F(a,b)| at the polar coordinate position (rcosθ, rsinθ). Since rcosθ and rsinθ are usually not integers, and the amplitude spectrum |F(a,b)| is defined on discrete integer coordinates (a,b), it is necessary to obtain the amplitude at this position through interpolation (e.g., bilinear interpolation).

[0027] The shape of the A(r) curve differs under different process parameters. The energy of large-grain zinc spangles is concentrated in the low-frequency region, and the A(r) value is larger when r is small, and it decays rapidly as r increases; the energy of small-grain zinc spangles diffuses to high frequencies, and the A(r) decays more slowly.

[0028] Extract the feature parameters of A(r), including the peak frequency r. p A(r) represents the frequency radius at which it reaches its maximum value; bandwidth w b A(r) represents the frequency width at which A(r) drops 3 dB from its peak value; the attenuation coefficient β is the attenuation constant obtained by exponentially fitting A(r) to the high-frequency band. Different process parameters and characteristic parameters (r) are established. p w b The mapping relationship between β and β was used to complete the construction of the zinc flower texture spectrum model.

[0029] Step 2: Using a random noise vector and a crystal grain size parameter vector sampled from the spectrum model as joint inputs, a simulated zinc flower background image with specified crystal grain size distribution characteristics is generated through a generative adversarial network.

[0030] It should be noted that the generative adversarial network includes a generator G and a discriminator D.

[0031] G takes the random noise vector z and the crystal grain size parameter vector c as joint inputs and outputs the simulated zinc flower background image I. bg :

[0032] Where z is sampled from a standard normal distribution, and c contains r sampled from the spectral model. p w b β. G is composed of stacked transposed convolutional layers, progressively upsampled to the target resolution.

[0033] D receives real zinc flower images I real and the generated I bg The system outputs the discrimination result and predicts the crystal grain size parameters corresponding to the input image through an auxiliary head to ensure that the crystal grain size of the generated image is consistent with the input condition c.

[0034] The training loss function is: Among them, L adv To combat losses; L cond λ is the mean square error between the crystal grain size parameter predicted by D and the input condition c; λ is the balance coefficient, λ>0.

[0035] After training, G can generate a simulated zinc flower background image with corresponding zinc flower texture features based on any specified crystal grain size parameter vector c.

[0036] Step 3: Obtain the three-dimensional geometric morphology parameters of the target defect, and construct a micro-element bidirectional reflection distribution function model describing the reflected light intensity distribution of the target defect under different incident angles and viewing angles.

[0037] It is important to note that for the target defect, a laser confocal microscope is used to acquire its three-dimensional morphology, and corresponding geometric morphology parameters are extracted based on its morphological type. The morphological types of the target defect include: concave defects, such as pits and depressions; the extracted parameters include the concave depth h. d Opening radius r a For raised defects, such as zinc nodules and bubbles, the extracted parameters include the raised height h. r Bottom radius r b For peeling defects, such as flaking, the extracted parameters include the peeling thickness h. p , peeling area s p For crack-like defects, such as cracks, the extracted parameters include the crack depth d. c Length l c Opening width w c For coating defects, such as incomplete coating, the extracted parameters include the area s of the incomplete coating region. e Edge transition width w e .

[0038] Based on the aforementioned geometric topography parameters, a surface height function H(x,y) is constructed for each defect, where (x,y) represents the local coordinates of the defect surface. The H(x,y) for different defect morphology types is determined by their geometric topography parameters. For example: Defects such as dents:

[0039] The origin of the coordinate system is located at the center of the defect; Raised defects:

[0040] The origin of the coordinate system is located at the center of the defect; Cracking defects:

[0041] Where x is a local coordinate perpendicular to the crack propagation direction.

[0042] The roughness parameter σ of the defect surface is determined by the root mean square value of the local gradient of H(x,y):

[0043] Where A is the area of ​​the defect surface region; ▽H(x,y) is the gradient vector of H(x,y), calculated by the following formula:

[0044] The calculated roughness parameter σ is substituted into the following bidirectional reflection distribution function (Cook-Torrance BRDF model) to describe the reflection characteristics of the defect:

[0045] Where, ω i ω is the incident light direction vector; o θ is the direction vector of the emitted light; i θ is the angle of incidence; o F is the angle of departure; sch The Fresnel reflection coefficient is given by the Schlick approximation:

[0046] Among them, F o For normal incident reflectance, the zinc coating region F o ≈0.45; θ h ω is the direction of the incident light i With respect to the direction of emitted light ω o Half of the included angle, that is: D ggx The distribution function of the micro-element normals is given by the GGX model:

[0047] G sm The geometric decay factor is determined using the Smith model:

[0048]

[0049] Where θ is the incident angle θ i Or the angle of emission θ o ; ω is the corresponding ω i or ω o Step 4: Based on the preset arrangement parameters of multi-angle light sources in the actual production line, simulate at least two different incident directions of illumination, and generate the reflection image of the micro-surface element bidirectional reflection distribution function model under the simulated illumination.

[0050] It is important to obtain the light source arrangement parameters corresponding to the image acquisition devices in the actual production line, including the incident angle, azimuth angle, and light source type of each light source. In a typical layout, each image acquisition device is equipped with two LED line light sources, which illuminate the device at incident angles of 15° and 45°, respectively.

[0051] A virtual scene is created in the rendering engine, comprising: a 3D geometric model of a square tube, covering planar and edge regions; a defect model with assigned material properties to the BRDF; and at least two virtual light sources consistent with the production line layout. A path tracing algorithm is used on each virtual light source to simulate the physical process of light reflecting off the defective BRDF surface and reaching the virtual camera sensor, generating a defect reflection image I. defect .

[0052] Step 5: Perform physical rendering fusion of the reflected image and the background image. During fusion, the embedding position of the target defect, the pixel-level contour mask and the defect type are recorded simultaneously, and a synthetic training image set with ground truth annotations is automatically generated.

[0053] It is important to note that the defect reflection image generated in step 4 is used as the foreground, and the simulated zinc flower background image generated in step 2 is used as the background for physical rendering fusion. During fusion, ensure that the lighting conditions of the foreground and background are consistent, and that the defect placement position is consistent with the preset spatial coordinates. Alpha blending is used at the boundary between the foreground and background to achieve a smooth transition.

[0054] Fusion Image I fused The values ​​of each pixel p are:

[0055] Where γ(p) is the fusion weight at pixel p, γ(p) = 1 inside the defect contour, and gradually changes from 1 to 0 at the edge transition zone; I defect (p) is the defect reflection image I defect The pixel value at pixel p; I bg (p) is a simulated zinc flower background image I bg The pixel value at pixel p.

[0056] During the fusion process, ground truth labels are automatically generated simultaneously, recording the embedding location of defects, pixel-level contour masks, and defect types, forming a synthetic training image set with attached ground truth labels.

[0057] Optionally, after pre-training the defect detection network with the synthetic training image set, a domain adaptation step is also included: Collect a small number of standard images of the surface of real square tubes from the production line to form a target domain sample set; The pre-trained network is fine-tuned on the target domain sample set; Alternatively, a domain adversarial training strategy can be adopted in the pre-training stage. A gradient reversal layer and a domain classifier are introduced into the feature extraction network. The gradient reversal layer performs an identity mapping during forward propagation and multiplies the gradient by -1 during back propagation, causing the feature extraction network to update in the direction of increasing the domain classification error. This extracts features that are invariant to the synthetic domain and the real domain, thereby improving the generalization ability of the pre-trained network to real production line scenarios.

[0058] S2. Obtain multi-view temporal image sequences surrounding the square tube, predict the feature spatial displacement field and non-rigid deformation caused by the movement of the square tube, and register and align the multi-view feature maps.

[0059] like Figure 3 As shown, this invention provides a method for multi-view acquisition and spatiotemporal feature alignment; Specifically: Step 1: Acquire two adjacent image acquisition devices arranged around the square tube, and acquire the first and second images sequentially with a preset acquisition time difference during the movement of the square tube.

[0060] It should be noted that multiple image acquisition devices are arranged around the square tube, with their fields of view covering the upper surface, lower surface, left side and right side of the square tube, respectively, and the fields of view of adjacent image acquisition devices overlap in the edge area of ​​the square tube.

[0061] During the high-speed movement of the square tube along the production line, a first image I1 and a second image I2 are acquired sequentially by two adjacent image acquisition devices with a preset acquisition time difference Δt. Δt is determined based on the square tube's speed v and the spatial baseline distance B between the adjacent image acquisition devices, i.e.:

[0062] Step 2: Input the first image and the second image into a convolutional coding network with shared weights, extract the corresponding first feature map and second feature map, and concatenate them to obtain a concatenated feature map.

[0063] It is important to note that I1 and I2 are input into a convolutional coding network E with shared weights to extract the corresponding first feature map F1 and second feature map F2:

[0064]

[0065] In this context, F1 and F2 are both three-dimensional tensors of size H×W×C, where H is the feature map height, W is the feature map width, and C is the number of feature channels.

[0066] F1 and F2 are concatenated along the channel dimension to obtain the concatenated feature map F. cat The dimensions are H×W×2C.

[0067] Step 3: Input the stitched feature map into the displacement field prediction network and output a dense displacement field. Each element in the dense displacement field represents the amount of spatial displacement required for the corresponding pixel in the first feature map to align with the second feature map.

[0068] It should be noted that the displacement field prediction network P adopts an encoder-decoder structure. The encoder gradually reduces F. cat The spatial resolution is increased and the number of channels is increased to extract multi-scale contextual information; the decoder gradually restores the spatial resolution through upsampling and skip connections. The last layer of P outputs a two-channel feature map, representing the horizontal and vertical displacement components, respectively, which is the dense displacement field. .

[0069] The size is H×W×2, where Φ=(△x,△y) means that the pixel at coordinate (x,y) in the first feature map F1 needs to be moved △x pixels horizontally and △y pixels vertically to align with the corresponding position in the second feature map F2.

[0070] To improve the alignment accuracy of edge regions, the displacement field prediction network P predicts dense displacement fields. At that time, a pre-defined spatial attention weight map W is introduced. att W att The size of W is consistent with the spatial size H×W of the feature map, and its response value at the pixel position in the corresponding square tube edge region is higher than that in the planar region. Specifically, W att The weights in the edge region are set to η (η > 1), and the weights in the planar region are set to 1. In the loss function of P, W... att Element-wise multiplication with the pixel-by-pixel alignment error allows the network to prioritize optimizing the alignment accuracy of edge regions during training.

[0071] Among them, L align Let P be the alignment loss function value for the displacement field prediction network. Let P be the displacement vector of the dense displacement field predicted by the displacement prediction network P at coordinates (x, y). The displacement vector of the real displacement field at coordinates (x, y) is generated by the known geometric transformation parameters used when S1 synthesizes the training image; This represents the L2 norm of a vector.

[0072] Step 4: Decouple the dense displacement field into a global rigid transformation matrix and a residual displacement field.

[0073] It is important to note that By performing least-squares fitting on the displacement vectors of all pixels in the dense displacement field, the global rigid transformation matrix T representing rotation and translation is estimated. rigid :

[0074] Where δ is the rotation angle of the square tube around the optical axis during the acquisition time Δt; t x t y These are the translation components in the horizontal and vertical directions, respectively; T rigid Used to compensate for the overall rigid body motion of the square tube during the acquisition time difference.

[0075] Will The displacement vector of each pixel (x, y) minus T rigid The residual displacement field is obtained by applying the displacement component to the pixel. :

[0076] in, Let be the displacement vector of the dense displacement field at coordinates (x, y); [x, y, 1] T Homogeneous coordinates for pixel coordinates; For dense fields, and With the same dimensions, this characterizes the non-rigid offset of the square tube's edge region caused by perspective distortion and local deformation. In the planar region... The value approaches zero, exhibiting local nonlinear changes in the edge region.

[0077] Step 5: Superimpose the global rigid transformation matrix and the residual displacement field, apply them to the first feature map, and perform mesh sampling on it through a differentiable spatial transformation layer to generate a transformed first feature map that is registered and aligned with the second feature map in the feature space.

[0078] It is important to note that T rigid and The superposition is applied to the first feature map F1, and a transformed first feature map aligned with the second feature map F2 is generated through a spatial transformation layer. . Specifically: Generate a sampling grid U with the same size as the F1 space. The coordinates of the sampling point at each coordinate position (x, y) in U are obtained by superimposing the original coordinates T. rigid and The displacement is obtained as follows:

[0079] in, The coordinates of the sampling point at coordinates (x, y) in the sampling grid U are transformed. Let x be the horizontal displacement component of the residual displacement field at coordinates (x, y); Let be the vertical displacement component of the residual displacement field at coordinates (x, y).

[0080] Based on the coordinates in the sampling grid U Bilinear interpolation sampling is performed on F1 to obtain the transformed first feature map. :

[0081] The grid_sample(·) function represents a bilinear interpolation operation, which performs a linear weighted summation of the values ​​of the surrounding four integer coordinate pixels to obtain the final sampled value.

[0082] After the above spatial transformation, The two are registered and aligned in the feature space with F2, and their edge region features correspond precisely in space.

[0083] For each pair of images acquired by adjacent image acquisition devices arranged around the square tube, steps 1 to 5 are performed to obtain multi-view feature maps that are registered and aligned between all adjacent viewpoints.

[0084] S3. Input the registered and aligned multi-view feature map into the defect detection network, aggregate the features of the overlapping areas of adjacent view fields to reconstruct the defects in the edge area of ​​the square tube, and output the defect type, physical size and cross-plane continuous coordinates based on the three-dimensional mapping established by the alignment process.

[0085] like Figure 4 As shown, the present invention provides a method for cross-view aggregation and edge defect reconstruction; Specifically: Step 1: Obtain the transformed first feature map and second feature map after S2 registration and alignment.

[0086] It should be noted that for each pair of adjacent image acquisition devices arranged around the square tube, the corresponding F2 is obtained, resulting in four pairs of adjacent viewpoint aligned feature maps, corresponding to the four adjacent viewpoints: upper surface-right side, right side-lower surface, lower surface-left side, and left side-upper surface.

[0087] Step 2: Based on the known overlapping areas of the fields of view of the two adjacent image acquisition devices, extract the edge local feature maps corresponding to the overlapping areas of the fields of view from the transformed first feature map and the second feature map, respectively.

[0088] It is important to note that the overlapping area of ​​the field of view is predetermined during the calibration phase. During calibration, a square tube calibration piece with a checkerboard pattern of known size is used to image each image acquisition device, determining the intrinsic and extrinsic parameter matrices of each image acquisition device. Then, the boundary of the overlapping area of ​​the field of view of adjacent image acquisition devices on the surface of the square tube is calculated. This overlapping area corresponds to the edge area of ​​the square tube and a strip area extending a certain width (e.g., 5mm-10mm) on both sides.

[0089] Let the pixel range of the overlapping field of view in the feature map space be: horizontal direction [x a x b ], vertical direction [y a y b ].from By extracting feature patches within this range, the local feature map F of the first-view edge is obtained. 1,edge :

[0090] Among them, [x a :x b y a y b, ,:] indicates taking the x-th value in the height dimension of the feature map. a To x b The row and width dimensions are taken as the y-th dimension. a To y b Column, retain all channels. Similarly, extract feature patches with the same coordinate range from F2 to obtain the second-view edge local feature map F. 2,edge : Among them, F 2,edge and F 1,edge The dimensions are the same, all are H e ×W e ×C,H e H represents the height of the local feature map of the edge. e =x b -x a W e W represents the width of the local feature map of the edge. e =y b -y a C represents the number of feature channels.

[0091] Step 3: Concatenate the edge local feature map of the transformed first feature map with the edge local feature map of the second feature map, input the cross-view aggregation sub-network, and output the cross-surface fusion feature map. The cross-surface fusion feature map represents the continuity feature of the defect between adjacent surfaces within the field of view overlap area.

[0092] It should be noted that F 1,edge and F 2,edge Perform cross-view aggregation and output a cross-face fusion feature map F for defect detection. fuse ; while retaining F 1,edge and F 2,edge . Specifically: F 1,edge and F 2,edge The concatenation is performed along the channel dimension to obtain the concatenated feature map F.cat .

[0093] F cat Input cross-view aggregation sub-network A, output cross-view fusion feature map F fuse :

[0094] Among them, F fuse The size is H e ×W e ×C f C f Let F be the number of feature channels after fusion. A consists of several stacked convolutional layers used to learn the complementary relationship between features from two adjacent viewpoints, so that F... fuse The feature vector at each spatial location in the model simultaneously fuses semantic information from two perspectives, representing the continuity of defects between adjacent surfaces within the overlapping field of view. An example structure for A is: a 3×3 convolutional layer (2C input channels, C output channels). f (Step size 1, padding 1), followed by a batch normalization layer and a ReLU activation function.

[0095] Step 4: Input the cross-surface fusion feature map into the detection head of the defect detection network, reconstruct the defect representation of the square tube edge region, and determine whether the defect is a continuous defect that spans adjacent surfaces.

[0096] It is important to note that the detection head includes a classification branch and a segmentation branch. The classification branch outputs a defect category probability vector, the dimension of which is equal to the preset number of defect categories (dents, bulges, peelings, cracks, coating defects, etc.); the segmentation branch outputs a pixel-level defect prediction mask M. edge Dimensions are H e ×W e The value for each pixel is the confidence level that the pixel belongs to a defect.

[0097] For M edge For each candidate defect region detected, it is determined whether it is a continuous defect spanning adjacent surfaces. Specifically, for M... edge In a certain connected region, from F 1,edge and F 2,edge The local feature vector sets corresponding to the connected region are extracted from each set, and the similarity between the feature vectors at corresponding spatial locations in the two sets is calculated, such as by calculating cosine similarity. The average cosine similarity of all pixels in the connected region is calculated. If the average value is higher than a preset threshold (e.g., 0.7), the defect is determined to be a continuous defect that spans adjacent surfaces; otherwise, it is determined to be a discontinuous defect that exists only on a single surface.

[0098] For instances determined to be continuous defects, the F retained in step 3 is used.1,edge and F 2,edge M edge The regions corresponding to the defects are projected back to F respectively. 1,edge and F 2,edge The local mask M is obtained from each viewpoint. 1,d and M 2,d .

[0099] The 3D mapping output based on the alignment process includes the defect type, physical size, and cross-plane continuous coordinates. Step 1: Obtain the mapping transformation matrix between the two-dimensional pixel coordinates and the three-dimensional surface points of the square tube, which is synchronously established during the registration and alignment process from the first feature map to the second feature map in S2.

[0100] It is important to note that during the calibration phase, in addition to determining the intrinsic and extrinsic parameter matrices of each image acquisition device, a three-dimensional geometric model of the square tube was also established. This three-dimensional geometric model of the square tube is a cuboid of known dimensions, comprising four planar regions (upper surface, lower surface, left side, and right side) and four edge regions.

[0101] For each image acquisition device, its imaging satisfies the pinhole camera model:

[0102] Where (u, v) are pixel coordinates; s is the scale factor; (X, Y, Z) are the spatial coordinates of points on the 3D surface of the square tube; and K is a 3×3 intrinsic parameter matrix.

[0103] Among them, f x f is the equivalent focal length in the horizontal direction. y c is the equivalent focal length in the vertical direction; x c y These are the coordinates of the principal point of the image.

[0104] [R|t] is a 3×4 extrinsic parameter matrix:

[0105] Where, r 11 to r 33 t are elements of the rotation matrix R; x t y t z These are the three components of the translation vector t.

[0106] Using the pinhole camera models of two adjacent image acquisition devices, two projection equations are established simultaneously. Then, using the pixel correspondence between the first feature map F1 and the second feature map F2 established in S2, the coordinates (X, Y, Z) of the corresponding three-dimensional surface point of the square tube for each pixel pair are solved. A mapping transformation matrix M from two-dimensional pixel coordinates (u, v) to three-dimensional surface point coordinates (X, Y, Z) is established. 3D The mapping transformation matrix is ​​stored in the form of a lookup table, and for each pixel position in the feature map space, the coordinates of the corresponding three-dimensional surface point of the square tube are recorded.

[0107] Step 2: When the defect detection network detects a defect instance, obtain the pixel-level segmentation mask corresponding to the defect instance in the transformed first feature map or second feature map.

[0108] It should be noted that for instances determined to be a single surface defect, the value at that location should be directly taken. Or the corresponding M in F2 edge The region, denoted as M, serves as a pixel-level segmentation mask for this viewpoint. single For instances determined to be continuous defects, M will be... 1,d and M 2,d According to its origin Alternatively, the coordinate offset during the truncation in F2 can be mapped back to the complete feature map space to obtain... On the mask M 1,D and the mask M on F2 2,D .

[0109] Step 3: Using the mapping transformation matrix, the pixel-level segmentation mask is mapped to a unified three-dimensional space, and the complete three-dimensional representation of the defect instance is obtained by stitching it together in the three-dimensional space.

[0110] It should be noted that for a single surface defect, M single Mapped to a 3D defect point cloud P def .

[0111] For continuous defects, M 1,D and M 2,D Mapped to P respectively 1,def and P 2,def After converting to the same world coordinate system, the coordinates are stitched together. For any 3D point within the overlapping region of the two viewpoints, a distance-weighted fusion method is used to calculate its fused position confidence level c. fuse :

[0112] Where c1 is the confidence level of the 3D point in the first viewpoint, and M is the confidence level of the point. 1,D The mask value of the corresponding pixel is obtained by bilinear interpolation; c2 is the confidence level of the 3D point in the second view, determined by M.2,D The mask value of the corresponding pixel is obtained by bilinear interpolation; φ1 is the angle between the optical axis of the first-view image acquisition device and the normal of the square tube surface where the three-dimensional point is located; φ2 is the angle between the optical axis of the second-view image acquisition device and the normal of the square tube surface where the three-dimensional point is located; cosφ1 and cosφ2 are weighting factors. The smaller the angle, the closer the observation direction is to the normal direction, the higher the observation confidence of the viewpoint, and the greater the weight.

[0113] After splicing, a complete three-dimensional representation P of the defect instance is obtained. def , is a collection of point clouds in three-dimensional space, where each point contains spatial coordinates (X, Y, Z) and a fused confidence score c. fuse .

[0114] Step 4: Calculate the physical dimensions of the defect instance in the three-dimensional space, and determine the surfaces and edges spanned by the defect instance in combination with the splicing shape, and output the continuous coordinates of the spanning surfaces and the defect type.

[0115] It is important to note that in P def The physical dimensions of defects are calculated. For area-type defects such as plating defects or peeling, the total surface area is calculated on the reconstructed triangular mesh; for linear defects such as scratches or cracks, the length is obtained by integrating along the three-dimensional skeleton; for point-type defects such as pits or dents, the major and minor axes of the smallest circumscribed ellipsoid are taken.

[0116] Output cross-surface coordinates based on defect type. For single-surface defects, the cross-surface coordinates are a sequence of local three-dimensional coordinates within that surface. For continuous defects, combine P... def Based on the splicing shape and three-dimensional geometric model of the square tube, determine the surface and edges crossed by the defect, and output the continuous coordinates of the cross surface: {surface s edge1, surface m ,...,edge k ,surface e ,{P1,...,P n}}. Among them, surface s For the starting face identifier, surface e For the termination face identifier, surface m The edge is the identifier for the face it passes through. k To identify the edges that are crossed, {P1,...,P n} represents the sequence of path point coordinates of the defect on the three-dimensional surface of the square tube.

[0117] At the same time, the defect type determined by the classification branch in step 4 is output.

[0118] S4. Instance-based input of defect type, physical size, and cross-plane continuous coordinates into a pre-built domain knowledge graph containing process tolerance rules, perform compliance reasoning, output defect level, and generate corresponding post-processing control instructions.

[0119] It should be noted that the domain knowledge graph includes the following entity types and their relationships.

[0120] The defined entities include: defect type entity, defect location entity, defect size entity, tolerance threshold entity, defect level entity, and post-processing entity.

[0121] The attributes of each entity include: The defect type entity contains a defect category attribute, with values ​​of dent, bulge, peel, crack, and coating loss, which correspond to the defect types output by S3. The defect location entity includes surface region type attributes, and is at least divided into planar regions and edge regions, determined by the cross-plane continuous coordinates output by S3; A defect size entity, which includes size attributes, including the surface area value of an area-type defect, the length value of a linear defect, or the major and minor axis values ​​of a point-type defect; The tolerance threshold entity contains tolerance threshold attributes, including area threshold, length threshold, and depth threshold; Defect level entity, which includes a level attribute, with values ​​of acceptable, repairable, or scrap; The post-processing entity includes process attributes, with values ​​of no processing, grinding and repair, inkjet marking, and rejection.

[0122] The relationships between entities include: The tolerance threshold entity is associated with both the defect type entity and the defect location entity, indicating that the same defect type has different tolerance thresholds in planar and edge regions. For example, the allowable area for plating defects in planar regions is no more than 3 mm. 2 No area is allowed in the edge region.

[0123] The defect level entity is associated with the defect type entity, defect location entity, defect size entity, and defect tolerance threshold entity, indicating that the determination of the defect level is jointly determined by the defect type, location, size, and corresponding tolerance threshold.

[0124] The post-processing entity is associated with the defect level entity, indicating that the selection of the post-processing process is determined by the defect level. Among them, the acceptable level corresponds to no processing, the repairable level corresponds to grinding and repair, and the scrap level corresponds to inkjet marking or rejection.

[0125] like Figure 5 As shown, this invention provides a compliance reasoning method; Specifically: Step 1: Instantiate the defect type, physical size, and cross-plane continuous coordinates output by S3 into the defect type entity, defect size entity, and defect location entity in the domain knowledge graph.

[0126] It should be noted that an instance of a defect type entity is created, and its defect category attribute is assigned the defect type output by S3; an instance of a defect size entity is created, and its size attribute is assigned the physical size value output by S3; an instance of a defect location entity is created, and its surface area type attribute will be assigned a value after step 2 is determined.

[0127] The three entity instances after instantiation are interconnected through defect relationship edges belonging to the same type, indicating that they jointly describe the same defect detection result.

[0128] Step 2: Determine the type of the square tube surface region where the defect is located based on the cross-plane continuous coordinates, and characterize it through the defect location entity.

[0129] It is important to note that the continuous coordinates across surfaces include the surface identifier sequence {surface}. s ,...,surface e} and edge identifier sequence {edge1,...,edge k If the edge identifier sequence is empty (i.e., k=0), the surface defect is located on only a single surface, and the defect is determined to be located in a planar region. The surface region type attribute of the entity at the defect location is then assigned to planar region. If the edge identifier sequence is not empty (i.e., k≥1), it indicates that the defect crosses at least one edge. The defect is determined to involve an edge region, and the surface region type attribute of the entity at the defect location is then assigned to edge region. At the same time, the list of involved edge identifiers is recorded.

[0130] Step 3: Based on the association between the tolerance threshold entity and the defect type entity and defect location entity, query the tolerance threshold corresponding to the instantiated defect type and defect location.

[0131] It is important to note that the query conditions are the defect category attribute of the defect type entity instance and the surface region type attribute of the defect location entity instance, which are used to match the corresponding defect tolerance threshold entity instances in the graph. During matching, a joint query is performed along the two associated edges: defect tolerance threshold entity → defect type entity and defect tolerance threshold entity → defect location entity. This retrieves defect tolerance threshold entity instances that simultaneously satisfy both defect type matching and location type matching, and their tolerance threshold attribute values ​​are extracted.

[0132] Step 4: Based on the association relationship of the defect level entities, determine the defect level of the defect by the instantiated defect type, defect location, defect size and the tolerance threshold.

[0133] It is important to note that the physical dimensions of the defect size entity instance are compared with the tolerance threshold, and the attribute values ​​of the defect level entity are determined according to a preset level classification rule. An example of the level classification rule is as follows: If the physical size is less than or equal to 50% of the tolerance threshold, it is considered acceptable. If the tolerance threshold × 50% < physical size ≤ tolerance threshold, and the defect type is repairable (such as scratch), it is determined to be repairable; If the physical size exceeds the tolerance threshold, it is considered a defective product. If the defect is located in an edge area and the physical size is greater than the tolerance threshold × 50%, it is judged as a scrap.

[0134] The determination result is assigned to the level attribute of the defect level entity instance.

[0135] Step 5: Based on the association between the post-processing process entity and the defect level entity, generate corresponding post-processing control instructions according to the defect level.

[0136] It is important to note that the post-processing entity instance associated with the defect level entity instance is queried to obtain the corresponding process attribute value. If the process attribute is "no processing," a release instruction is generated; if the process attribute is "grinding and repair," a sequence of path points {P1,...,P} containing repair coordinates (taken from the continuous coordinates across the surface) is generated. n The system generates grinding control instructions for the repair depth (taken from the depth value in the physical dimensions), and for marking control instructions for the process attribute of inkjet marking, including the marking coordinates and defect level. If the process attribute is rejection, it generates rejection instructions including the defect location. The generated instructions are packaged in a standard industrial communication protocol format and sent to the corresponding post-processing execution device.

[0137] A machine vision-based online detection system for surface defects of square tubes used in solar energy brackets employs any of the aforementioned machine vision-based online detection methods for surface defects of square tubes used in solar energy brackets. The system includes: a multi-view image acquisition module for acquiring multi-view temporal image sequences surrounding the square tube; a physical-guided synthesis training module for executing S1; a spatiotemporal feature alignment module for executing S2; a cross-view aggregation and 3D quantization module for executing S3; and a knowledge graph reasoning and decision-making module for executing S4.

[0138] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended embodiments and their equivalents.

Claims

1. A machine vision-based online detection method for surface defects of square tubes of a solar support, characterized in that, include: S1. Collect normal zinc flower texture of square tube, construct a zinc flower texture spectrum model to characterize the difference in crystal grain size distribution, and generate a simulated background image; A two-way reflection distribution function model of the micro-facet of the target defect is established. The target defect is embedded into the background image through physical rendering to generate a synthetic training image set and pre-train the defect detection network. S2. Obtain multi-view temporal image sequences surrounding the square tube, predict the characteristic spatial displacement field and non-rigid deformation caused by the movement of the square tube, and register and align the multi-view feature maps. S3. Input the registered and aligned multi-view feature map into the defect detection network, aggregate the features of the overlapping areas of adjacent view fields to reconstruct the defects in the edge area of ​​the square tube, and output the defect type, physical size and cross-plane continuous coordinates based on the three-dimensional mapping established by the alignment process. S4. Instance-based input of defect type, physical size, and cross-plane continuous coordinates into a pre-built domain knowledge graph containing process tolerance rules, perform compliance reasoning, output defect level, and generate corresponding post-processing control instructions.

2. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: S1 includes: S11. Collect normal zinc flower texture samples generated under different galvanizing process parameters, perform Fourier transform on each sample to extract its amplitude spectrum in the spatial frequency domain, and establish a zinc flower texture spectrum model based on the correspondence between different process parameters and the amplitude spectrum. S12. Using a random noise vector and a crystal grain size parameter vector sampled from the spectrum model as joint inputs, a simulated zinc flower background image with specified crystal grain size distribution characteristics is generated through a generative adversarial network. S13. Obtain the three-dimensional geometric morphology parameters of the target defect, and construct a micro-element bidirectional reflection distribution function model describing the reflected light intensity distribution of the target defect under different incident angles and viewing angles; S14. Based on the preset arrangement parameters of multi-angle light sources in the actual production line, simulate at least two different incident directions of illumination, and generate the reflection image of the micro-surface element bidirectional reflection distribution function model under the simulated illumination. S15. The reflected image and the background image are physically rendered and fused. During the fusion, the embedding position of the target defect, the pixel-level contour mask and the defect type are recorded simultaneously, and a synthetic training image set with truth labels is automatically generated.

3. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 2, characterized in that: S11 further includes: performing unsupervised clustering on the acquired initial zinc flower images, removing abnormal samples in the clustering results that are close to known defect feature clusters, and retaining only samples that characterize the zinc flower crystallization features without surface defects.

4. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: S2 includes: S21. Acquire two adjacent image acquisition devices arranged around the square tube, and acquire the first image and the second image sequentially with a preset acquisition time difference during the movement of the square tube. S22. Input the first image and the second image into a convolutional coding network with shared weights, extract the corresponding first feature map and second feature map, and concatenate them to obtain a concatenated feature map. S23. Input the stitched feature map into the displacement field prediction network and output a dense displacement field. Each element in the dense displacement field represents the amount of spatial displacement required for the corresponding pixel in the first feature map to align with the second feature map. S24. Decouple the dense displacement field into a global rigid transformation matrix and a residual displacement field. The global rigid transformation matrix is ​​used to compensate for the overall rigid motion of the square tube within the acquisition time difference. The residual displacement field is used to characterize the non-rigid offset of the square tube edge region caused by perspective distortion and local deformation. S25. The global rigid transformation matrix and the residual displacement field are superimposed and applied to the first feature map. The first feature map is then meshed and sampled through a differentiable spatial transformation layer to generate a transformed first feature map that is registered and aligned with the second feature map in the feature space.

5. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 4, characterized in that: When predicting the dense displacement field, the displacement field prediction network introduces a preset spatial attention weight map. The spatial attention weight map has a higher response value in the corresponding square tube edge region than in the planar region, so as to guide the displacement prediction network to prioritize optimizing the alignment accuracy of the edge region.

6. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: The defects in the reconstructed square tube edge region in S3 include: S31. Obtain the transformed first feature map and second feature map after registration and alignment in S2; S32. Based on the known overlapping regions of the fields of view of two adjacent image acquisition devices, extract the edge local feature maps corresponding to the overlapping regions of the fields of view from the transformed first feature map and the second feature map, respectively. S33. The edge local feature map of the transformed first feature map is concatenated with the edge local feature map of the second feature map, and input into the cross-view aggregation sub-network to output the cross-surface fusion feature map. The cross-surface fusion feature map represents the continuity features of the defects between adjacent surfaces within the overlapping area of ​​the field of view. S34. Based on the cross-surface fusion feature map, reconstruct the defect representation of the square tube edge region, and determine whether the defect is a continuous defect that spans adjacent surfaces.

7. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: S3 includes: S35. Obtain the mapping transformation matrix between the two-dimensional pixel coordinates and the three-dimensional surface points of the square tube, which was synchronously established during the registration and alignment process from the first feature map to the second feature map in S2. S36. When the defect detection network detects a defect instance, the pixel-level segmentation mask corresponding to the defect instance in the transformed first feature map or second feature map is obtained. S37. Using the mapping transformation matrix, the pixel-level segmentation mask is mapped to a unified three-dimensional space, and the complete three-dimensional representation of the defect instance is obtained by stitching it together in the three-dimensional space. S38. Calculate the physical dimensions of the defect instance in the three-dimensional space, and determine the surfaces and edges crossed by the defect instance in combination with the splicing shape, and output the continuous coordinates of the cross surface and the defect type.

8. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: The construction of the domain knowledge graph in S4 includes: defining defect type entity, defect location entity, defect size entity, tolerance threshold entity, defect level entity, and post-processing process entity; the tolerance threshold entity is associated with the defect type entity and the defect location entity; the defect level entity is associated with the defect type entity, defect location entity, defect size entity, and defect tolerance threshold entity; and the post-processing process entity is associated with the defect level entity.

9. The online detection method for surface defects of square tube solar panel brackets based on machine vision according to claim 1, characterized in that: The compliance reasoning in S4 includes: S41. Instantiate the defect type, physical size and cross-plane continuous coordinates output by S3 into the defect type entity, defect size entity and defect location entity in the domain knowledge graph. S42. Determine the type of the square tube surface region where the defect is located based on the cross-plane continuous coordinates, and characterize it through the defect location entity; S43. Based on the association between the tolerance threshold entity and the defect type entity and defect location entity, query the tolerance threshold corresponding to the instantiated defect type and defect location; S44. Based on the association relationship of the defect level entities, the defect level of the defect is determined by the instantiated defect type, defect location, defect size and the tolerance threshold. S45. Based on the association between the post-processing process entity and the defect level entity, generate corresponding post-processing control instructions according to the defect level.

10. An online detection system for surface defects of square tube solar panel brackets based on machine vision, characterized in that, The method for online detection of surface defects in square tube solar panel brackets based on machine vision, as described in any one of claims 1-9, includes: A multi-view image acquisition module is used to acquire multi-view time-domain image sequences surrounding the square tube; The physical-guided synthesis training module is used to collect normal zinc flower textures of square tubes, construct a zinc flower texture spectrum model that characterizes the differences in crystal grain size distribution, and generate a simulated background image; and to establish a micro-facet bidirectional reflection distribution function model of the target defect, embed the target defect into the background image through physical rendering, generate a set of synthetic training images, and pre-train the defect detection network. The spatiotemporal feature alignment module is used to predict the feature spatial displacement field and non-rigid deformation caused by the motion of the square tube, and to register and align the feature maps from multiple perspectives. The cross-view aggregation and 3D quantization module is used to input the registered and aligned multi-view feature maps into the defect detection network, aggregate the features of the overlapping areas of adjacent view fields to reconstruct the defects in the edge area of ​​the square tube, and output the defect type, physical size and cross-plane continuous coordinates based on the 3D mapping established by the alignment process. The knowledge graph reasoning and decision-making module is used to input the defect type, physical size, and cross-plane continuous coordinates into a pre-built domain knowledge graph containing process tolerance rules, perform compliance reasoning, output the defect level, and generate corresponding post-processing control instructions.