A three-dimensional model defect detection method and system based on image recognition
Patent Information
- Application Number
- CN202610843133.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]传统检测方法在获取二维影像并比对标准投影矩阵标定位置时,由于成像阶段高曲率边缘区域发生显著投影压缩使得表面微小缺陷对应极少像素且反光大幅增强,物理变化导致缺陷纹理压缩成弱信号被高亮边界掩盖,常规框架难区分几何投影引发的特征畸变易将微弱异常误判为正常边缘导致候选响应不足,二维坐标受高亮边界牵引偏离真实中心在回映至三维空间时产生位置偏移,制约缺陷检出率与三维位置绑定精度
空间定位模块,基于所述候选缺陷二维坐标映射的三维初始坐标构建空间搜索线段,并反向投影生成局部纹理法向向量,计算所述空间搜索线段与内部三角面片交点的面片法向量与局部纹理法向向量之间的点积数值,提取所述点积数值大于预设正向阈值且空间欧氏距离最小的几何交点作为有效交点,并基于所述有效交点生成三维缺陷绑定位置。
Smart Images

Figure CN122675809A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a method and system for detecting defects in three-dimensional models based on image recognition. Background Technology
[0002] The field of image recognition technology mainly involves the core matters of extracting, analyzing, and classifying specific target feature patterns or states in input images. This field covers a complete system from image acquisition and preprocessing to feature extraction and target classification. It uses computers to transform multidimensional visual signals into measurable numerical matrices, and obtains the inherent spatial structure and pixel brightness distribution patterns of images through specific mathematical operations such as edge extraction, texture analysis, or pixel gradient calculation, thereby constructing a mapping relationship between computer vision and real physical objects. Traditional image recognition-based 3D model defect detection methods refer to the technical matters of identifying and judging abnormal states such as damage, deformation, or holes on the surface and structure of 3D solid models during manufacturing or generation. Traditional methods usually use industrial cameras set up at multiple angles to take 2D images of the 3D object under test from multiple perspectives. Then, edge detection operators or gray-level co-occurrence matrices are used to extract surface texture and contour feature pixels from the captured 2D images. Next, the extracted current pixel array is compared pixel by pixel with the 2D reference pixel matrix generated by the projection of a pre-set defect-free standard 3D model. The location of defects in the 3D model is marked according to the specific pixel area where the pixel coordinate deviation or gray-level difference between the two exceeds a set threshold.
[0003] Traditional detection methods suffer from several drawbacks when acquiring two-dimensional images and comparing them to standard projection matrices for positioning. During the imaging stage, significant projection compression occurs in high-curvature edge regions, resulting in minimal pixel counts for minor surface defects and a substantial increase in reflectivity. This physical change causes the defect texture to be compressed into a weak signal that is masked by the bright edges. Conventional frameworks struggle to distinguish feature distortions caused by geometric projections, easily misjudging weak anomalies as normal edges, leading to insufficient candidate responses. Furthermore, the two-dimensional coordinates are pulled away from the true center by the bright edges, resulting in positional shifts when projected back into three-dimensional space. These limitations restrict the defect detection rate and the accuracy of three-dimensional position binding. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method and system for detecting defects in three-dimensional models based on image recognition.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting defects in a three-dimensional model based on image recognition, comprising the following steps: S1: Parse the 3D model file of the metal structure to extract the internal triangular facets and vertex set, fit the spatial coordinates to generate the principal curvature of the vertex, merge the vertices whose absolute value of the principal curvature is greater than the preset curvature threshold to generate a high curvature influence area, and aggregate the dihedral angles between adjacent triangular facets and the extreme points of the internal principal curvature within the high curvature influence area to generate a priori feature lines. S2: Obtain two-dimensional images of the metal structural component and the camera extrinsic matrix from each acquisition viewpoint, construct a normalized line-of-sight vector corresponding to the spatial coordinates, calculate the dot product of the normalized line-of-sight vector and the average normal vector of the vertex to generate the projection compression ratio, select high curvature influence areas with an average projection compression ratio greater than or equal to the compression threshold as target areas, extract vertices inside the target areas to generate a set of high curvature vertices to be unfolded, and extract the local pixel range corresponding to the target areas in the two-dimensional image to generate a two-dimensional image to be unfolded. S3: Fit the set of high curvature vertices to be unfolded to generate a local parameterized surface, generate a three-dimensional mesh matrix and a two-dimensional mesh matrix through discretization sampling, and generate an unfolded image mapping table by associating the two-dimensional pixel coordinates of the two-dimensional mesh matrix with the three-dimensional spatial coordinate index of the three-dimensional mesh matrix. S4: Based on the unfolded image mapping table and the prior feature lines, expand to both sides to generate a prior edge mask, call the multi-branch feature decoupling neural network to extract features from the prior edge mask, calculate differential features and generate an anomaly probability map based on the differential features, segment the anomaly probability map and filter target connected regions whose area exceeds the preset area resolution scale threshold, calculate the centroid coordinates of the target connected regions to generate candidate defect two-dimensional coordinates. S5: Construct a spatial search line segment based on the three-dimensional initial coordinates of the two-dimensional coordinate mapping of the candidate defect, and generate a local texture normal vector by back projection. Calculate the dot product value between the normal vector of the facet at the intersection of the spatial search line segment and the internal triangular facet and the local texture normal vector. Extract the geometric intersection point with the dot product value greater than a preset positive threshold and the smallest spatial Euclidean distance as the effective intersection point, and generate the three-dimensional defect binding position based on the effective intersection point.
[0006] The present invention is improved in that the specific steps of S1 are as follows: S111: Obtain the 3D model file of the metal structural component, parse the 3D model file to extract the internal triangular facets and vertex set, traverse the vertex set to obtain the adjacent facets of each vertex and extract the area and normal vector of the adjacent facets, calculate the weighted average value based on the area and the normal vector of the facets, and construct the vertex average normal vector. S112: Call the average normal vector of the vertex, traverse the vertex set, take each vertex as the target vertex and establish a local coordinate system on its tangent plane, transform the spatial coordinates of the adjacent vertices of the target vertex to the local coordinate system, fit the spatial coordinates inside the local coordinate system based on the least squares method to extract the quadratic surface equation and solve the Weingarten mapping matrix, calculate the eigenvalues corresponding to the Weingarten mapping matrix, and generate the principal curvature of the vertex. S113: Call the preset curvature threshold and the principal curvature of the vertex, compare the absolute values of the curvature threshold and the principal curvature of the vertex, filter out high curvature vertices whose absolute values are greater than the curvature threshold, determine the spatial connectivity of high curvature vertices based on the region growing algorithm, and perform a merging operation on the connected vertices to generate a high curvature influence region; S114: Obtain the normal vectors of each adjacent triangular facet, calculate the angle data between the normal vectors to extract the dihedral angle, call the preset angle threshold, filter the common intersection lines corresponding to the dihedral angle values being greater than the angle threshold, extract the connection lines of the principal curvature extreme points inside the high curvature influence area, aggregate the common intersection line data and the extreme point connection line data items, and generate the prior feature line.
[0007] The present invention is improved in that the specific steps of S2 are as follows: S211: Obtain two-dimensional images and camera extrinsic matrix from each acquisition viewpoint, deduce the absolute spatial coordinates of the camera optical center in the three-dimensional world coordinate system based on the camera extrinsic matrix, extract the spatial coordinates of each vertex inside the high curvature influence area, construct the line-of-sight vector between the absolute spatial coordinates and the spatial coordinates of each vertex, and perform normalization processing on the line-of-sight vector to generate a normalized line-of-sight vector. S212: Call the normalized line-of-sight vector and the vertex average normal vector of the corresponding vertex, calculate the dot product between the normalized line-of-sight vector and the vertex average normal vector, extract the dot product as the cosine value of the angle between the line of sight and the local surface obliqueness, perform ratio calculation based on constant 1 and the cosine value of the angle, extract the ratio calculation results corresponding to each vertex, and generate the projection compression ratio. S213: Extract the projection compression ratio of all vertices within the high curvature influence region, perform an average calculation operation on the projection compression ratio of all vertices, extract the calculation result to generate an average projection compression ratio, call the compression threshold, compare the compression threshold with the average projection compression ratio, select the high curvature influence region corresponding to the average projection compression ratio being greater than or equal to the compression threshold as the target region, extract the vertices inside the target region to generate a set of high curvature vertices to be unfolded, and extract the local pixel range corresponding to the target region in the two-dimensional image according to the camera perspective relationship to generate a two-dimensional image to be unfolded.
[0008] The present invention is improved in that the process of selecting the high curvature influence region corresponding to the average projection compression ratio being greater than or equal to the compression threshold as the target region is specifically as follows: Obtain the mean calculation instruction, read the projection compression ratio of all vertices in the high curvature influence area, and perform cumulative calculation based on the projection compression ratio of each vertex to obtain the cumulative sum value; The total number of vertices contained in the high curvature influence region is read, and the average projection compression ratio is obtained by performing a division operation based on the accumulated sum value and the total number of vertices. Obtain the maximum optical distortion coefficient corresponding to the camera intrinsic parameter matrix and the minimum resolvable pixel size corresponding to the two-dimensional image; The initial compression tolerance limit is set by performing a function mapping calculation based on the maximum optical distortion coefficient and the minimum resolvable pixel size; Obtain the local maximum principal curvature value corresponding to the high curvature influence area, and adjust the initial compression tolerance limit nonlinearly based on the local maximum principal curvature value to set the compression threshold. The compression threshold and the average projection compression ratio are compared according to a preset numerical comparison rule. Based on the comparison results where the average projection compression ratio is greater than or equal to the compression threshold, the corresponding high curvature influence area is extracted and marked as the target area.
[0009] The present invention is improved in that the specific steps of S3 are as follows: S311: Obtain the camera intrinsic parameter matrix, calculate the spatial coordinates of the vertices inside the set of high curvature vertices to be unfolded according to the spatial coordinates of the vertices inside the set of high curvature vertices to be unfolded, calculate the spatial coordinates along the principal curvature direction and the minimum curvature direction using the B-spline surface fitting algorithm, generate a local parameterized surface, discretize and sample the local parameterized surface according to the equal arc length distribution condition, and generate a three-dimensional mesh lattice and the corresponding two-dimensional mesh lattice. S312: Call the camera intrinsic and extrinsic matrix of the corresponding viewpoint camera, perform perspective projection transformation on the three-dimensional grid dot matrix, obtain the set of pixel coordinates inside the two-dimensional image to be unfolded, extract the grayscale information of the adjacent integer pixels around the set of pixel coordinates inside the two-dimensional image to be unfolded, calculate the sub-pixel grayscale interpolation data using the bicubic interpolation algorithm based on the grayscale information as the sub-pixel grayscale interpolation result, assign the sub-pixel grayscale interpolation result to the pixel position corresponding to the two-dimensional grid dot matrix, and generate the unfolded image; S313: Extract the row and column indices of the corresponding two-dimensional grid points inside the unfolded image, extract the corresponding three-dimensional coordinates inside the local parameterized surface, establish the mapping relationship between the row and column indices and the three-dimensional coordinates, generate mapping data items, and summarize the mapping data items to generate an unfolded image mapping table.
[0010] The present invention is improved in that the specific steps of S4 are as follows: S411: Obtain the point spread function and diffuse reflection attenuation model, call the unfolded image mapping table to reversely retrieve the corresponding pixel trajectory of the prior feature line inside the unfolded image, calculate the Gaussian dilation weight value corresponding to the expansion of the pixel trajectory to both sides according to the point spread function and the diffuse reflection attenuation model, and generate a prior edge mask. S412: For the unfolded image, a multi-branch feature decoupling neural network is invoked to extract the corresponding basic edge gradient and multi-scale texture features. The basic edge gradient and multi-scale texture features are integrated to establish a basic feature map. Based on the autoencoder branch inside the multi-branch feature decoupling neural network and combined with the prior edge mask indication corresponding to the spatial attention mechanism, the normal edge highlight response corresponding to the basic feature map is calculated to establish a background feature map. For the basic feature map and the background feature map, pixel-by-pixel gray-level difference value calculation is performed to extract the difference features and generate an anomaly probability map. S413: Invoke the adaptive Otsu method that combines local neighborhood mean to segment the anomaly probability map and extract connected regions. Compare the area values of connected regions according to a preset area resolution scale threshold, and select connected regions whose area values exceed the preset area resolution scale threshold as target connected regions. Calculate the centroid coordinates of the target connected regions in the unfolded image and generate two-dimensional coordinates of candidate defects.
[0011] The present invention is improved in that the process of generating the two-dimensional coordinates of the candidate defect is specifically as follows: Calculate the mean probability of local windows of pixels in the anomaly probability map and adjust the global threshold of the Otsu method to set an adaptive segmentation threshold; Based on the adaptive segmentation threshold, the anomaly probability map is binarized to generate a binary image containing connected regions. The area of the connected region is generated by counting the number of pixels in the connected region. The preset area resolution scale threshold is set according to the preset tolerance defect physical area and spatial mapping scale parameters; Extract connected regions whose area values are greater than the preset area resolution scale threshold as target connected regions; The zeroth-order spatial moment and the first-order spatial moment of the target connected region within the unfolded image are calculated, and a division operation is performed to generate two-dimensional coordinates of the candidate defect.
[0012] The present invention is improved in that the specific steps of S5 are as follows: S511: Obtain the preset tolerance distance range, call the unfolded image mapping table, use the bilinear interpolation algorithm to retrieve the candidate defect two-dimensional coordinates and map them to the corresponding three-dimensional initial coordinates and parameterized surface normal vector inside the local parameterized surface, take the three-dimensional initial coordinates as the starting node, establish reverse search line segments along the positive and negative directions corresponding to the parameterized surface normal vector within the tolerance distance range, and generate spatial search line segments. S512: Extract the corresponding principal gray-level gradient direction inside the unfolded image for the two-dimensional coordinates of the candidate defect, call the unfolded image mapping table to perform back projection parameter transformation processing operation, project the principal gray-level gradient direction into the three-dimensional world coordinate system, and generate a local texture normal vector. S513: Obtain a preset positive threshold, call the spatial search line segment and the internal triangular facet, calculate the facet normal vector corresponding to the geometric intersection point generated by the two, perform the corresponding dot product calculation operation between the facet normal vector and the local texture normal vector, compare the positive threshold with the dot product value, filter the valid intersection points corresponding to the dot product value being greater than the positive threshold, perform spatial feature measurement operation on the valid intersection points and the three-dimensional initial coordinates, extract the spatial Euclidean distance between the valid intersection points and the three-dimensional initial coordinates, retrieve the valid intersection point corresponding to the smallest spatial Euclidean distance value, and generate the three-dimensional defect binding position.
[0013] The present invention is improved in that the process of generating the three-dimensional defect binding position is specifically as follows: Extract the allowable deflection angle of the surface normal and calculate the corresponding cosine value of the allowable deflection angle of the surface normal; extract the patch fitting error parameters corresponding to the three-dimensional model analysis stage; combine the cosine value and the patch fitting error parameters to perform subtraction calculation to set the preset positive threshold; Extract the internal triangular facets corresponding to the bounding box coverage area of the spatial search line segment; calculate the spatial coordinates of the intersection between the spatial search line segment and the internal triangular facets to generate a geometric intersection point; extract the unit perpendicular vector corresponding to the internal triangular facet at the geometric intersection point to generate a facet normal vector; extract the facet normal vector and the local texture normal vector in the three-dimensional coordinate system corresponding to the three coordinate axis dimensional components; For the corresponding dimensional component values, perform product calculation and summation to generate a dot product value; extract the geometric intersection points corresponding to the dot product values that are greater than the preset positive threshold to generate valid intersection points; calculate the coordinate differences between the valid intersection points and the three-dimensional initial coordinates in the three coordinate axes; perform square calculation and summation on the coordinate differences in the three coordinate axes to generate a distance sum of squares value; perform square root calculation on the distance sum of squares value to generate spatial Euclidean distance; compare the values of the spatial Euclidean distances corresponding to all valid intersection points; extract the valid intersection points corresponding to the spatial Euclidean distances in the minimum value state to generate three-dimensional defect binding positions.
[0014] A 3D model defect detection system based on image recognition, wherein the image recognition-based 3D model defect detection system is used to implement the above-mentioned image recognition-based 3D model defect detection method, and the system includes: The model parsing module parses the 3D model file of the metal structure to extract the internal triangular facets and vertex set, fits the spatial coordinates to generate the principal curvature of the vertex, merges the vertices whose absolute value of the principal curvature is greater than the preset curvature threshold to generate a high curvature influence area, and aggregates the dihedral angles between adjacent triangular facets and the extreme points of the internal principal curvature within the high curvature influence area to generate a priori feature lines. The distortion assessment module acquires two-dimensional images of the metal structure and camera extrinsic matrix from various acquisition angles, constructs a normalized line-of-sight vector corresponding to the spatial coordinates, calculates the dot product of the normalized line-of-sight vector and the average normal vector of the vertex to generate the projection compression ratio, selects high curvature influence areas with an average projection compression ratio greater than or equal to the compression threshold as target areas, extracts vertices inside the target areas to generate a set of high curvature vertices to be unfolded, and extracts the local pixel range corresponding to the target areas in the two-dimensional images to generate a two-dimensional image to be unfolded. The image reconstruction module fits the set of high curvature vertices to be unfolded to generate a local parameterized surface, generates a three-dimensional mesh matrix and a two-dimensional mesh matrix through discretization sampling, and generates an unfolded image mapping table by associating the two-dimensional pixel coordinates of the two-dimensional mesh matrix with the three-dimensional spatial coordinate index of the three-dimensional mesh matrix. The defect identification module expands to both sides based on the unfolded image mapping table and prior feature lines to generate a prior edge mask, calls a multi-branch feature decoupling neural network to extract features from the prior edge mask, calculates differential features and generates an anomaly probability map based on the differential features, segments the anomaly probability map and filters target connected regions whose area exceeds a preset area resolution scale threshold, and calculates the centroid coordinates of the target connected regions to generate two-dimensional coordinates of candidate defects. The spatial positioning module constructs a spatial search line segment based on the three-dimensional initial coordinates mapped from the two-dimensional coordinates of the candidate defect, and generates a local texture normal vector by back projection. It calculates the dot product between the normal vector of the facet at the intersection of the spatial search line segment and the internal triangular facet and the local texture normal vector. It extracts the geometric intersection point with the dot product value greater than a preset positive threshold and the smallest spatial Euclidean distance as the effective intersection point, and generates the three-dimensional defect binding position based on the effective intersection point.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a priori feature lines are generated by analyzing a 3D model to construct the geometric foundation. The projection compression ratio is generated by calculating the dot product of the line of sight and the normal vector to screen the target area and avoid the feature distortion problem caused by projection. The high curvature area is fitted into a parameterized surface and mapped with 2D pixels to solve the limitation that the defect texture is compressed into a weak signal. The prior mask is generated by combining the mapping table and feature lines, and a neural network is called to distinguish between defects and bright boundaries and generate an anomaly probability map to overcome the defect of reflection masking weak response. The spatial search line segment is constructed using candidate coordinates and the effective intersection point is calculated by back projection to generate the 3D defect binding position to correct the positional offset deviation caused by the 2D coordinate deviating from the true center, thereby improving the defect detection rate and binding accuracy. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of S1 of the present invention; Figure 3 This is a flowchart of S2 of the present invention; Figure 4 This is a flowchart of S3 of the present invention; Figure 5 This is a flowchart of S4 of the present invention; Figure 6 This is a flowchart of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0018] This embodiment provides a method and system for detecting defects in 3D models based on image recognition. During the detection process of 3D modeling and multi-view image acquisition of metal structural components, surface transitions, hole and slot boundaries, bending transitions, and concave connection areas in the 3D model will cause edge overlap, projection compression, and texture occlusion in the 2D image. This embodiment continuously processes the 3D model data, 2D image data, and camera extrinsic data of the same metal structural component. It first determines the high curvature influence area and prior feature lines from the 3D model, then selects the target area based on the viewpoint compression state, and performs unfolding reconstruction, anomaly identification, and 3D binding localization on the target area, forming a method embodiment and a system embodiment associated with this method. The method embodiment serves as a detailed implementation object, while the system embodiment illustrates the correspondence between its functional structure and the aforementioned method flow. Method Implementation Examples
[0019] Please see Figure 1 and Figure 2 S1: The process involves parsing the 3D model file of a metal structural component to extract internal triangular facets and vertex sets. Spatial coordinates are fitted to generate vertex principal curvature. Vertices whose principal curvature meets preset curvature criteria are merged to generate a high curvature influence zone. The dihedral angles between adjacent triangular facets within the high curvature influence zone and the lines connecting the extreme points of the internal principal curvature are aggregated to generate prior feature lines. Internal triangular facets refer to the triangular facet records in the 3D model file used to represent the detectable geometric boundaries of the metal structural component's surface and interior. Each record contains facet vertex indices, facet spatial positions, and facet orientation information. Vertex sets are the set of spatial points obtained from the parsed 3D model file and indexed by the internal triangular facets. Principal curvature refers to the surface bending state field obtained based on the neighborhood geometry of the target vertex; this field distinguishes between smooth regions, turning regions, and regions with abrupt local morphological changes. The high curvature influence zone is a local region composed of vertices and related facets that meet the curvature criteria and are spatially contiguous. Prior feature lines refer to linear reference data formed by geometric turning boundaries and curvature change lines within the high curvature influence area, which are subsequently used to constrain the generation of edge masks in the unfolded image.
[0020] Please see Figure 1 and Figure 2S111: Obtain the 3D model file of the metal structural component, parse the 3D model file to extract the internal triangular facets and vertex set, traverse the vertex set to obtain the adjacent facets of each vertex and extract the area and normal vector of the adjacent facets, calculate the weighted average value based on the area and facet normal vector, and construct the vertex average normal vector. The facet normal vector refers to the unit orientation field of the plane where the facet is located. Its orientation is determined according to the unified winding order of the facet vertices in the 3D model file. When there are facets with inconsistent winding orders in the same model, they are first uniformized according to the external direction of the entity on both sides of the shared edge of the adjacent facets, and then proceed to the subsequent averaging process. The area field comes from the spatial unfolding result of the facet geometric boundary and is used to represent the degree of influence of adjacent facets on the local orientation of the target vertex. When the area participates in the weighting, the facet with a larger coverage area in the neighborhood has a stronger influence on the direction of the vertex average normal vector, and the facet with a smaller coverage area has a weaker influence on the direction. The constructed vertex average normal vector is bound to the corresponding vertex index and used as input for subsequent line-of-sight vector dot product judgment, local coordinate system establishment, and intersection validity screening.
[0021] Please see Figure 1 and Figure 2 S112: The process involves calling the average normal vector of the vertex, traversing the vertex set, using each vertex as the target vertex, establishing a local coordinate system on its tangent plane, transforming the spatial coordinates of adjacent vertices of the target vertex to the local coordinate system, fitting the spatial coordinates within the local coordinate system using the least squares method to extract the quadratic surface equation, solving the Weingarten mapping matrix, calculating the corresponding eigenvalues of the Weingarten mapping matrix, and generating the principal curvatures of the vertex. Here, the local coordinate system refers to a temporary geometric coordinate description using the target vertex as the spatial reference, the average normal vector of the vertex as the normal direction, and the stable tangent direction within the neighborhood as the tangent plane direction. The Weingarten mapping matrix is a set of curvature description fields formed by the local surface normal transformation relationship; its eigenvalues are used to generate the curvature state of the target vertex in the two principal directions. During least squares fitting, the input is the transformed local spatial coordinates of the target vertex and its adjacent vertices; the processing result is the local surface parameters that can describe the trend of surface change in the neighborhood. If the target vertex has duplicate vertices, missing patches, or the neighboring points cannot form a stable tangent plane, then the vertex is marked as having undetermined curvature and will only be used as an adjacent supplementary point when merging high curvature influence areas, without triggering the generation of high curvature regions independently.
[0022] Please see Figure 1 and Figure 2S113: The preset curvature threshold and the principal curvature of the vertices are compared. The relationship between the curvature threshold and the principal curvature of the vertices is analyzed. Vertices whose curvature meets the curvature threshold criteria are selected as high-curvature vertices. Based on the region growing algorithm, the spatial connectivity of high-curvature vertices is determined, and a merging operation is performed on connected vertices to generate a high-curvature influence region. The preset curvature threshold is derived from the process calibration record, historical qualified model statistics record, or inspection task configuration record before the inspection of metal structural parts. The field means the curvature boundary used to distinguish between smooth surfaces and the curvature boundary of the area to be expanded for identification. The region growing algorithm starts with high-curvature vertices and expands connected vertices according to the shared edge relationship of triangular facets and the adjacency relationship of vertices. When adjacent high-curvature vertices can be continuously reached through internal triangular facets, they are included in the same high-curvature influence region. When there is no shared edge continuity between two groups of high-curvature vertices, or when they are only connected by points with undetermined curvature and cannot form a stable facet link, they are retained as different high-curvature influence regions. Each high curvature influence region outputs a region identifier, vertex index, patch index, local maximum principal curvature state, and boundary vertex index for S2 to perform viewpoint compression filtering.
[0023] Please see Figure 1 and Figure 2 S114: Obtain the normal vectors of each adjacent triangular facet, calculate the angle data between the normal vectors to extract dihedral angles, call the preset angle threshold, filter the common intersection lines corresponding to the dihedral angle states that meet the angle threshold judgment conditions, extract the connection lines of the principal curvature extreme points within the high curvature influence zone, aggregate the common intersection line data and the extreme point connection line data items, and generate the prior feature lines. The preset angle threshold comes from the process boundary calibration records or the turn boundary statistical records of qualified structural parts in the model analysis stage, and the field meaning is the angle judgment condition used to identify the geometric polyline boundary. The connection lines of the principal curvature extreme points are obtained by connecting the vertices in the same high curvature influence zone where the curvature state is at a local extreme change position according to the facet connection path. When connecting, priority is given to generating linear trajectories along the shared edge path and the continuous curvature change path. The data items of the prior feature lines include the identifier of the high curvature influence zone, the vertex index passed by the line segment, the associated facet index, the order relationship of the line segment in three-dimensional space, and the mapping state during subsequent unfolding. By simultaneously extracting the common intersection line of dihedral angles and the line connecting the extreme points of principal curvature during the model analysis stage, subsequent two-dimensional image recognition can obtain prior constraints consistent with the three-dimensional geometric boundary, thereby reducing the situation of misidentifying normal structural boundaries as defect candidate regions.
[0024] Please see Figure 1 and Figure 3S2: Acquire 2D images of the metal structural component and the camera extrinsic matrix from various acquisition perspectives. Construct a normalized line-of-sight vector corresponding to the spatial coordinates. Calculate the dot product of the normalized line-of-sight vector and the vertex average normal vector to generate the projection compression ratio. Select high-curvature influence areas whose average projection compression ratio meets the compression threshold as target regions. Extract vertices within the target regions to generate a set of high-curvature vertices to be unfolded. Extract the corresponding local pixel range of the target region in the 2D image to generate the 2D image to be unfolded. Here, the camera extrinsic matrix refers to the data field describing the position and orientation relationship between the camera coordinate system and the 3D world coordinate system at the acquisition perspective. The normalized line-of-sight vector is the direction field pointing from the camera optical center to the target vertex. After uniformizing the direction length, it is used for comparison with the vertex average normal vector. The projection compression ratio is a dimensionless state field representing the degree of oblique compression of the target region at the current perspective. Its input comes from the directional relationship between the normalized line-of-sight vector and the vertex average normal vector. The 2D image to be unfolded refers to the local image data cropped from the original 2D image according to the projection range of the target region.
[0025] Please see Figure 1 and Figure 3 S211: Acquire 2D images and camera extrinsic parameter matrices from each acquisition viewpoint. Based on the camera extrinsic parameter matrix, deduce the absolute spatial coordinates of the camera's optical center in the 3D world coordinate system. Extract the spatial coordinates of each vertex within the high curvature influence area. Construct a line-of-sight vector between the absolute spatial coordinates and the spatial coordinates of each vertex, and perform normalization processing on the line-of-sight vector to generate a normalized line-of-sight vector. The absolute spatial coordinates of the camera's optical center are jointly determined by the position and attitude fields in the extrinsic parameters, serving as the starting point of the line-of-sight for each acquisition viewpoint. The normalization process unifies the line-of-sight direction at different spatial distances into a field that only indicates direction, avoiding the influence of distance differences on projection compression judgment. If an acquisition viewpoint lacks an extrinsic parameter field, or if the extrinsic parameter field cannot correspond to the 2D image at the viewpoint, then that viewpoint will not participate in target area filtering. If the same high curvature influence area can be projected from multiple effective viewpoints, then a set of normalized line-of-sight vectors under the viewpoint identifier is generated for each viewpoint.
[0026] Please see Figure 1 and Figure 3S212: The normalized view vector and the average vertex normal vector of the corresponding vertex are called. The dot product between the normalized view vector and the average vertex normal vector is calculated. This dot product is extracted as the cosine of the angle between the view vector and the local surface obliqueness. A ratio calculation is performed based on a constant and the cosine of the angle, and the ratio calculation results for each vertex are extracted to generate the projection compression ratio. The dot product value serves as the orientation consistency field. The closer the orientation is to the surface's normal viewing state, the weaker the projection compression state; the further the orientation deviates from the surface's normal viewing state, the stronger the projection compression state. The projection compression ratio field is recorded according to the vertex index and viewpoint identifier, and is subsequently used for the average state determination within the same high curvature influence area. If the orientation relationship causes the compression state to be unstable, the corresponding vertex is marked as having an abnormal view state. This vertex is not used as a separate criterion for determining the target region, but is still retained in the region vertex set for use in spatial connectivity relationships.
[0027] Please see Figure 1 and Figure 3 S213: Extract the projection compression ratio of all vertices within the high curvature influence region, perform mean calculation on the projection compression ratio of all vertices, extract the calculation results to generate the average projection compression ratio, call the compression threshold, compare the compression threshold with the average projection compression ratio, select the high curvature influence region corresponding to the average projection compression ratio reaching the compression threshold judgment condition as the target region, extract the vertices inside the target region to generate a set of high curvature vertices to be unfolded, and extract the local pixel range of the target region in the 2D image according to the camera perspective relationship to generate the 2D image to be unfolded. The average projection compression ratio refers to the summary field of the projection compression state of each effective vertex in the same high curvature influence region under the same acquisition viewpoint. The compression threshold is jointly determined by the optical distortion field in the camera intrinsic parameter matrix, the minimum resolvable pixel field of the 2D image, and the local maximum principal curvature field of the high curvature influence region. The closer the maximum optical distortion coefficient is to the boundary state in the lens calibration record, the tighter the initial compression tolerance limit; the weaker the texture resolution corresponding to the minimum resolvable pixel size, the tighter the initial compression tolerance limit; the closer the local maximum principal curvature is to the abrupt structure state, the stronger the attenuation of the initial compression tolerance limit. After the compression threshold is set, it is bound and saved with viewpoint and region identifiers, and used to determine whether the corresponding high curvature influence area needs to be expanded under that viewpoint. The local pixel range is determined according to the coverage boundary of all projection points of the target area in the two-dimensional image, and the boundary expansion state is preserved to accommodate the neighboring pixels required for edge mask expansion and abnormal region segmentation.
[0028] In one embodiment, when screening target regions, the projection compression ratio field of each vertex within the high curvature influence area is first read, and vertices with abnormal viewing conditions are excluded. Then, the effective vertex count status of the high curvature influence area is read, and the compression ratio status corresponding to the effective vertices is summarized into an average projection compression ratio. Subsequently, the maximum optical distortion field corresponding to the camera intrinsic parameter matrix and the minimum resolvable pixel field corresponding to the 2D image are read, and an initial compression tolerance limit is generated according to the mapping rules in the detection task configuration. Then, the local maximum principal curvature field corresponding to the high curvature influence area is read, and curvature attenuation adjustment is performed on the initial compression tolerance limit to obtain the compression threshold for the current region and the current viewing angle. If the average projection compression ratio reaches the compression threshold judgment condition, the high curvature influence area is marked as a target region; if not, its region resolution record is retained but it does not enter the S3 unfolding process. By converting optical distortion, pixel resolution, and local curvature status into a compression threshold, target region screening can simultaneously reflect imaging conditions and geometric complexity, thereby concentrating subsequent unfolding processing on regions prone to edge compression and defect confusion.
[0029] Please see Figure 1 and Figure 4 S3: Fit the set of high-curvature vertices to be unfolded to generate a locally parametric surface. Discretize and sample to generate a 3D and 2D mesh matrix. Associate the 2D pixel coordinates of the 2D mesh matrix with the 3D spatial coordinate indices of the 3D mesh matrix to generate an unfolded image mapping table. The locally parametric surface refers to a local continuous surface description formed by fitting the set of high-curvature vertices to be unfolded, used to establish a correspondence between the 3D surface and the unfolded 2D plane. The 3D mesh matrix refers to the spatial point sequence sampled along the locally parametric surface. The 2D mesh matrix refers to the unfolded plane point sequence corresponding to the 3D mesh matrix in row and column indices. The unfolded image mapping table is a data table recording the correspondence between the unfolded image row and column indices, 2D pixel coordinates, 3D spatial coordinates, and the identifier of the target region. It is subsequently used for prior feature line projection, candidate defect coordinate backmapping, and 3D binding localization.
[0030] Please see Figure 1 and Figure 4S311: Obtain the camera intrinsic parameter matrix. Based on the spatial coordinates of the vertices within the set of high-curvature vertices to be unfolded, calculate the spatial coordinates along the principal curvature direction and the minimum curvature direction using a B-spline surface fitting algorithm to generate a local parametric surface. Discretize the local parametric surface according to the equal arc length distribution condition to generate a 3D mesh lattice and the corresponding 2D mesh lattice. The B-spline surface fitting algorithm receives the set of high-curvature vertices to be unfolded, the principal curvature direction of the vertices, and the relationship between adjacent facets as input, and fits the transition region of the surface into a continuous parametric surface. The principal curvature direction is used to determine the sampling direction most sensitive to the curvature change of the surface, and the minimum curvature direction is used to determine the unfolding direction that matches the principal direction. The equal arc length distribution condition organizes the sampling points according to the consistent path length on the local surface, so that the surface intervals corresponding to adjacent mesh points in the unfolded image remain continuous. If there are local holes or boundary gaps in the set of high-curvature vertices to be unfolded, the boundary vertex extension state of adjacent facets is used to supplement the sampling boundary, and the supplemented points are marked as boundary auxiliary states to avoid them being output separately as defect candidate points.
[0031] Please see Figure 1 and Figure 4 S312: The corresponding camera intrinsic and extrinsic parameter matrices are invoked. A perspective projection transformation is performed on the 3D mesh matrix to obtain the set of pixel coordinates within the corresponding 2D image to be unfolded. Gray-level information is extracted from the surrounding integer pixels of the corresponding pixel coordinate set within the 2D image to be unfolded. Based on this gray-level information, a bicubic interpolation algorithm is used to calculate sub-pixel gray-level interpolation data as the sub-pixel gray-level interpolation result. This sub-pixel gray-level interpolation result is assigned to the corresponding pixel position in the 2D mesh matrix to generate the unfolded image. The perspective projection transformation takes the 3D mesh matrix, camera intrinsic and extrinsic parameter matrices as input, mapping each 3D mesh point to a pixel position in the 2D image to be unfolded. The bicubic interpolation algorithm reads the gray-level field within the neighborhood of the target pixel position and determines the sub-pixel gray-level interpolation result based on the continuity of gray-level changes in the neighborhood. If the projection position falls outside the boundary of the two-dimensional image to be unfolded, the corresponding two-dimensional grid point is marked as a missing state; if the grayscale information around the projection position is incomplete, it is supplemented by the continuous grayscale state of adjacent valid points in the same grid row and column, and the supplementation mark is retained. In subsequent abnormal segmentation, the priority of this position as an independent candidate defect is reduced.
[0032] Please see Figure 1 and Figure 4S313: Extract the row and column indices of the corresponding 2D grid points inside the unfolded image, extract the corresponding 3D coordinates inside the local parametric surface, establish the mapping relationship between the row and column indices and the 3D coordinates, generate mapping data items, and summarize the mapping data items to generate an unfolded image mapping table. The mapping data items include the target region identifier, acquisition viewpoint identifier, unfolded image row and column indices, original 2D pixel coordinates, 3D spatial coordinates, parametric surface normal vector, sampling valid state, and boundary auxiliary state. After generating the unfolded image mapping table, prior feature lines can be traced back to their pixel trajectories in the unfolded image according to the 3D coordinate indices, and candidate defect 2D coordinates can be traced back to their initial 3D coordinates according to the row and column indices. By establishing the unfolded image mapping table, a traceable index connection is formed between the 2D anomaly region and the 3D geometric position, thus providing a common coordinate basis for subsequent prior mask generation, anomaly probability map segmentation, and 3D defect binding.
[0033] Please see Figure 1 and Figure 5 S4: Based on the unfolded image mapping table and prior feature lines, a priori edge masks are generated by expanding to both sides. A multi-branch feature decoupling neural network is called to extract features from the priori edge masks, calculate differential features, and generate an anomaly probability map based on the differential features. The anomaly probability map is segmented, and target connected regions with areas exceeding a preset area resolution scale threshold are selected. The centroid coordinates of the target connected regions are calculated to generate candidate defect two-dimensional coordinates. Here, the priori edge mask refers to the edge constraint region field formed after mapping the three-dimensional prior feature lines onto the unfolded image, and its range covers the pixel trajectory of the prior feature lines and the neighboring areas affected by imaging diffusion on both sides. The multi-branch feature decoupling neural network is an image feature processing network composed of a basic edge branch, a texture branch, and an autoencoder background branch. The anomaly probability map is an image field that records the abnormal response state of each pixel position relative to the normal edge background in the unfolded image. The candidate defect two-dimensional coordinates refer to the centroid position of the target connected region in the unfolded image, used for three-dimensional binding in S5.
[0034] Please see Figure 1 and Figure 5S411: Obtain the point spread function and diffuse attenuation model, call the unfolded image mapping table to retrieve the corresponding pixel trajectory of the prior feature line inside the unfolded image, calculate the Gaussian dilation weight values corresponding to the expansion of the pixel trajectory to both sides based on the point spread function and diffuse attenuation model, and generate a prior edge mask. The point spread function refers to the pixel spread description field from the camera imaging calibration record, used to represent the diffusion state of edge imaging in neighboring pixels. The diffuse attenuation model refers to the brightness attenuation description field formed by the surface reflection state of the metal structure and the viewing angle relationship, used to adjust the mask coverage intensity of pixels on both sides of the prior edge. During the reverse retrieval, first look up the row and column indices in the unfolded image mapping table based on the 3D vertex index in the prior feature line, and then connect adjacent row and column indices to form the pixel trajectory. When expanding along both sides of the pixel trajectory, pixels closer to the trajectory are assigned a stronger prior edge state, and pixels farther from the trajectory and less affected by diffuse reflection are assigned a weaker prior edge state. The generated prior edge mask has the same row and column index structure as the unfolded image and serves as the input for the subsequent spatial attention mechanism.
[0035] Please see Figure 1 and Figure 5 S412: For the unfolded image, a multi-branch feature decoupling neural network is invoked to extract the corresponding basic edge gradients and multi-scale texture features. The basic edge gradients and multi-scale texture features are integrated to establish a basic feature map. Based on the autoencoder branch within the multi-branch feature decoupling neural network, combined with the prior edge mask indicating the corresponding spatial attention mechanism, the background feature map is established by calculating the normal edge highlight response corresponding to the basic feature map. Pixel-by-pixel gray-level difference calculations are performed on the basic feature map and the background feature map to extract differential features and generate an anomaly probability map. The basic edge branch receives the gray-level change field of the unfolded image and outputs the edge gradient state; the texture branch reads the gray-level texture continuity, texture direction consistency, and local texture abrupt change states at different neighborhood scales and outputs multi-scale texture features; the autoencoder background branch receives the basic feature map and the prior edge mask. During the encoding stage, it compresses the common features of normal edges and normal textures, and during the decoding stage, it reconstructs a background feature map consistent with the prior edges. The spatial attention mechanism increases the response weight of normal structural edge positions in background reconstruction based on the prior edge mask and suppresses isolated texture abrupt changes inconsistent with the prior edges. During pixel-wise differencing, responses in the base feature map that cannot be interpreted by the background feature map are transformed into anomaly state fields in the anomaly probability map. During network training, training images are derived from qualified surface images of metal structural components, labeled defect images, and unfolded images corresponding to prior feature lines. Labeling rules distinguish between normal geometric edges, normal texture regions, and defect regions. Parameter adjustments are made based on the deviation direction between the network output and the labeled state until the output can stably distinguish between normal edge backgrounds and non-prior anomalous responses before being used for inference processing.
[0036] Please see Figure 1 and Figure 5 S413: The adaptive Otsu method, which combines the local neighborhood mean with the segmentation of the anomaly probability map, is invoked to extract connected regions. The area values of the connected regions are compared with a preset area resolution scale threshold. Connected regions whose area values exceed the preset area resolution scale threshold are selected as target connected regions. The centroid coordinates of the target connected region within the unfolded image are calculated to generate two-dimensional coordinates of candidate defects. The adaptive Otsu method first reads the anomaly state of each pixel in the anomaly probability map and adjusts the segmentation threshold based on the average state of the anomaly response in the pixel's neighborhood. The more continuous the local neighborhood anomaly response, the more likely the location is to be classified as an anomaly region. When the local neighborhood anomaly response is scattered and close to the normal background state, the location is not considered the core of the connected region. The connected region area field is determined by the number of connected pixels and the unfolded image mapping scale. The preset area resolution scale threshold comes from the configuration records of the tolerable defect physical area, unfolded image mapping scale, and imaging resolution capability in the detection task. The centroid coordinates of the target connected region are determined by the spatial moment state of the pixels within the connected region. The output fields include the unfolded image row and column coordinates, the target region identifier, the viewpoint identifier, and the valid state of the connected region.
[0037] In one embodiment, when generating candidate defect 2D coordinates, the mean probability of local windows of pixels in the anomaly probability map is first calculated, and the global threshold of the Otsu method is adjusted to set an adaptive segmentation threshold. Then, the anomaly probability map is binarized based on the adaptive segmentation threshold to generate a binary image containing connected regions. Subsequently, the number of pixels in the connected regions is counted to generate the area state of the connected regions, and a preset area resolution scale threshold is set according to the preset tolerance defect physical area and spatial mapping scale parameters. When the area state of the connected regions meets the area resolution scale judgment condition, it is extracted as the target connected region. Then, the zero-order spatial moment and the first-order spatial moment of the target connected region are read to determine the centroid position of the target connected region in the unfolded image, and candidate defect 2D coordinates are generated. If the connected region in the binary image coincides with the boundary of the 2D image to be unfolded and lacks a complete neighborhood, it is marked as a boundary pending confirmation state, and the 3D defect binding position is output only when a valid 3D intersection point can be obtained in S5. By combining the prior edge mask, background feature map and adaptive segmentation results, normal structural edges can be reconstructed as background, while abnormal textures inconsistent with the prior edges can form candidate defect 2D coordinates, thereby reducing the interference of complex curved surface edges on defect recognition.
[0038] Please see Figure 1 and Figure 6S5: Based on the 3D initial coordinates mapped from the 2D coordinates of the candidate defect, a spatial search line segment is constructed, and a local texture normal vector is generated by back-projection. The dot product between the normal vector of the facet at the intersection of the spatial search line segment and the internal triangular facet and the local texture normal vector is calculated. The geometric intersection point with the minimum Euclidean distance that meets the preset positive threshold is extracted as the valid intersection point, and the 3D defect binding position is generated based on the valid intersection point. Here, the 3D initial coordinates refer to the spatial position of the local parametric surface obtained by back-looking up the 2D coordinates of the candidate defect through the unfolded image mapping table. The spatial search line segment refers to a finite search line segment established with the 3D initial coordinates as the starting state and along the positive and negative directions of the parametric surface normal. The local texture normal vector refers to the texture orientation field obtained by back-projecting the principal grayscale gradient direction of the candidate defect neighborhood in the unfolded image onto the 3D world coordinate system. The 3D defect binding position refers to the spatial coordinates of the candidate defect finally bound to the internal triangular facet of the 3D model of the metal structure.
[0039] Please see Figure 1 and Figure 6 S511: Obtain the preset tolerance distance range, call the unfolded image mapping table, and use the bilinear interpolation algorithm to retrieve the corresponding 3D initial coordinates and parametric surface normal vectors mapped to the 2D coordinates of the candidate defect within the local parametric surface. Using the 3D initial coordinates as the starting node, establish reverse search segments along the positive and negative directions corresponding to the parametric surface normal vectors within the tolerance distance range to generate spatial search segments. The preset tolerance distance range comes from the detection task configuration record of 3D model analytical error, image projection error, and surface fitting error. The field means the spatial boundary that allows the 2D coordinates of the candidate defect to perform intersection point search near the surface of the 3D model. The bilinear interpolation algorithm reads the neighboring row and column indices in the unfolded image mapping table around the 2D coordinates of the candidate defect and generates the 3D initial coordinates and parametric surface normal vectors according to the positional relationship between the candidate coordinates and the neighboring grids. The spatial search segments extend simultaneously in the positive and negative normal directions to cover the deviation between the parametric surface fitting position and the original internal triangular facet.
[0040] Please see Figure 1 and Figure 6S512: Extract the corresponding principal gray-level gradient direction inside the unfolded image based on the two-dimensional coordinates of the candidate defect. Call the unfolded image mapping table to perform back projection parameter transformation processing, projecting the principal gray-level gradient direction into the three-dimensional world coordinate system to generate a local texture normal vector. The principal gray-level gradient direction refers to the direction field where the gray-level change is most concentrated in the neighborhood of the candidate defect, used to describe the local orientation of the texture abrupt change. The back projection parameter transformation processing first reads the neighboring three-dimensional coordinates of the candidate defect's two-dimensional coordinates in the unfolded image mapping table, and then, based on the correspondence between the two-dimensional mesh direction and the three-dimensional parameterized surface direction, converts the two-dimensional gray-level gradient direction into the tangential change state in the three-dimensional world coordinate system, and combines it with the parameterized surface normal direction to generate a local texture normal vector. If the gray-level change direction in the neighborhood of the candidate defect is unstable, the overall texture direction of the target connected region is used as a supplementary input; if the overall texture direction still cannot be stably formed, the candidate defect is marked as a state where the texture direction is pending confirmation, and the binding position is only output when the geometric intersection meets a more stringent positive consistency.
[0041] Please see Figure 1 and Figure 6S513: Obtain the preset positive threshold, call the spatial search line segment and internal triangular facet, calculate the facet normal vector corresponding to the geometric intersection point generated by the two, perform the dot product calculation operation between the facet normal vector and the local texture normal vector, compare the positive threshold with the dot product value, filter the valid intersection points corresponding to the positive threshold judgment condition, perform spatial feature measurement operation on the valid intersection points and the 3D initial coordinates, extract the spatial Euclidean distance between the valid intersection points and the 3D initial coordinates, search for the valid intersection point with the smallest spatial Euclidean distance value, and generate the 3D defect binding position. The preset positive threshold is jointly determined by the cosine state corresponding to the allowable deflection angle of the surface normal and the facet fitting error parameter. The allowable deflection angle of the surface normal comes from the configuration record of the consistency between the texture orientation and the facet orientation of the detection task, and the facet fitting error parameter comes from the error state record of the S1 model analysis and S3 surface fitting process. When searching for intersections between spatial search line segments and internal triangular faces, the internal triangular faces within the bounding box of the search line segment are first read. Then, it is determined whether the line segment passes through the internal region of the face boundary, and a geometric intersection point is generated. The face normal vector is generated according to the unit perpendicular vector of the internal triangular face. If the normals of adjacent faces are inconsistent, they are unified according to the external direction of the entity. For each geometric intersection point, the directional components of the face normal vector and the local texture normal vector in the 3D coordinate system are read to form a directional consistency dot product value. Geometric intersection points that meet the preset positive threshold judgment condition are taken as valid intersection point candidates. Then, the spatial Euclidean distance state between each valid intersection point and the 3D initial coordinates is read, and the valid intersection point with the distance state closest to the 3D initial coordinates is selected as the 3D defect binding position. If there is no geometric intersection point that meets the positive threshold judgment condition, the 2D coordinates and anomaly probability map region identifier of the candidate defect are retained, and the candidate defect is marked as unbound to avoid outputting a position inconsistent with the 3D model surface.
[0042] In one embodiment, when generating the 3D defect binding location, firstly, the directional consistency benchmark corresponding to the allowable deflection angle of the surface normal is extracted. Then, the patch fitting error parameters corresponding to the 3D model analysis stage are read, and a preset positive threshold is set by combining the two. Subsequently, the internal triangular patches within the bounding box of the spatial search line segment are extracted, and it is sequentially determined whether the spatial search line segment and the internal triangular patches produce geometric intersections located inside the patch boundaries. Then, the unit perpendicular vector corresponding to the internal triangular patch where the geometric intersection is located is extracted to generate the patch normal vector, and the directional consistency is compared with the local texture normal vector. When the directional consistency state of the geometric intersection reaches the preset positive threshold judgment condition, it is taken as a valid intersection. Finally, the spatial distance state between all valid intersections and the initial 3D coordinates is compared, and the valid intersection with the smallest distance state is extracted to generate the 3D defect binding location. By using the spatial search line segment to limit the candidate region, and then using the patch normal vector and the local texture normal vector for positive consistency screening, the 2D anomaly points can be bound to the position in the 3D model that is consistent with the actual texture direction and geometric surface, thereby forming a spatial coordinate result that can be used for model defect annotation.
[0043] In this embodiment, S1 forms a high curvature influence zone and prior feature lines, enabling normal geometric edges in complex surfaces to enter the two-dimensional recognition constraints in advance; S2 combines the viewpoint compression state to filter the target area, allowing areas with projection compression and edge aliasing to enter the unfolding process; S3 establishes an unfolded image mapping table to maintain an index correspondence between the two-dimensional unfolded image and the three-dimensional spatial coordinates; S4 uses prior edge masks and multi-branch feature decoupling neural networks to separate normal edge responses and abnormal responses, ensuring that the two-dimensional coordinates of candidate defects originate from abnormal areas inconsistent with prior geometric edges; S5 binds the two-dimensional coordinates of candidate defects to the effective intersection points on the internal triangular facets, allowing the defect recognition results to return to their spatial positions in the three-dimensional model of the metal structure, thus completing the full execution of the three-dimensional model defect detection method based on image recognition.
[0044] System-related Implementation Examples Please see Figure 7This embodiment also provides an image recognition-based 3D model defect detection system associated with the aforementioned method. This system implements the processing flow described in the aforementioned method embodiments and includes a model parsing module, a distortion assessment module, an image reconstruction module, a defect identification module, and a spatial localization module. The model parsing module performs the 3D model file parsing, vertex principal curvature generation, high curvature influence region merging, and prior feature line generation processes in S1, and provides the high curvature influence region, vertex average normal vector, and prior feature lines to the distortion assessment module. The distortion assessment module performs the 2D image and camera extrinsic parameter matrix reading, normalized line-of-sight vector generation, projection compression ratio generation, target region filtering, and 2D image extraction processes in S2, and outputs the set of high curvature vertices to be unfolded and the 2D image to be unfolded to the image reconstruction module. The image reconstruction module performs the local parametric surface fitting, 3D and 2D mesh point matrix generation, unfolded image generation, and unfolded image mapping table generation processes in S3, and provides the unfolded image and unfolded image mapping table to the defect identification module. The defect identification module performs the following steps in S4: prior edge mask generation, multi-branch feature decoupling, anomaly probability map generation, connected region filtering, and candidate defect 2D coordinate generation. It then outputs the candidate defect 2D coordinates and their associated target region identifier to the spatial positioning module. The spatial positioning module performs the following steps in S5: 3D initial coordinate retrieval, spatial search line segment construction, local texture normal vector generation, effective intersection point filtering, and 3D defect binding location generation. It then outputs the defect location bound to the internal triangular facets of the 3D model of the metal structural component.
[0045] In the system-associated embodiment, the functional boundaries of each module correspond one-to-one with the aforementioned method steps. The model parsing module does not perform 2D image anomaly segmentation, the distortion assessment module does not generate 3D defect binding locations, the image reconstruction module does not determine defect connectivity regions, the defect identification module does not modify the 3D model patch structure, and the spatial positioning module does not regenerate prior feature lines. The data transmitted between modules consists of fields or results already generated in the aforementioned method embodiments, including high curvature influence areas, vertex average normal vectors, prior feature lines, the 2D image to be unfolded, the unfolded image mapping table, anomaly probability maps, candidate defect 2D coordinates, and 3D defect binding locations. Through this correspondence, the system can support the continuous data flow and judgment process of the aforementioned method, thereby ensuring that the 3D model defect detection results maintain a clear source, continuous state, and traceable location across modules.
[0046] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for defect detection in a 3D model based on image recognition, characterized in that, Includes the following steps: S1: Parse the 3D model file of the metal structure to extract the internal triangular facets and vertex set, fit the spatial coordinates to generate the principal curvature of the vertex, merge the vertices whose absolute value of the principal curvature is greater than the preset curvature threshold to generate a high curvature influence area, and aggregate the dihedral angles between adjacent triangular facets and the extreme points of the internal principal curvature within the high curvature influence area to generate a priori feature lines. S2: Obtain two-dimensional images of the metal structural component and the camera extrinsic matrix from each acquisition viewpoint, construct a normalized line-of-sight vector corresponding to the spatial coordinates, calculate the dot product of the normalized line-of-sight vector and the average normal vector of the vertex to generate the projection compression ratio, select high curvature influence areas with an average projection compression ratio greater than or equal to the compression threshold as target areas, extract vertices inside the target areas to generate a set of high curvature vertices to be unfolded, and extract the local pixel range corresponding to the target areas in the two-dimensional image to generate a two-dimensional image to be unfolded. S3: Fit the set of high curvature vertices to be unfolded to generate a local parameterized surface, generate a three-dimensional mesh matrix and a two-dimensional mesh matrix through discretization sampling, and generate an unfolded image mapping table by associating the two-dimensional pixel coordinates of the two-dimensional mesh matrix with the three-dimensional spatial coordinate index of the three-dimensional mesh matrix. S4: Based on the unfolded image mapping table and the prior feature lines, expand to both sides to generate a prior edge mask, call the multi-branch feature decoupling neural network to extract features from the prior edge mask, calculate differential features and generate an anomaly probability map based on the differential features, segment the anomaly probability map and filter target connected regions whose area exceeds the preset area resolution scale threshold, calculate the centroid coordinates of the target connected regions to generate candidate defect two-dimensional coordinates. S5: Construct a spatial search line segment based on the three-dimensional initial coordinates of the two-dimensional coordinate mapping of the candidate defect, and generate a local texture normal vector by back projection. Calculate the dot product value between the normal vector of the facet at the intersection of the spatial search line segment and the internal triangular facet and the local texture normal vector. Extract the geometric intersection point with the dot product value greater than a preset positive threshold and the smallest spatial Euclidean distance as the effective intersection point, and generate the three-dimensional defect binding position based on the effective intersection point.
2. The method for detecting defects in a three-dimensional model based on image recognition according to claim 1, characterized in that, The specific steps of S1 are as follows: S111: Obtain the 3D model file of the metal structural component, parse the 3D model file to extract the internal triangular facets and vertex set, traverse the vertex set to obtain the adjacent facets of each vertex and extract the area and normal vector of the adjacent facets, calculate the weighted average value based on the area and the normal vector of the facets, and construct the vertex average normal vector. S112: Call the average normal vector of the vertex, traverse the vertex set, take each vertex as the target vertex and establish a local coordinate system on its tangent plane, transform the spatial coordinates of the adjacent vertices of the target vertex to the local coordinate system, fit the spatial coordinates inside the local coordinate system based on the least squares method to extract the quadratic surface equation and solve the Weingarten mapping matrix, calculate the eigenvalues corresponding to the Weingarten mapping matrix, and generate the principal curvature of the vertex. S113: Call the preset curvature threshold and the principal curvature of the vertex, compare the absolute values of the curvature threshold and the principal curvature of the vertex, filter out high curvature vertices whose absolute values are greater than the curvature threshold, determine the spatial connectivity of high curvature vertices based on the region growing algorithm, and perform a merging operation on the connected vertices to generate a high curvature influence region; S114: Obtain the normal vectors of each adjacent triangular facet, calculate the angle data between the normal vectors to extract the dihedral angle, call the preset angle threshold, filter the common intersection lines corresponding to the dihedral angle values being greater than the angle threshold, extract the connection lines of the principal curvature extreme points inside the high curvature influence area, aggregate the common intersection line data and the extreme point connection line data items, and generate the prior feature line.
3. The method for detecting defects in a three-dimensional model based on image recognition according to claim 1, characterized in that, The specific steps of S2 are as follows: S211: Obtain two-dimensional images and camera extrinsic matrix from each acquisition viewpoint, deduce the absolute spatial coordinates of the camera optical center in the three-dimensional world coordinate system based on the camera extrinsic matrix, extract the spatial coordinates of each vertex inside the high curvature influence area, construct the line-of-sight vector between the absolute spatial coordinates and the spatial coordinates of each vertex, and perform normalization processing on the line-of-sight vector to generate a normalized line-of-sight vector. S212: Call the normalized line-of-sight vector and the vertex average normal vector of the corresponding vertex, calculate the dot product between the normalized line-of-sight vector and the vertex average normal vector, extract the dot product as the cosine value of the angle between the line of sight and the local surface obliqueness, perform ratio calculation based on constant 1 and the cosine value of the angle, extract the ratio calculation results corresponding to each vertex, and generate the projection compression ratio. S213: Extract the projection compression ratio of all vertices within the high curvature influence region, perform an average calculation operation on the projection compression ratio of all vertices, extract the calculation result to generate an average projection compression ratio, call the compression threshold, compare the compression threshold with the average projection compression ratio, select the high curvature influence region corresponding to the average projection compression ratio being greater than or equal to the compression threshold as the target region, extract the vertices inside the target region to generate a set of high curvature vertices to be unfolded, and extract the local pixel range corresponding to the target region in the two-dimensional image according to the camera perspective relationship to generate a two-dimensional image to be unfolded.
4. The method for detecting defects in a three-dimensional model based on image recognition according to claim 3, characterized in that, The process of selecting high curvature influence regions corresponding to average projection compression ratios greater than or equal to the compression threshold as target regions is as follows: Obtain the mean calculation instruction, read the projection compression ratio of all vertices in the high curvature influence area, and perform cumulative calculation based on the projection compression ratio of each vertex to obtain the cumulative sum value; The total number of vertices contained in the high curvature influence region is read, and the average projection compression ratio is obtained by performing a division operation based on the accumulated sum value and the total number of vertices. Obtain the maximum optical distortion coefficient corresponding to the camera intrinsic parameter matrix and the minimum resolvable pixel size corresponding to the two-dimensional image; The initial compression tolerance limit is set by calculating the function mapping based on the maximum optical distortion coefficient and the minimum resolvable pixel size. Obtain the local maximum principal curvature value corresponding to the high curvature influence area, and adjust the initial compression tolerance limit nonlinearly based on the local maximum principal curvature value to set the compression threshold. The compression threshold and the average projection compression ratio are compared according to a preset numerical comparison rule. Based on the comparison results where the average projection compression ratio is greater than or equal to the compression threshold, the corresponding high curvature influence area is extracted and marked as the target area.
5. The method for detecting defects in a three-dimensional model based on image recognition according to claim 1, characterized in that, The specific steps of S3 are as follows: S311: Obtain the camera intrinsic parameter matrix, calculate the spatial coordinates of the vertices inside the set of high curvature vertices to be unfolded according to the spatial coordinates of the vertices inside the set of high curvature vertices to be unfolded, calculate the spatial coordinates along the principal curvature direction and the minimum curvature direction using the B-spline surface fitting algorithm, generate a local parameterized surface, discretize and sample the local parameterized surface according to the equal arc length distribution condition, and generate a three-dimensional mesh lattice and the corresponding two-dimensional mesh lattice. S312: Call the camera intrinsic and extrinsic matrix of the corresponding viewpoint camera, perform perspective projection transformation on the three-dimensional grid dot matrix, obtain the set of pixel coordinates inside the two-dimensional image to be unfolded, extract the grayscale information of the adjacent integer pixels around the set of pixel coordinates inside the two-dimensional image to be unfolded, calculate the sub-pixel grayscale interpolation data using the bicubic interpolation algorithm based on the grayscale information as the sub-pixel grayscale interpolation result, assign the sub-pixel grayscale interpolation result to the pixel position corresponding to the two-dimensional grid dot matrix, and generate the unfolded image; S313: Extract the row and column indices of the corresponding two-dimensional grid points inside the unfolded image, extract the corresponding three-dimensional coordinates inside the local parameterized surface, establish the mapping relationship between the row and column indices and the three-dimensional coordinates, generate mapping data items, and summarize the mapping data items to generate an unfolded image mapping table.
6. The method for detecting defects in a three-dimensional model based on image recognition according to claim 1, characterized in that, The specific steps of S4 are as follows: S411: Obtain the point spread function and diffuse reflection attenuation model, call the unfolded image mapping table to reversely retrieve the corresponding pixel trajectory of the prior feature line inside the unfolded image, calculate the Gaussian dilation weight value corresponding to the expansion of the pixel trajectory to both sides according to the point spread function and the diffuse reflection attenuation model, and generate a prior edge mask. S412: For the unfolded image, a multi-branch feature decoupling neural network is invoked to extract the corresponding basic edge gradient and multi-scale texture features. The basic edge gradient and multi-scale texture features are integrated to establish a basic feature map. Based on the autoencoder branch inside the multi-branch feature decoupling neural network and combined with the prior edge mask indication corresponding to the spatial attention mechanism, the normal edge highlight response corresponding to the basic feature map is calculated to establish a background feature map. For the basic feature map and the background feature map, pixel-by-pixel gray-level difference value calculation is performed to extract the difference features and generate an anomaly probability map. S413: Invoke the adaptive Otsu method that combines local neighborhood mean to segment the anomaly probability map and extract connected regions. Compare the area values of connected regions according to a preset area resolution scale threshold, and select connected regions whose area values exceed the preset area resolution scale threshold as target connected regions. Calculate the centroid coordinates of the target connected regions in the unfolded image and generate two-dimensional coordinates of candidate defects.
7. The method for detecting defects in a three-dimensional model based on image recognition according to claim 6, characterized in that, The process of generating the two-dimensional coordinates of the candidate defects is as follows: Calculate the mean probability of local windows of pixels in the anomaly probability map and adjust the global threshold of the Otsu method to set an adaptive segmentation threshold; Based on the adaptive segmentation threshold, the anomaly probability map is binarized to generate a binary image containing connected regions. The area of the connected region is generated by counting the number of pixels in the connected region. The preset area resolution scale threshold is set according to the preset tolerance defect physical area and spatial mapping scale parameters; Extract connected regions whose area values are greater than the preset area resolution scale threshold as target connected regions; The zero-order spatial moment and first-order spatial moment of the target connected region within the unfolded image are calculated, and a division operation is performed to generate two-dimensional coordinates of the candidate defect.
8. The method for detecting defects in a three-dimensional model based on image recognition according to claim 1, characterized in that, The specific steps of S5 are as follows: S511: Obtain the preset tolerance distance range, call the unfolded image mapping table, use the bilinear interpolation algorithm to retrieve the candidate defect two-dimensional coordinates and map them to the corresponding three-dimensional initial coordinates and parameterized surface normal vector inside the local parameterized surface, take the three-dimensional initial coordinates as the starting node, establish reverse search line segments along the positive and negative directions corresponding to the parameterized surface normal vector within the tolerance distance range, and generate spatial search line segments. S512: Extract the corresponding principal gray-level gradient direction inside the unfolded image for the two-dimensional coordinates of the candidate defect, call the unfolded image mapping table to perform back projection parameter transformation processing operation, project the principal gray-level gradient direction into the three-dimensional world coordinate system, and generate a local texture normal vector. S513: Obtain a preset positive threshold, call the spatial search line segment and the internal triangular facet, calculate the facet normal vector corresponding to the geometric intersection point generated by the two, perform the corresponding dot product calculation operation between the facet normal vector and the local texture normal vector, compare the positive threshold with the dot product value, filter the valid intersection points corresponding to the dot product value being greater than the positive threshold, perform spatial feature measurement operation on the valid intersection points and the three-dimensional initial coordinates, extract the spatial Euclidean distance between the valid intersection points and the three-dimensional initial coordinates, retrieve the valid intersection point corresponding to the smallest spatial Euclidean distance value, and generate the three-dimensional defect binding position.
9. The method for detecting defects in a three-dimensional model based on image recognition according to claim 8, characterized in that, The process of generating the three-dimensional defect binding location is as follows: Extract the allowable deflection angle of the surface normal and calculate the corresponding cosine value of the allowable deflection angle of the surface normal; extract the patch fitting error parameters corresponding to the three-dimensional model analysis stage; combine the cosine value and the patch fitting error parameters to perform subtraction calculation to set the preset positive threshold; Extract the internal triangular facets corresponding to the bounding box coverage area of the spatial search line segment; calculate the spatial coordinates of the intersection between the spatial search line segment and the internal triangular facets to generate a geometric intersection point; extract the unit perpendicular vector corresponding to the internal triangular facet at the geometric intersection point to generate a facet normal vector; extract the facet normal vector and the local texture normal vector in the three-dimensional coordinate system corresponding to the three coordinate axis dimensional components; Perform product calculations on the corresponding dimension component values and then perform cumulative summation to generate dot product values; Extract the geometric intersections whose dot product values are greater than the preset positive threshold to generate valid intersections; Calculate the coordinate difference between the effective intersection point and the three-dimensional initial coordinates in the three coordinate axes; The system performs a square calculation and a summation calculation on the coordinate differences corresponding to the three coordinate axes to generate a distance sum of squares; it then performs a square root calculation on the distance sum of squares to generate a spatial Euclidean distance; it compares the values of the spatial Euclidean distances corresponding to all valid intersection points; and it extracts the valid intersection points corresponding to the spatial Euclidean distances in the minimum value state to generate the three-dimensional defect binding position.
10. A three-dimensional model defect detection system based on image recognition, characterized in that, The system is used to implement the image recognition-based three-dimensional model defect detection method as described in any one of claims 1-9, the system comprising: The model parsing module parses the 3D model file of the metal structure to extract the internal triangular facets and vertex set, fits the spatial coordinates to generate the principal curvature of the vertex, merges the vertices whose absolute value of the principal curvature is greater than the preset curvature threshold to generate a high curvature influence area, and aggregates the dihedral angles between adjacent triangular facets and the extreme points of the internal principal curvature within the high curvature influence area to generate a priori feature lines. The distortion assessment module acquires two-dimensional images of the metal structure and camera extrinsic matrix from various acquisition angles, constructs a normalized line-of-sight vector corresponding to the spatial coordinates, calculates the dot product of the normalized line-of-sight vector and the average normal vector of the vertex to generate the projection compression ratio, selects high curvature influence areas with an average projection compression ratio greater than or equal to the compression threshold as target areas, extracts vertices inside the target areas to generate a set of high curvature vertices to be unfolded, and extracts the local pixel range corresponding to the target areas in the two-dimensional images to generate a two-dimensional image to be unfolded. The image reconstruction module fits the set of high curvature vertices to be unfolded to generate a local parameterized surface, generates a three-dimensional mesh matrix and a two-dimensional mesh matrix through discretization sampling, and generates an unfolded image mapping table by associating the two-dimensional pixel coordinates of the two-dimensional mesh matrix with the three-dimensional spatial coordinate index of the three-dimensional mesh matrix. The defect identification module expands to both sides based on the unfolded image mapping table and prior feature lines to generate a prior edge mask, calls a multi-branch feature decoupling neural network to extract features from the prior edge mask, calculates differential features and generates an anomaly probability map based on the differential features, segments the anomaly probability map and filters target connected regions whose area exceeds a preset area resolution scale threshold, and calculates the centroid coordinates of the target connected regions to generate two-dimensional coordinates of candidate defects. The spatial positioning module constructs a spatial search line segment based on the three-dimensional initial coordinates mapped from the two-dimensional coordinates of the candidate defect, and generates a local texture normal vector by back projection. It calculates the dot product between the normal vector of the facet at the intersection of the spatial search line segment and the internal triangular facet and the local texture normal vector. It extracts the geometric intersection point with the dot product value greater than a preset positive threshold and the smallest spatial Euclidean distance as the effective intersection point, and generates the three-dimensional defect binding position based on the effective intersection point.