Deep learning-based automatic detection method for noble metal casting sand hole defects

CN122737451APending Publication Date: 2026-09-11HANGZHOU SHENGYUAN JEWELS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610648054.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

若提高检测网络的灵敏度,会导致对正常工艺设计产生海量误报;若降低灵敏度,则会引发严重的漏检

Benefits of technology

1.本发明通过构建多模态时序光度张量并结合灰度波动率生成方差权重图,实现了非线性的反照率调制。这一机制能够有效过滤光源干扰分量,从而逆向解构出反映材质本征属性的伪漫反射图谱和相对拓扑形貌图谱。该技术方案有效降低了贵金属表面反光特性对目标特征提取的干扰,提升了算法模型在复杂光照环境下的鲁棒性,使得微观几何特征和真实表面材质的提取更为清晰、准确。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737451A_ABST
    Figure CN122737451A_ABST
Patent Text Reader

Abstract

This invention relates to the field of visual inspection technology, specifically to an automatic detection method for sand hole defects in precious metal castings based on deep learning. The invention acquires multi-directional temporal exposure images and constructs a multimodal temporal photometric tensor; calculates pixel-level grayscale volatility to generate a variance weight map; performs nonlinear albedo modulation on the tensor to deconstruct a pseudo-diffuse reflection map and a relative topological morphology map; transforms the CAD model into an implicit neural field and aligns it with the morphology map, rendering a Gaussian soft boundary mask covering a compliant texture; extracts deep features from the pseudo-diffuse reflection map and performs element-wise multiplication suppression with the mask to obtain a pure feature tensor; inputs the tensor into a detection network and outputs the detection results. This invention utilizes variance to decouple optical noise and combines CAD priors to shield complex design interference, significantly reducing false alarms and missed detections under highly reflective complex curved surfaces, achieving accurate extraction of micron-level sand holes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual inspection technology, specifically to an automatic detection method for sand hole defects in precious metal castings based on deep learning. Background Technology

[0002] The forming process of precious metals has long relied on precision investment casting technology. During the physicochemical process of casting, factors such as fluctuations in melting temperature, sand spalling, or residual non-metallic inclusions can easily lead to "sand hole" defects on the surface and inside of castings, severely compromising the aesthetics and structural density of high-end consumer products. With the development of intelligent manufacturing, deep learning-based automated optical inspection technology has gradually been introduced to replace manual visual inspection. However, when faced with the extremely unique application scenario of "highly reflective, complex curved precious metal surfaces," existing deep learning inspection methods reveal significant limitations in their underlying technology.

[0003] First, precious metal surfaces exhibit extremely strong specular reflectivity. Under structured light projection during automated inspection, the curved metal surface produces numerous artifacts with extremely high brightness and deep shadows. Existing surface defect detection technologies typically attempt to identify defect differences by extracting the maximum / minimum grayscale value features of the image from multi-angle illumination. However, the specular reflection of precious metals is highly dynamic, and the position of the false spots shifts drastically with changes in the light source angle. The edges of these moving spots still generate strong high-frequency gradients, making them highly susceptible to misclassification as holes by deep learning networks, resulting in a very high false alarm rate on actual production lines. Existing technologies have failed to achieve a complete mathematical decoupling of "dynamic physical lighting noise" from "real geometric topological concavity features" at the underlying computational level. Second, precious metal jewelry and other castings are essentially complex three-dimensional topologies woven from highly irregular free-form surfaces, intricate micro-carved textures, and deep hollows. Existing multi-scale feature extraction or contrast analysis algorithms, when expanding the receptive field to extract micron-level sand hole features, easily obscure tiny shallow defect features by the surrounding strong and complex process textures. More critically, in terms of physical scale and contrast, the "compliant micro-engraving gaps" deliberately left by designers are extremely similar to "illegal sand holes." Increasing the sensitivity of the detection network would lead to a massive number of false alarms for normal process designs; decreasing the sensitivity would result in serious missed detections. Existing detection networks lack a tensor-level forced suppression mechanism capable of actively sensing and shielding compliant design textures.

[0004] In summary, how to overcome the interference of strong optical noise and inherent complex textures in a highly reflective background with complex topological surfaces to achieve accurate extraction and identification of sand hole defects is a technical problem that urgently needs to be solved in this field.

[0005] To address this, a deep learning-based automatic detection method for sand hole defects in precious metal castings is proposed. Summary of the Invention

[0006] The purpose of this invention is to provide an automatic detection method for sand hole defects in precious metal castings based on deep learning. This invention acquires multi-directional temporal exposure images and constructs a multimodal temporal photometric tensor; calculates pixel-level grayscale volatility to generate a variance weight map, performs nonlinear albedo modulation on the tensor, and deconstructs a pseudo-diffuse reflection map and a relative topological morphology map; transforms the CAD model into an implicit neural field and aligns it with the morphology map, rendering a Gaussian soft boundary mask covering a compliant texture; extracts deep features from the pseudo-diffuse reflection map and performs element-wise multiplication suppression with the mask to obtain a pure feature tensor; inputs the tensor into a detection network and outputs the detection results. This invention utilizes variance to decouple optical noise and combines CAD priors to shield complex design interference, significantly reducing false alarms and missed detections under highly reflective complex curved surfaces, achieving accurate extraction of micron-level sand holes.

[0007] To achieve the above objectives, the present invention provides the following technical solution: An automatic detection method for sand hole defects in precious metal castings based on deep learning includes: Acquire temporal exposure image sequences of precious metal castings under illumination from light sources in different orientations, and stack them in the channel dimension to construct a multimodal temporal photometric tensor; The gray-level fluctuation rate of each spatial coordinate of the multimodal temporal photometric tensor in the channel dimension is calculated to generate a variance weight map. The variance weight map is used to perform nonlinear albedo modulation on the multimodal temporal photometric tensor. Based on the gray-level fluctuation rate, the light source interference components that cause specular reflection are filtered out, and the components are deconstructed in reverse into a pseudo diffuse reflection spectrum that characterizes the intrinsic properties of the material and a relative topological morphology spectrum that characterizes the micro-geometric features. The standard CAD design model of precious metal castings is transformed into an implicit neural field representation and spatially affine aligned with the relative topological map. The resulting rendering generates a topological constraint mask tensor that covers the compliant process texture. The deep feature map of the pseudo diffuse reflection map is extracted, and the topological constraint mask tensor is element-wise multiplied with the deep feature map to suppress it, so as to obtain a pure feature tensor that filters out compliant process textures; the pure feature tensor is input into the visual defect detection network for feature extraction, and the sand hole detection result is output.

[0008] Preferably, acquiring a sequence of temporal exposure images of a precious metal casting under illumination from different directional light sources, and constructing a multimodal temporal photometric tensor, includes: controlling a dome light source, which serves as the light source, to sequentially illuminate individual illumination zones at a preset illumination angle; triggering an industrial camera to acquire fixed exposure parameters of the precious metal casting under each illumination zone illumination state, and acquiring single-channel grayscale images containing different physical shadow distributions; and stitching and fusing the pixel matrices of all single-channel grayscale images in the depth channel direction according to the illumination time sequence of the light source zones, outputting a multimodal temporal photometric tensor with a three-dimensional matrix structure, wherein the width and height of the tensor correspond to the physical pixel size of the image, and the channel depth of the tensor corresponds to the total number of illumination zones.

[0009] Preferably, the calculation of grayscale volatility of each spatial coordinate in the multimodal temporal photometric tensor along the channel dimension and the generation of a variance weight map includes: traversing the pixel spatial coordinates in the multimodal temporal photometric tensor; extracting the pixel values ​​of a single pixel spatial coordinate at all channel depths and calculating the arithmetic mean and variance as grayscale volatility; inputting the grayscale volatility corresponding to all pixel spatial coordinates into a nonlinear activation function for normalization mapping; subtracting the normalized grayscale volatility from the value one to generate a variance weight map corresponding to each pixel with values ​​ranging from zero to one.

[0010] Preferably, the nonlinear albedo modulation of the multimodal temporal photometric tensor is performed using a variance weighting map, and then deconstructed into a pseudo diffuse reflectance map and a relative topological map. This includes: performing a pixel-by-pixel Hadamard product operation on the variance weighting map and the multimodal temporal photometric tensor to attenuate specular reflection interference components and outputting a modulated photometric tensor; calculating the integral mean of all pixel values ​​along the channel depth direction of the modulated photometric tensor and outputting a single-channel two-dimensional pixel array as a pseudo diffuse reflectance map; extracting the photometric partial derivatives between adjacent pixel coordinates in the modulated photometric tensor, calculating the three-dimensional surface normal vector for each spatial coordinate, mapping all three-dimensional surface normal vectors to a two-dimensional plane, and generating the relative topological map.

[0011] Preferably, converting a standard CAD design model of a precious metal casting into an implicit neural field representation includes: performing spatial sampling within the three-dimensional bounding box of the CAD design model to obtain a three-dimensional coordinate set; inputting the three-dimensional coordinate set into a multilayer perceptron network containing position encoding; training the multilayer perceptron network to fit the signed distance function value from each three-dimensional coordinate to the surface of the CAD design model; and using the weight parameters of the trained multilayer perceptron network as the implicit neural field representation.

[0012] Preferably, spatial affine alignment with the relative topological map is performed to render and generate a topological constraint mask tensor covering the compliant process texture. This includes: extracting a feature point array whose normal vector change rate exceeds a preset gradient from the relative topological map; constructing a spatial pose optimization objective function, with the goal of minimizing the sum of squares of the signed distance function values ​​output after the feature point array is input to the implicit neural field representation, and iteratively optimizing the rotation and translation transformation matrix through a gradient descent algorithm to complete the spatial affine alignment; in the aligned state, emitting virtual rays along the virtual optical center viewpoint of the industrial camera towards the implicit neural field representation, and recording the coordinates of the intersection points between the rays and the surface of the implicit neural field; allocating spatial weights according to the structural attributes to which the intersection coordinates belong, and projecting to generate a two-dimensional mask matrix as the topological constraint mask tensor.

[0013] Preferably, the deep feature map of the pseudo diffuse reflection map is extracted, and the topological constraint mask tensor is suppressed by element-wise multiplication with the deep feature map, including: inputting the pseudo diffuse reflection map into the backbone feature extraction network to output a multi-channel deep feature map; scaling the spatial resolution of the topological constraint mask tensor to match the pixel size of the deep feature map using a bilinear interpolation algorithm; copying and expanding the dimension of the scaled topological constraint mask tensor in the channel direction to make its number of channels equal to the number of channels in the deep feature map, generating a suppression tensor; performing a pixel-by-pixel multiplication operation on the suppression tensor and the deep feature map to clear the activation values ​​of neurons corresponding to the compliant process texture regions, and outputting the clean feature tensor.

[0014] Preferably, the pure feature tensor is input into the visual defect detection network for feature extraction, and the trachoma detection result is output. This includes: dividing the pure feature tensor into non-overlapping local windows, calculating the self-attention relationship between pixel features within each local window, and obtaining local context-dependent features; performing a window movement operation to create a spatial overlap intersection between adjacent non-overlapping local windows, and calculating the cross-window attention relationship matrix between pixel feature vectors again within the overlap intersection to obtain global topological dependency features; inputting the local context-dependent features and the global topological dependency features into a feature pyramid fusion module, and outputting a trachoma detection result including trachoma bounding box coordinates, category confidence scores, and instance segmentation boundaries.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention achieves nonlinear albedo modulation by constructing a multimodal temporal photometric tensor and combining it with grayscale volatility to generate a variance weight map. This mechanism effectively filters out light source interference components, thereby inversely deconstructing pseudo-diffuse reflectance and relative topological morphology maps that reflect the intrinsic properties of the material. This technical solution effectively reduces the interference of the reflective properties of noble metal surfaces on target feature extraction, improves the robustness of the algorithm model in complex lighting environments, and makes the extraction of microscopic geometric features and real surface materials clearer and more accurate.

[0016] 2. This invention introduces an implicit neural field representation based on a standard CAD design model and aligns it with the topological topography map space to generate a topological constraint mask tensor, thereby suppressing deep feature maps through element-wise multiplication. This mechanism, which integrates prior physical design intent with deep neural networks, effectively shields the interference of inherent complex openwork and carving textures on the surface of jewelry, allowing the visual inspection network to focus more on randomly occurring abnormal areas, thus effectively reducing the probability of misjudging normal process design features as defects.

[0017] 3. This invention provides a high signal-to-noise ratio (SNR) clean feature tensor for the subsequent visual defect detection network through pre-processing optical signal decoupling and structural feature purification, optimizing the feature extraction path for micro-targets. Under low background noise input conditions, the detection network can more sensitively capture the local context-dependent and global topological dependent features of micron-sized pinholes. This approach effectively improves the overall accuracy of the system in identifying micro-defects, while also enhancing the detection model's generalization ability and adaptability when facing products with different styles and curvature surfaces. Attached Figure Description

[0018] Figure 1 A flowchart of an automatic detection method for sand hole defects in precious metal casting based on deep learning, provided in an embodiment of the present invention; Figure 2 This is a data flow diagram of photometric decoupling and albedo modulation provided in an embodiment of the present invention; Figure 3 The diagram illustrates the principle of implicit neural field alignment and feature multiplication inhibition provided in this embodiment of the invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1 to 3This invention provides an automatic detection method for sand hole defects in precious metal castings based on deep learning. The technical solution is as follows: An automatic detection method for sand hole defects in precious metal castings based on deep learning includes: Acquire temporal exposure image sequences of precious metal castings under illumination from light sources in different orientations, and stack them in the channel dimension to construct a multimodal temporal photometric tensor; The gray-level fluctuation rate of each spatial coordinate of the multimodal temporal photometric tensor in the channel dimension is calculated to generate a variance weight map. The variance weight map is used to perform nonlinear albedo modulation on the multimodal temporal photometric tensor. Based on the gray-level fluctuation rate, the light source interference components that cause specular reflection are filtered out, and the components are deconstructed in reverse into a pseudo diffuse reflection spectrum that characterizes the intrinsic properties of the material and a relative topological morphology spectrum that characterizes the micro-geometric features. The standard CAD design model of precious metal castings is transformed into an implicit neural field representation and spatially affine aligned with the relative topological map. The resulting rendering generates a topological constraint mask tensor that covers the compliant process texture. The deep feature map of the pseudo diffuse reflection map is extracted, and the topological constraint mask tensor is element-wise multiplied with the deep feature map to suppress it, so as to obtain a pure feature tensor that filters out compliant process textures; the pure feature tensor is input into the visual defect detection network for feature extraction, and the sand hole detection result is output.

[0021] Example 1: This example demonstrates an application scenario in the automated quality inspection production line of an 18K gold intricate openwork ring at a high-end jewelry manufacturing company. This particular ring features a large area of ​​mirror-polished surface, along with meticulously detailed micron-level carvings and deep-hole prongs intentionally left by the designer. In traditional inspection methods, the mirror-like reflections and complex carvings result in a high false alarm rate for traditional visual algorithms, making it impossible to accurately identify minute pinholes caused by defects in the lost-wax casting process.

[0022] As one embodiment of the present invention, refer to Figure 1 Flowchart of an automatic detection method for sand hole defects in precious metal casting based on deep learning, refer to Figure 2 Photometric decoupling and albedo modulation data flow diagram, see reference Figure 3 Schematic diagram of implicit neural field alignment and feature multiplication inhibition principle.

[0023] Furthermore, a time-series exposure image sequence of the precious metal casting under illumination from different directional light sources is acquired, and a multimodal temporal photometric tensor is constructed. This includes: controlling the dome light source, which serves as the light source, to sequentially illuminate individual illumination zones at a preset illumination angle; triggering an industrial camera to acquire fixed exposure parameters of the precious metal casting under each illumination zone illumination state, thereby acquiring single-channel grayscale images containing different physical shadow distributions; and stitching and fusing the pixel matrices of all single-channel grayscale images in the depth channel direction according to the illumination time sequence of the light source zones, outputting a multimodal temporal photometric tensor with a three-dimensional matrix structure, wherein the width and height of the tensor correspond to the physical pixel size of the image, and the channel depth of the tensor corresponds to the total number of illumination zones.

[0024] Specifically, in the actual quality inspection station of this embodiment, a hemispherical programmable LED dome light source with 16 independently controllable zones is used. The control system sends a trigger signal to sequentially illuminate individual zones in a clockwise direction, with the illumination duration of each zone set to 15 milliseconds. Simultaneously, an industrial camera equipped with a telecentric lens is triggered to perform fixed-exposure shooting, with the camera's exposure time strictly locked at 10 milliseconds to avoid motion blur and ambient light interference. After one complete lighting cycle, 16 single-channel grayscale images (e.g., with a resolution of 4000×3000) are acquired. These 16 images record the brightness changes and highlight drift state of the ring under different incident light angles. Subsequently, these 16 two-dimensional pixel matrices are stacked and assembled in the depth direction according to chronological order via the memory data bus. Finally, a three-dimensional data matrix with a width of 4000, a height of 3000, and a channel depth of 16 is constructed in the video memory, namely the multimodal temporal photometric tensor. Each spatial coordinate point in this tensor contains a one-dimensional vector of length 16, which fully records the temporal characteristics of the optical response of the physical point under different lighting conditions, providing basic data support for subsequent optical noise separation.

[0025] This invention ensures that optical response data of the same physical location under different lighting conditions can be accurately captured and structured by strictly controlling the timing of lighting angle, exposure time, and image matrix assembly process. This technique effectively avoids the loss of local shadow blind spots or overexposure caused by single lighting, providing a rich and informative raw tensor for subsequent in-depth analysis, and effectively improving the system's data fault tolerance in complex optical environments.

[0026] Further, the grayscale volatility of each spatial coordinate in the multimodal temporal photometric tensor along the channel dimension is calculated to generate a variance weight map. This includes: traversing the pixel spatial coordinates in the multimodal temporal photometric tensor; extracting the pixel values ​​of a single pixel spatial coordinate at all channel depths and calculating the arithmetic mean and variance as the grayscale volatility; inputting the grayscale volatility corresponding to all pixel spatial coordinates into a nonlinear activation function for normalization mapping; and subtracting the normalized grayscale volatility from the value one to generate a variance weight map corresponding to each pixel with values ​​ranging from zero to one.

[0027] Specifically, for the constructed 4000×3000×16 photometric tensor, the system processor executes a double-nested loop to traverse all 12 million spatial coordinates. For any given coordinate point, 16 grayscale values ​​are extracted from its depth channel. First, the algebraic sum of these 16 values ​​is calculated and divided by 16 to obtain the arithmetic mean baseline. Next, the deviation of each of the 16 independent grayscale values ​​from this baseline is calculated, the deviation values ​​are squared to eliminate the influence of positive and negative signs, and then the average of these 16 squared values ​​is calculated to obtain the statistical variance of the pixel. The physical meaning of this variance is that real microscopic pinholes exhibit stable dark area scattering under various lighting conditions, and their variance is usually less than 15; while the specular reflections on metal surfaces undergo drastic spatial displacement with the switching of light source zones, causing the grayscale value at the same coordinate to jump from 0 to 255, and their variance is often greater than 8000. To transform this vast range of variance values ​​into a soft mask recognizable by the neural network, all variance values ​​are input into a sigmoid nonlinear activation function. For highly varianced reflective points, the function maps them to values ​​close to 1; for low-variance imperfections or realistic textures, the function maps them to values ​​close to 0. Finally, an inversion operation is performed, subtracting the mapping result from the constant 1. After processing, specular artifact regions receive suppression weights close to 0, while realistic texture regions receive preservation weights close to 1. The final output is a single-channel floating-point weight matrix of size 4000×3000.

[0028] This invention utilizes the representational properties of statistical variance in the channel dimension to transform the complex physical problem of specular reflection into a purely mathematical tensor calculation process. Through normalization and inverse mapping, the location and intensity of dynamic specular artifacts are autonomously located and locked without relying on any prior physical reflection model. This mechanism effectively avoids the oversegmentation phenomenon that traditional thresholding methods are prone to under dynamic lighting, providing precise mathematical gating weights for subsequent purification of real defect features without lighting noise.

[0029] Furthermore, nonlinear albedo modulation is performed on the multimodal temporal photometric tensor using a variance weighting map, and then deconstructed into a pseudo diffuse reflectance map and a relative topological map. This includes: performing a pixel-by-pixel Hadamard product operation on the variance weighting map and the multimodal temporal photometric tensor to attenuate specular reflection interference components and outputting a modulated photometric tensor; calculating the integral mean of all pixel values ​​along the channel depth direction of the modulated photometric tensor and outputting a single-channel two-dimensional pixel array as a pseudo diffuse reflectance map; extracting the photometric partial derivatives between adjacent pixel coordinates in the modulated photometric tensor, calculating the three-dimensional surface normal vector for each spatial coordinate, mapping all three-dimensional surface normal vectors to a two-dimensional plane, and generating the relative topological map.

[0030] Solving for the three-dimensional surface normal vector for each spatial coordinate includes: constructing a two-way reflection distribution function model of micro-element containing Fresnel reflection terms, geometric occlusion terms, and surface roughness distribution terms; The micro-area bidirectional reflection distribution function model specifically adopts the Cook-Torrance model architecture, and its expression is as follows: ; Among them, the surface roughness distribution term The GGX normal distribution function is adopted based on the set empirical coefficient of 18K gold surface roughness; Fresnel reflection term. The Schlick approximation is used to calculate reflectivity attenuation at different viewing angles; geometric occlusion term. The Smith masking-shadowing function is employed. The roughness parameters of the model are calibrated by pre-scanning a standard 18K gold sample. Then, the photometric partial derivative of the modulated photometric tensor is substituted into the micro-element bidirectional reflectance distribution function model to calculate the nonlinear photometric distortion compensation value caused by the uneven surface roughness of the precious metal. The nonlinear photometric distortion compensation value is then jointly optimized with the solution results of the basic photometric solid equation to correct and output the three-dimensional surface normal vector.

[0031] Specifically, the 4000×3000 variance weight map generated in the previous stage is extracted and used as a multiplication operator. It is then multiplied element-wise with each of the 16 original temporal grayscale images using Hadamard products that are spatially aligned. Since the values ​​of the highlight regions in the weight map are close to 0, this multiplication operation physically suppresses and clears the overexposed areas in the original images caused by specular reflection, thus outputting a 16-channel modulated photometric tensor that eliminates glaring spots. Next, all pixels are accumulated along these 16 channels, and the average value is calculated. This integration process effectively homogenizes the diffuse light intensity in multi-light source environments, removing the unevenness of light and shadow caused by directional illumination, generating a grayscale image similar to one taken under uniform, shadowless lighting—a pseudo-diffuse reflection map.

[0032] Simultaneously, to restore the three-dimensional shape, the spatial photometric partial derivatives between adjacent pixels are calculated within the modulated photometric tensor. Although the pre-applied Hadamard product has cleared macroscopic specular overexposure points, at the microscopic scale, the uneven roughness of the complex curved surfaces of precious metals still leads to residual nonlinear photometric distortion. Therefore, the basic three-dimensional linear photometric equations are first invoked, and the initial values ​​of the basic three-dimensional normals are obtained by combining 16 known light source azimuth angles. Subsequently, a micro-element bidirectional reflectance distribution function model for precious metal materials is constructed. This model rigorously defines the Fresnel reflection term (characterizing the change of reflectivity with the observation angle), the geometric occlusion term (characterizing the mutual occlusion effect between micro-protrusions), and the surface roughness distribution term (characterizing the statistical probability distribution of the micro-surface normals). Substituting the extracted photometric partial derivatives into this BRDF model, the compensation values ​​for nonlinear photometric distortion caused by local uneven roughness and micro-occlusion are accurately calculated. Next, a joint optimization algorithm is employed, using the distortion compensation value as a priori constraint, to perform nonlinear iterative correction on the initial value of the basic normal, thereby calculating the pixel-level 3D surface normal vector that completely eliminates microscopic optical distortion. Finally, the X, Y, and Z components of these high-precision 3D normal vectors are linearly mapped onto the RGB color channels to generate a color topological map containing extremely faithful information on the surface's microscopic concavity and convexity slope.

[0033] This invention successfully deconstructs the extremely complex nonlinear optical response of noble metals through a deep integration of rigorous tensor-level Hadamard multiplication operations and the BRDF micro-surface element physical compensation mechanism. This method not only restores the true intrinsic albedo characteristics of the material surface but also overcomes the photometric distortion problems caused by microscopic surface roughness and geometric occlusion without adding additional 3D scanning hardware. High-fidelity microscopic 3D normal morphology is reconstructed using low-cost 2D images, enabling the detection network to cross-verify defects from two orthogonal dimensions: pure material albedo and precise topological geometry. This significantly improves the accuracy and recognition of hidden, weak pinholes and high-curvature surface defects.

[0034] Furthermore, the standard CAD design model of the precious metal casting is transformed into an implicit neural field representation, including: spatial sampling within the three-dimensional bounding box of the CAD design model to obtain a three-dimensional coordinate set; inputting the three-dimensional coordinate set into a multilayer perceptron network containing position encoding; training the multilayer perceptron network to fit the signed distance function value from each three-dimensional coordinate to the surface of the CAD design model, and using the weight parameters of the trained multilayer perceptron network as the implicit neural field representation.

[0035] Specifically, the engineering design drawings (usually in STEP or STL format) of the 18K gold openwork ring are pre-imported. The parsing engine reads the 3D coordinates of hundreds of thousands of triangular facet vertices and their topological connections within the file. Then, a multilayer perceptron network with 8 layers and 256 hidden nodes per layer is initialized in memory, and high-frequency sine and cosine functions are added to the input to achieve coordinate position encoding, enhancing the network's ability to capture high-frequency, complex engraving details. A virtual 3D cube bounding box is constructed based on the ring's maximum outline, and 3 million sampling coordinate points are uniformly generated within the box using a random dithering strategy. For each sampling point, the shortest vertical distance from it to the CAD facet surface is calculated using traditional computational geometry algorithms. If the point is inside the model, the distance is marked as negative; if it is outside, it is marked as positive; and if it is on the surface, it is marked as 0. This constitutes a set of signed SDF (distance data). Next, these 3 million sampling coordinates are used as input, and their corresponding SDF values ​​are used as training labels. The backpropagation algorithm drives the multilayer perceptron for iterative learning. After approximately 500 training iterations, the neural network learns the spatial geometric distribution patterns of the ring. At this point, the massive original CAD triangular facet data is discarded, and the entire multilayer perceptron containing hundreds of thousands of floating-point weight parameters is preserved as an implicit neural field.

[0036] This invention abandons traditional discrete 3D mesh matching methods and innovatively compresses the geometric boundaries of CAD design models into a continuous and differentiable neural network weight. This implicit neural field representation method not only significantly reduces the storage memory required for 3D topological features, but more importantly, it enables continuous querying of the model at arbitrary spatial resolutions. This provides a mathematical foundation for seamless pixel-level alignment of 3D design intent with 2D high-resolution images, effectively bridging the gap between different data modalities.

[0037] Further, spatial affine alignment is performed with the relative topological map, and a topological constraint mask tensor covering the compliant process texture is generated. This includes: extracting a feature point array whose normal vector change rate exceeds a preset gradient from the relative topological map; constructing a spatial pose optimization objective function, with the goal of minimizing the sum of squares of the signed distance function values ​​output after the feature point array is input to the implicit neural field representation, and iteratively optimizing the rotation and translation transformation matrix through a gradient descent algorithm to complete the spatial affine alignment; in the aligned state, emitting virtual rays along the virtual optical center viewpoint of the industrial camera towards the implicit neural field representation, and recording the coordinates of the intersection points between the rays and the implicit neural field surface; allocating spatial weights according to the structural attributes to which the intersection coordinates belong, and projecting to generate a two-dimensional mask matrix as the topological constraint mask tensor.

[0038] Specifically, after obtaining the relative topological map derived from actual shooting, gradient calculations are performed, and coordinate points with normal vector deflection angles greater than 15 degrees are specifically selected. These points typically correspond to significant structural features such as the physical edges of the ring and the prong outline. Subsequently, a spatial pose optimization module based on differentiable rendering is initiated. First, a six-DOF pose transformation matrix containing rotation and translation variables is initialized. The extracted feature point array is multiplied by this transformation matrix and mapped to the standard coordinate system of the implicit neural field. Next, the mapped 3D coordinates of the feature points are input one by one into a pre-trained multilayer perceptron network for forward propagation, outputting the symbolic SDF prediction value for each point. To achieve accurate alignment, a geometric alignment loss function is constructed, with the sum of squares of the SDF prediction values ​​of all feature points as its core term. Since the implicit neural field (multilayer perceptron) is inherently end-to-end continuously differentiable, the system directly utilizes gradient descent algorithms (such as the Adam optimizer) to calculate the partial derivatives of the loss function with respect to the six degrees of freedom parameters in the pose transformation matrix and performs backpropagation parameter updates. As the iterations proceed, the pose matrix is ​​continuously optimized, driving the feature point array to gradually approach and closely adhere to the isosurface where the SDF output value is 0 in mathematical space. After approximately several hundred fast gradient iterations, the optimization stops when the loss function converges to a preset small threshold. At this point, the pose transformation matrix represents the precise spatial pose of the actual workpiece in the camera coordinate system.

[0039] Spatial weight allocation is performed based on the structural attributes to which the intersection coordinates belong, including: extracting the three-dimensional boundary of the compliant process texture region, calculating the Euclidean distance decay gradient in the smooth surface region outward from the three-dimensional boundary; applying Gaussian smoothing filter along the Euclidean distance decay gradient to generate a Gaussian soft boundary distance field with continuously gradually changing values; projecting the Gaussian soft boundary distance field onto a two-dimensional plane to generate a continuous gradient matrix with mask values ​​between zero and one, which serves as the topological constraint mask tensor.

[0040] After registration, a virtual camera rendering engine based on a ray-stepping algorithm is constructed. Starting from a virtual optical center, this engine emits 12 million probe rays into the aligned implicit neural field, according to the actual camera's focal length and pixel density. As the probe rays advance in 3D space and iteratively query the local SDF value to approach an implicit surface with an SDF value of 0, the local average curvature of that physical point is calculated using the second-order spatial partial derivative of the implicit neural field at the intersection. When the absolute value of the local average curvature at the intersection is greater than a preset process fluctuation threshold, the intersection is automatically identified as a micro-carved groove or a high-frequency hollowed-out edge preset by the designer; otherwise, it is identified as a smooth surface, thus accurately extracting the 3D geometric boundaries of these compliant texture areas in an unsupervised manner. Using these boundaries as a reference, the Euclidean distance field is calculated to the surrounding smooth surface areas. For intersections directly falling inside the carving, a weight value of 0 is assigned; for intersections in smooth areas far from the carving boundary, a weight value of 1 is assigned. For intersections at the boundary transition zones, a Gaussian smoothing operator with a standard deviation of 1.5 pixels is applied, and a continuous gradient weight between 0.1 and 0.9 is calculated based on the Euclidean distance to the boundary. Finally, through ray tracing projection traversal, a two-dimensional gradient matrix with "soft edges" is rendered. This mask tensor is no longer a rigid black-and-white segmentation, but rather possesses a smooth gray-scale transition zone at the boundary between complex textures and smooth surfaces, accurately simulating the blurring effect in optical imaging.

[0041] This invention addresses the feature abrupt change problem caused by binarized masks at the edges of complex curved surfaces by introducing a Gaussian soft boundary distance field mechanism. This continuous gradient weight allocation effectively prevents signal truncation when deep learning networks process carved edge features, allowing for the complete preservation of weak defect signals immediately adjacent to the texture edge while shielding against process texture interference. This scheme significantly improves the system's ability to capture minute pinhole defects in extremely complex topologies, avoiding false edge reports caused by minor alignment deviations.

[0042] Further, deep feature maps of the pseudo diffuse reflection map are extracted, and element-wise multiplication suppression is performed between the topological constraint mask tensor and the deep feature map. This includes: inputting the pseudo diffuse reflection map into a backbone feature extraction network to output a multi-channel deep feature map; scaling the spatial resolution of the topological constraint mask tensor to match the pixel size of the deep feature map using bilinear interpolation; copying and expanding the dimension of the scaled topological constraint mask tensor in the channel direction to make its number of channels equal to the number of channels in the deep feature map, generating a suppression tensor; performing a pixel-by-pixel multiplication operation between the suppression tensor and the deep feature map to clear the activation values ​​of neurons corresponding to the compliant process texture regions, and outputting the clean feature tensor.

[0043] Specifically, the pseudo diffuse reflection map (with specular noise removed) deconstructed earlier is input into a deep feature extraction backbone network based on the ResNet architecture. After downsampling and nonlinear feature encoding through three stages of residual convolutional blocks and pooling layers, the output size is reduced to one-sixteenth of the original image (i.e., 250 x 187), and the number of channels is expanded to a deep feature map of 256 channels. To enable the two-dimensional mask to be applied to the high-dimensional feature map, the nearest neighbor interpolation algorithm is first used to proportionally reduce the 4000 x 3000 binary topological constraint mask tensor to 250 x 187 to ensure strict alignment of spatial coordinates. Subsequently, a channel broadcasting operation is performed, copying this single-channel mask image 256 times in the depth direction to construct a stereo suppression tensor of 250 x 187 x 256. Finally, in the tensor operation layer of the deep learning framework, this suppression tensor is multiplied with the extracted deep feature map at a very low level. In mathematical logic, the regions with values ​​of 0 in the mask correspond to legitimate intricate carvings and openwork structures on the ring. After multiplication, the activation values ​​of neurons in all 256 channels of the feature map at these legitimate texture space locations are forcibly cleared to zero. Finally, a pure feature tensor is truncated and output, retaining response signals only in unexpected regions (i.e., real, randomly occurring sand hole locations).

[0044] This invention innovatively introduces mandatory gating of external CAD prior knowledge into the mid-stage of feature flow in deep neural networks. Through scale interpolation and multi-channel bitwise multiplication suppression, legitimate texture features conforming to product process design are completely shielded at the computational level. This processing mechanism prevents background textures from swallowing up weak defect features during large-scale convolution operations, effectively reducing the false alarm rate of the system and significantly improving the signal-to-noise ratio of the network for unknown defect signals.

[0045] Further, the pure feature tensor is input into the visual defect detection network for feature extraction, and the trachoma detection result is output. This includes: dividing the pure feature tensor into non-overlapping local windows, calculating the self-attention relationship between pixel features within each local window, and obtaining local context-dependent features; performing a window movement operation to create a spatial overlap intersection between adjacent non-overlapping local windows, and calculating the cross-window attention relationship matrix between pixel feature vectors again within the overlap intersection to obtain global topological dependency features; inputting the local context-dependent features and the global topological dependency features into the feature pyramid fusion module, and outputting the trachoma detection result including trachoma bounding box coordinates, category confidence scores, and instance segmentation boundaries.

[0046] The parameters of the visual defect detection network are pre-trained using supervised learning. The specific training process involves constructing a dataset of thousands of historical quality inspection images of precious metal castings, both with and without pinhole defects, and having professionals annotate real pinhole defects with polygonal instance masks. During training, the cross-entropy loss function is used to calculate the class confidence error, the GIoU loss function to calculate the bounding box coordinate regression error, and the Dice loss function to calculate the instance segmentation boundary error. These three errors are weighted and summed according to a preset ratio to obtain the total loss function. The AdamW optimizer is used to continuously update the network weights through backpropagation until the average accuracy of the model on the validation set reaches a set threshold.

[0047] Specifically, for the feature tensor purified by masking, a Swin-Transformer architecture is used as the detection head. To meet the divisibility requirements of the feature pyramid and self-attention window, a dynamic feature padding operation is performed before inputting the feature map into the visual defect detection network. The modulus of the width and height of the current deep feature map and the local window size 7 is calculated, and zero padding is performed on the right and lower boundaries of the deep feature map, expanding its width from 250 to 252 and its height from 187 to 189, ensuring that the spatial dimension of the padded pure feature tensor is completely divisible by the pixel size of the local window. Subsequently, the aligned pure feature tensor is cut into multiple non-overlapping local windows with an area of ​​7×7 pixels. Within each independent window, the inner product between all 49 feature vectors is calculated based on a query key-value mechanism to generate a local self-attention weight matrix, thereby capturing the dark gradient pattern of extremely small pinhole edges. After completing the local calculation, all windows are uniformly shifted to the right and down by 3 pixels, so that the originally separated local regions have overlapping intersections. Within these new windows containing cross-boundary features, self-attention operations are performed again to achieve information interaction between different local regions, obtaining global topological dependency characteristics covering the entire detection surface. Finally, these feature maps with local sensitivity and a global macroscopic view are fed into the feature pyramid network for multi-scale fusion. The regression branch calculates the specific two-dimensional coordinate bounding box containing the sand hole defect and its confidence score percentage for different defect types based on the fused features; simultaneously, the segmentation branch outputs instance polygon mask boundaries that accurately fit the irregular physical contour of the sand hole. Finally, the detection coordinates are sent to the production line rejection mechanism via the communication interface.

[0048] This invention employs a self-attention mechanism that combines local window isolation with alternating window sliding. This mechanism effectively establishes a high-dimensional connection between microscopic defect signals and the surrounding macroscopic background while suppressing the overall computational scale. This detection scheme ensures both high sensitivity in capturing micron-level pinholes and the ability to analyze overall spatial distribution characteristics, thereby outputting comprehensive defect identification results with extremely high confidence and morphological accuracy.

[0049] Example 2: This embodiment describes an automatic detection method for sand hole defects in precious metal castings with high reflectivity and complex topological surfaces. Addressing the technical bottlenecks faced by precious metal castings in conventional industrial visual inspection, such as high-temperature diffuse reflection noise, dynamic specular reflection artifacts, and interference from complex process textures, this embodiment constructs a complete detection process through deep coupling of optical, geometric, and deep learning feature spaces.

[0050] As one embodiment of the present invention, refer to Figure 1 Flowchart of an automatic detection method for sand hole defects in precious metal casting based on deep learning, refer to Figure 2 Photometric decoupling and albedo modulation data flow diagram, see reference Figure 3 Schematic diagram of implicit neural field alignment and feature multiplication inhibition principle.

[0051] After the inspection station is activated, the multi-zone programmable dome light source enters a timing control mode. The light source sequentially illuminates each independent LED illumination area with a fixed angular velocity and step size, with each area representing a specific physical incident orientation. During each pulse illumination time, a high-resolution industrial camera performs fixed exposure parameter acquisition to obtain the original image under the current lighting conditions. This process, through hardware-level synchronous triggering, ensures that, with the camera's viewing angle absolutely fixed, only shadow shifts and highlight drifts caused by changes in the light source's orientation occur. After acquisition, these single-channel grayscale image sequences are stitched and stacked in memory in chronological order along the depth channel dimension to construct a multimodal temporal photometric tensor. Each pixel in this tensor carries a temporal photometric vector reflecting its local reflectivity, containing complete information about the optical response at that location under different illumination directions.

[0052] After obtaining the multimodal temporal photometric tensor, the optical denoising stage begins. Each pixel's spatial coordinates within the tensor are traversed, and the grayscale volatility of that point across all illumination channels is calculated. Since pinhole defects are typical geometric depressions, their interiors maintain a relatively stable low brightness state under illumination from all angles due to occlusion effects, exhibiting extremely low grayscale volatility. In contrast, the dynamic highlights generated by specular reflection rapidly shift across the surface as the light source position changes, causing drastic fluctuations in the grayscale values ​​of the affected pixels, resulting in extremely high grayscale volatility.

[0053] Based on this physical characteristic, a variance weight map is generated and used as a nonlinear gating factor to perform Hadamard product operations on the original photometric tensor. This step achieves pixel-level suppression of the dynamic specular reflection component, resulting in a modulated photometric tensor. Subsequently, the modulated tensor is integrally mean-valued along the channel dimension to eliminate the brightness gradient caused by directional lighting, generating a pseudo-diffuse reflectance map that reflects the true albedo of the material. Simultaneously, the spatial photometric partial derivatives are extracted from the modulated tensor, and a BRDF model based on micro-surface element theory is introduced as a physical constraint. By substituting the photometric partial derivatives into the model, the nonlinear photometric distortion caused by uneven surface roughness is calculated and compensated. Using joint optimization, a high-fidelity true topological morphology map characterizing the microscopic geometric features is output.

[0054] To mask the inherent textures of the product's manufacturing process, a pre-stored standard CAD design model is retrieved. Through spatial sampling and multilayer perceptron training, the geometric boundaries of the CAD model are transformed into a continuously differentiable implicit neural field representation. Subsequently, a normal vector lattice with high gradient features is extracted from the relative topological map generated from real-world photography. Abandoning the traditional discrete mesh matching approach, the 3D coordinates of this normal vector lattice are transformed using an initial pose matrix and directly input into the implicit neural field to obtain distance prediction values. A geometric loss function is constructed with the goal of minimizing the sum of squares of these distance prediction values. Utilizing the continuously differentiable nature of the implicit field, the rotation and translation parameters of the pose matrix are iteratively optimized using a gradient descent algorithm. This allows the real-world feature points to seamlessly "slide" within the differentiable space and fit onto the zero isosurface in the implicit field, thereby calculating a high-precision spatial affine transformation matrix. After registration, virtual ray tracing technology is used to emit probe rays into the implicit neural field according to the camera's intrinsic parameter model. By identifying the geometric properties of the intersection of light and the surface of the neural field, it is possible to accurately determine which pixel coordinates correspond to the designer's preset compliant micro-carving, lettering, or hollowed-out areas, and which correspond to smooth surfaces.

[0055] After determining the distribution of compliant process textures, the Euclidean distance fields from these textured regions to the smooth regions are further calculated. To eliminate noise caused by minor alignment deviations or defocusing at the edges of optical imaging, Gaussian smoothing filters are applied at the texture boundaries to generate Gaussian soft boundary distance field masks with continuously varying values ​​between zero and one.

[0056] Simultaneously, the generated pseudo-diffuse reflection map is input into the backbone extraction layer of the deep neural network to extract deep feature maps with semantic information. To achieve the fusion of geometric priors and deep features, the soft boundary mask tensor is spatially scaled and broadcast-copied in the channel dimension to ensure it is perfectly aligned with the deep feature map in terms of dimension. Subsequently, element-wise multiplication is performed in the feature activation layer, using low weight values ​​(close to 0) in the mask to forcibly clear the neuronal activation responses in the region where the compliant process texture is located. This action removes background interference signals from the mathematical level, so that the feature tensor output by the network retains only the abnormally random sand-hole defect features that violate the design intent, i.e., a pure feature tensor.

[0057] The purified feature tensors are input into the backend visual defect detection network. This network utilizes a windowed self-attention mechanism to calculate the correlation between pixel features in non-overlapping local regions, capturing subtle topological contrasts between the shadows inside the pinholes and the surrounding metal matrix. Subsequently, by performing window movement operations, cross-regional information interaction is achieved, thereby establishing a global defect topological connectivity model. With the assistance of a feature pyramid, the detection network fuses purified features at different scales and drives regression and segmentation branches. The regression branch outputs the center coordinates and bounding box of the pinhole defect, while the segmentation branch outputs a pixel-level instance mask that precisely matches the edge of the physical pore of the pinhole. Finally, the identified defect location information is converted into control commands and fed back to the production line execution end, completing the precise interception of defective products.

[0058] The method described in this embodiment establishes a complete mapping system from the original physical world to the digital feature space through physical abstraction of optical reflection laws and tensor-based reconstruction of prior CAD knowledge. By employing two key tensor suppression techniques (photometric-level variance suppression and feature-level mask suppression), high reflectivity noise and complex texture interference are gradually stripped away, solving the problems of missed detections and false alarms in automated quality inspection of precious metal castings.

[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An automatic detection method for sand hole defects in precious metal castings based on deep learning, characterized in that, include: Acquire temporal exposure image sequences of precious metal castings under illumination from light sources in different orientations, and stack them in the channel dimension to construct a multimodal temporal photometric tensor; The gray-level fluctuation rate of each spatial coordinate of the multimodal temporal photometric tensor in the channel dimension is calculated to generate a variance weight map. The variance weight map is used to perform nonlinear albedo modulation on the multimodal temporal photometric tensor. Based on the gray-level fluctuation rate, the light source interference components that cause specular reflection are filtered out, and the components are deconstructed in reverse into a pseudo diffuse reflection spectrum that characterizes the intrinsic properties of the material and a relative topological morphology spectrum that characterizes the micro-geometric features. The standard CAD design model of precious metal castings is transformed into an implicit neural field representation and spatially affine aligned with the relative topological map. The resulting rendering generates a topological constraint mask tensor that covers the compliant process texture. The deep feature map of the pseudo diffuse reflection map is extracted, and the topological constraint mask tensor is element-wise multiplied with the deep feature map to suppress it, so as to obtain a pure feature tensor that filters out compliant process textures; the pure feature tensor is input into the visual defect detection network for feature extraction, and the sand hole detection result is output.

2. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, To acquire a sequence of temporal exposure images of a precious metal casting under illumination from different directional light sources, and to construct a multimodal temporal photometric tensor, the process includes: controlling a dome light source, which serves as the light source, to sequentially illuminate individual illumination zones at a preset illumination angle; triggering an industrial camera to acquire fixed exposure parameters of the precious metal casting under illumination zone illumination, thereby obtaining single-channel grayscale images containing different physical shadow distributions; and stitching and fusing the pixel matrices of all single-channel grayscale images along the depth channel direction according to the illumination time sequence of the light source zones, outputting a multimodal temporal photometric tensor with a three-dimensional matrix structure, wherein the width and height of the tensor correspond to the physical pixel size of the image, and the channel depth of the tensor corresponds to the total number of illumination zones.

3. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, The calculation of grayscale volatility of each spatial coordinate in the multimodal temporal photometric tensor along the channel dimension and the generation of a variance weight map includes: traversing the pixel spatial coordinates in the multimodal temporal photometric tensor; extracting the pixel values ​​of a single pixel spatial coordinate at all channel depths and calculating the arithmetic mean and variance as grayscale volatility; inputting the grayscale volatility corresponding to all pixel spatial coordinates into a nonlinear activation function for normalization mapping; subtracting the normalized grayscale volatility from the value one to generate a variance weight map corresponding to each pixel with values ​​ranging from zero to one.

4. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, The method involves performing nonlinear albedo modulation on a multimodal temporal photometric tensor using a variance weighting map, and then deconstructing it into a pseudo diffuse reflectance map and a relative topological map. This includes: performing a pixel-by-pixel Hadamard product operation on the variance weighting map and the multimodal temporal photometric tensor to attenuate specular reflection interference components and outputting a modulated photometric tensor; calculating the integral mean of all pixel values ​​along the channel depth direction of the modulated photometric tensor to output a single-channel two-dimensional pixel array as a pseudo diffuse reflectance map; extracting the photometric partial derivatives between adjacent pixel coordinates in the modulated photometric tensor, calculating the three-dimensional surface normal vector for each spatial coordinate, mapping all three-dimensional surface normal vectors to a two-dimensional plane, and generating the relative topological map.

5. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, The process of converting a standard CAD design model of a precious metal casting into an implicit neural field representation includes: spatial sampling within the three-dimensional bounding box of the CAD design model to obtain a three-dimensional coordinate set; inputting the three-dimensional coordinate set into a multilayer perceptron network containing position encoding; training the multilayer perceptron network to fit the signed distance function value from each three-dimensional coordinate to the surface of the CAD design model; and using the weight parameters of the trained multilayer perceptron network as the implicit neural field representation.

6. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, Spatial affine alignment with a relative topological map is performed, and a topological constraint mask tensor covering compliant process textures is generated. This includes: extracting a feature point array whose normal vector change rate exceeds a preset gradient from the relative topological map; constructing a spatial pose optimization objective function, aiming to minimize the sum of squares of the signed distance function values ​​output after the feature point array is input to the implicit neural field representation, and iteratively optimizing the rotation and translation transformation matrix through a gradient descent algorithm to complete the spatial affine alignment; in the aligned state, emitting virtual rays along the virtual optical center viewpoint of the industrial camera towards the implicit neural field representation, and recording the coordinates of the intersection points between the rays and the implicit neural field surface; assigning spatial weights according to the structural properties to which the intersection coordinates belong, and projecting to generate a two-dimensional mask matrix as the topological constraint mask tensor.

7. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, The process involves extracting deep feature maps from a pseudo diffuse reflection map and performing element-wise multiplication suppression on the topological constraint mask tensor and the deep feature map. This includes: inputting the pseudo diffuse reflection map into a backbone feature extraction network to output a multi-channel deep feature map; scaling the spatial resolution of the topological constraint mask tensor to match the pixel size of the deep feature map using bilinear interpolation; copying and expanding the dimensionality of the scaled topological constraint mask tensor along the channel direction to make its channel count equal to the channel count of the deep feature map, generating a suppression tensor; performing a pixel-by-pixel multiplication operation on the suppression tensor and the deep feature map to clear the activation values ​​of neurons corresponding to compliant process texture regions, and outputting the clean feature tensor.

8. The method for automatic detection of sand hole defects in precious metal casting based on deep learning according to claim 1, characterized in that, The pure feature tensor is input into a visual defect detection network for feature extraction, and the trachoma detection result is output. This includes: dividing the pure feature tensor into non-overlapping local windows, calculating the self-attention relationship between pixel features within each local window to obtain local context-dependent features; performing a window movement operation to create a spatial overlap intersection between adjacent non-overlapping local windows, and calculating the cross-window attention relationship matrix between pixel feature vectors again within the overlap intersection to obtain global topological dependency features; inputting the local context-dependent features and the global topological dependency features into a feature pyramid fusion module to output trachoma detection results including trachoma bounding box coordinates, category confidence scores, and instance segmentation boundaries.