Intelligent visual inspection method and system for hardware stamping die production
Through intelligent visual inspection methods, the hardware stamping molds are detected, which solves the problem of difficult identification of small defects in traditional methods, and efficient and accurate detection of the mold status is achieved. Potential defects are discovered early, avoiding mold failure and production line damage.
Patent Information
- Application Number
- CN202510876835.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively detect minor defects in the production of hardware stamping molds, resulting in mold failure and production line damage. The traditional detection methods are inefficient, subjective and prone to missed inspections.
The intelligent visual detection method is adopted to obtain geometric alignment by acquiring the current state image of the mold and the defect-free reference image for geometric alignment, calculate the difference graph to obtain the significance distribution of potential defect areas, use a pyramid network for multi-scale differential feature hierarchical encoding, and introduce a feature receptive field anchoring mechanism, and finally use a lightweight classifier for decision-making.
It realizes sub-mm-level abnormality recognition of slight changes in hardware stamping molds, and early detection of defects has been discovered, which improves the sensitivity and accuracy of detection and reduces false detection and missed detection.
Smart Images

Figure CN120388019A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of hardware stamping die production, and more particularly, in the embodiments of this application, it relates to an intelligent vision detection method and system for hardware stamping die production. Background Art
[0002] In the field of hardware stamping die production, the precision and surface integrity of the die directly determine the forming quality of the stamped parts. Due to the die being under long-term high pressure, high-frequency impact, and material friction, micro-cracks, pitting, local deformation, and other minor defects are likely to occur on its surface. Although these defects are difficult to detect initially, they will gradually expand during the production process, ultimately leading to die failure and even damage to the production line equipment. Therefore, the early detection of minor defects is one of the core requirements of industrial quality control. Traditional detection mainly relies on manual visual inspection or contact flaw detection, which has problems such as low efficiency, strong subjectivity, and difficulty in quantification. In addition, some existing machine vision methods based on threshold segmentation or edge detection are prone to false detection and missed detection under complex texture backgrounds, light fluctuations, and interference from oil stains on the die surface.
[0003] Therefore, an optimized intelligent vision detection solution for hardware stamping die production is desired. Summary of the Invention
[0004] To solve the above technical problems, this application is proposed. The embodiments of this application provide an intelligent vision detection method and system for hardware stamping die production, which obtain the current state image of the hardware stamping die to be detected and the defect-free die state reference image and perform geometric alignment, then calculate the difference map to obtain the saliency distribution of potential defect regions, and use a pyramid network for hierarchical encoding of multi-scale difference features to separate normal wear artifacts and real defects. Further, a feature receptive field anchoring mechanism is introduced to dynamically adjust the attention weight according to the die geometry topology. Finally, a lightweight classifier is used to make a decision on the enhanced difference features to achieve sub-millimeter-level anomaly recognition. In this way, the sensitivity to minor changes in the hardware stamping die can be effectively improved, which helps to detect defects early.
[0005] According to one aspect of this application, an intelligent vision detection method for hardware stamping die production is provided, which includes:
[0006] Place the hardware stamping die to be detected on the detection station;
[0007] Collect the current state image of the hardware stamping die to be detected through a camera;
[0008] Extract the die state reference image marked as defect-free from the background database;
[0009] After geometrically aligning the reference die state image and the current die state image, calculate the pixel-level difference between the images to obtain a die state change significance map;
[0010] Extract local visual difference features from the die state change significance map to obtain a local visual semantic encoding feature map of die state differences;
[0011] Perform local visual semantic feature saliency on the local visual semantic encoding feature map of die state differences to obtain a local visual semantic enhanced encoding feature map of die state differences;
[0012] Based on the local visual semantic enhanced encoding feature map of die state differences, determine whether there are defects in the hardware stamping die to be detected.
[0013] According to another aspect of the present application, there is provided an intelligent vision detection system for hardware stamping die production, which includes:
[0014] A hardware stamping die to be detected placement module for placing the hardware stamping die to be detected at the detection station;
[0015] A current die state image acquisition module for acquiring the current die state image of the hardware stamping die to be detected through a camera;
[0016] A defect-free die state reference image extraction module for extracting a die state reference image marked as defect-free from the background database;
[0017] A die state change significant encoding module for geometrically aligning the reference die state image and the current die state image and then calculating the pixel-level difference between the images to obtain a die state change significance map;
[0018] A local visual difference feature extraction module for extracting local visual difference features from the die state change significance map to obtain a local visual semantic encoding feature map of die state differences;
[0019] A local visual semantic feature saliency module for performing local visual semantic feature saliency on the local visual semantic encoding feature map of die state differences to obtain a local visual semantic enhanced encoding feature map of die state differences;
[0020] A module for determining the existence of defects in the hardware stamping die to be detected, which is used to determine whether there are defects in the hardware stamping die to be detected based on the local visual semantic enhanced encoding feature map of die state differences.
[0021] Compared with the prior art, the present application provides an intelligent visual inspection method and system for the production of metal stamping dies, which obtains the current state image of the metal stamping die to be inspected and the defect-free die state reference image and performs geometric alignment, then calculates the difference map to obtain the significance distribution of the potential defect area, and uses a pyramid network to perform multi-scale difference feature hierarchical encoding to separate normal wear artifacts from real defects. Furthermore, a feature receptive field anchoring mechanism is introduced to dynamically adjust the attention weight according to the geometric topology of the die. Finally, a lightweight classifier is used to make decisions on the enhanced difference features to achieve sub-millimeter level anomaly recognition. In this way, the sensitivity to small changes in the metal stamping die can be effectively improved, which helps to detect defects early. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 This is a flow chart of an intelligent visual inspection method for metal stamping die production according to an embodiment of the present application.
[0024] Figure 2 Schematic diagram of data flow of an intelligent visual inspection method for metal stamping die production according to an embodiment of the present application.
[0025] Figure 3 This is a flowchart of the intelligent visual inspection method for metal stamping mold production according to an embodiment of the present application, which geometrically aligns the mold state reference image and the mold current state image and then calculates the pixel-level difference between the images to obtain a mold state change significance map.
[0026] Figure 4 This is a flowchart of performing local visual semantic feature saliency on the mold state difference local visual semantic coding feature map in the intelligent visual inspection method for metal stamping mold production according to an embodiment of the present application to obtain a mold state difference local visual semantic enhanced coding feature map.
[0027] Figure 5 This is a system block diagram of an intelligent visual inspection system for metal stamping die production according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0029] In view of the problems in the above background art, in the technical solution of the present application, an intelligent vision detection method for the production of hardware stamping dies is proposed. It can take the dynamic benchmark difference enhancement of the die state as the core idea, and establish a closed-loop framework for defect evolution detection by constructing a vision state baseline for the entire life cycle of the die. First, based on the cascaded preprocessing of non-local mean filtering and global contrast normalization, high-frequency noises such as oil film reflection and processing texture on the die surface are suppressed, and at the same time, the influence of illumination difference on the gray distribution is balanced, so that subsequent algorithms focus on structural deformation rather than environmental interference. In the geometric alignment stage, a joint optimization strategy of feature point robust matching and perspective transformation is adopted to eliminate the pixel-level spatial misalignment caused by the die replacement pose offset, and ensure that the reference image and the current image are in a strictly homologous coordinate system. Then, after calculating the difference map to obtain the saliency distribution of potential defect regions, a pyramid network is used to hierarchically encode multi-scale difference features, separating the low-contrast artifacts caused by normal wear from the high-frequency edge responses of real defects. Further, a feature receptive field anchoring mechanism is introduced to dynamically adjust the attention weights of the convolution kernels according to the die geometry topology, strengthening the local semantic expression of targets such as micro-cracks and pitting, and at the same time suppressing background noises irrelevant to the die function. Finally, a lightweight classifier is used to make a decision on the enhanced difference features, realizing the reliable identification of sub-millimeter-level anomalies in the defect germination stage. Through the collaborative optimization of benchmark comparison and multi-scale feature decoupling, this solution highlights the tiny changes on the die surface (which may represent newly generated defects or the expansion of defects), breaking through the bottleneck of insufficient sensitivity of traditional single-frame detection to tiny defects, and helping to more effectively identify and detect abnormal defects of hardware stamping dies.
[0030] In view of the above technical problems, the present application proposes an intelligent vision detection method for the production of hardware stamping dies. Figure 1 The flowchart of the intelligent vision detection method for the production of hardware stamping dies according to an embodiment of the present application. Figure 2 The schematic diagram of data flow of the intelligent vision detection method for the production of hardware stamping dies according to an embodiment of the present application. As Figure 1 and Figure 2As shown, an intelligent vision detection method for the production of hardware stamping dies according to an embodiment of the present application includes: S110, placing the hardware stamping die to be detected on a detection station; S120, collecting an image of the current state of the die of the hardware stamping die to be detected through a camera; S130, extracting a die state reference image marked as defect-free from a background database; S140, performing geometric alignment on the die state reference image and the image of the current state of the die, and then calculating the pixel-level difference between the images to obtain a die state change significance map; S150, extracting local visual difference features from the die state change significance map to obtain a die state difference local visual semantic coding feature map; S160, performing local visual semantic feature saliency on the die state difference local visual semantic coding feature map to obtain a die state difference local visual semantic enhanced coding feature map; S170, determining whether there are defects in the hardware stamping die to be detected based on the die state difference local visual semantic enhanced coding feature map.
[0031] In the above intelligent vision detection method for the production of hardware stamping dies, in step S110, the hardware stamping die to be detected is placed on a detection station. It should be understood that in order to ensure the accuracy and consistency in the subsequent image acquisition process, precise positioning and fixation of the hardware stamping die to be detected are required. This means that when placing the die on the detection station, special fixtures can be designed or precise robotic arms can be used to avoid causing position deviations. Specifically, automated equipment with high repeat positioning accuracy can be considered, and its motion trajectory can be controlled through programming to ensure that the die can be accurately placed in the same position every time. In addition, considering the possible differences in different models and sizes of dies, adjustable clamping devices should also be configured to adapt to the structural characteristics of various different dies. At the same time, the interference of external light sources should also be minimized to prevent unnecessary shadows or reflections from appearing in the captured images. For this purpose, light-shielding plates can be set around the detection station, and professional lighting equipment can be used to provide uniform and sufficient light irradiation to create good conditions for the camera to capture high-quality die images. Furthermore, in actual operation, attention also needs to be paid to how to maximize the exposure of the key detection parts of the die. This usually involves adjusting the posture of the die so that the areas most likely to have defects are fully presented within the camera's field of view. For some dies with complex shapes or multiple important surfaces, the angle may need to be changed multiple times for a full-range scan. At this time, in addition to relying on the aforementioned automated equipment, computer-aided design (CAD) software can also be combined to pre-plan the best observation perspective in advance, and the relevant information can be fed back to the control system to guide the specific placement actions.
[0032] In the above intelligent vision inspection method for the production of metal stamping dies, in step S120, a current state image of the metal stamping die to be inspected is collected through a camera. It should be understood that in terms of hardware selection, considering that metal stamping dies usually have complex geometric shapes and may have subtle defect features, an industrial-grade camera with high resolution and low noise characteristics needs to be selected. Such cameras can capture minute changes on the die surface, which is crucial for identifying defects such as cracks or pitting at the sub-millimeter level. In addition, according to the requirements of the actual application scenario, different types of sensor technologies can also be selected, such as CMOS (Complementary Metal Oxide Semiconductor) or CCD (Charge Coupled Device). The former has more advantages in cost-effectiveness and integration, while the latter is known for its excellent image quality and stability. At the same time, in order to adapt to different detection distances and viewing angle requirements, a suitable lens system needs to be equipped, and it is also necessary to consider whether to add filters to reduce the impact of ambient light on the imaging quality. For example, using a polarizer can effectively reduce the reflection phenomenon and improve the image contrast. Secondly, when setting the camera parameters, the specific characteristics of the object to be photographed must be fully considered. Since the working environment of metal stamping dies is complex and variable, and may be affected by factors such as oil residue, processing texture, and uneven illumination, it is necessary to finely adjust the various parameters of the camera. For example, the exposure time should be set to ensure sufficient brightness without overexposure resulting in loss of details; the gain value adjustment needs to enhance the details in the dark areas while avoiding introducing too much noise. In addition, white balance is also one of the important factors affecting the final imaging effect. Especially in the case of mixed lighting from multiple light sources, the correct white balance setting can help restore the true color information of the die, thus more accurately reflecting its surface condition. Next, during the installation of the camera, it is particularly important to ensure its stable position and maintain an appropriate distance from the die to be inspected. This not only helps to obtain clear and sharp images but also avoids image blurring caused by vibration or displacement. Generally speaking, the camera can be fixed in a stable position through a precise mechanical device, and a three-axis adjustment mechanism can be used for fine-tuning so that the center of the lens can accurately align with the key detection area of the die. If different parts of the die need to be scanned, an electric slide table or a rotating platform can be used to automatically move the camera along a predetermined path to achieve full coverage without dead angles. Furthermore, for dies made of some special materials, such as metal parts with high reflectivity, a diffuser plate sometimes needs to be used in combination to make the light shine on the workpiece surface more softly and evenly, reducing the glare effect caused by direct reflection. It should be noted that in actual operation, customized solutions are often required for different types and sizes of metal stamping dies. This is because the differences between different dies are relatively large, including both structural differences and material differences, which will affect the imaging effect.Therefore, before starting each new inspection task, it is necessary to carefully evaluate the characteristics of the mold and formulate corresponding shooting plans accordingly. This includes, but is not limited to, determining the most suitable shooting angle, selecting an appropriate focal length range, and planning a reasonable scanning route, etc. Only in this way can it be ensured that each image collected records all relevant information of the mold as comprehensively as possible, providing comprehensive data support for subsequent defect detection.
[0033] In the above intelligent vision inspection method for the production of hardware stamping dies, in step S130, a die state reference image marked as defect-free is extracted from the background database. It should be understood that extracting a die state reference image marked as defect-free from the background database not only provides a necessary reference standard for subsequent image comparison and analysis, but also has decisive significance for accurately identifying subtle changes on the die surface. First, considering the diverse requirements in practical applications, it is also necessary to establish respective reference image libraries for different models and specifications of dies. The advantage of doing this is that when a specific die enters the inspection process, the most suitable reference sample can be quickly located from the corresponding reference image library, thereby improving the matching accuracy and efficiency. Next, in the process of extracting a die state reference image marked as defect-free from the background database, the architecture design of the data management system is particularly important. An ideal database should have efficient data retrieval capabilities and strong scalability to facilitate the storage and processing of a large amount of high-resolution image data. For this purpose, a distributed file system or a cloud storage solution can be used to build the underlying storage platform to ensure fast access even in the face of massive data. At the same time, for easy management and query, each reference image also needs to be carefully annotated, including but not limited to die number, shooting date, angle information, and quality assessment results, etc. These metadata can not only help quickly find the required image, but also provide rich context support for subsequent automated processing. Further, when specifically performing the extraction operation, first, according to the specific identification information (such as model, serial number, etc.) of the die to be inspected, the corresponding reference image set is located in the database. This process usually involves the application of complex index structures and search algorithms. For example, data structures such as B-trees or hash tables can help accelerate the search process, while full-text search engine technology can play an important role in fuzzy matching scenarios. Once the target image set is determined, one or more reference images most suitable for the current inspection task need to be further selected. This is usually carried out based on some predefined rules. For example, the image taken most recently and with the highest quality score is preferentially selected as the main reference; if there are multi-view images, the coverage range in each direction needs to be comprehensively considered to ensure that all important areas can be fully compared. It should be noted that in practical applications, due to the influence of various unforeseen factors, it may happen that no exact match can be found in the reference image library. At this time, certain compensation measures need to be taken to address this challenge. A common approach is to perform appropriate transformation operations on the existing reference images, such as rotation, scaling, or translation, etc., to make it as close as possible to the actual state of the die to be inspected. In addition, to ensure that the extracted reference images are always in the latest and optimal state, regularly updating the database is also an essential part.This means that whenever a new defect-free mold image is added or the quality of an existing image is significantly improved, it is necessary to promptly incorporate it into the reference image library and synchronously update the relevant metadata records. This not only helps improve the accuracy of the detection system but also enhances its adaptability and flexibility. At the same time, a sound data backup mechanism should be established to prevent the loss of important data in case of unexpected situations. By means of off-site disaster recovery backup or incremental backup, etc., the security and reliability of the data can be maximally ensured.
[0034] Figure 3 A flowchart for calculating the pixel-level difference between the mold state reference image and the current mold state image after geometric alignment to obtain the significance map of mold state changes in the intelligent vision detection method for hardware stamping mold production according to an embodiment of the present application. As Figure 3 shown, in the embodiment of the present application, step S140, calculating the pixel-level difference between the mold state reference image and the current mold state image after geometric alignment to obtain the significance map of mold state changes, includes: S141, performing image preprocessing on the mold state reference image and the current mold state image to obtain the preprocessed mold state reference image and the preprocessed current mold state image; S142, performing geometric alignment on the preprocessed mold state reference image and the preprocessed current mold state image to align the preprocessed current mold state image to the coordinate system of the preprocessed mold state reference image to obtain the aligned current mold state image; S143, calculating the pixel-level difference between the aligned current mold state image and the preprocessed mold state reference image to obtain the significance map of mold state changes.
[0035] In the embodiments of the present application, step S141, which preprocesses the die state reference image and the current die state image to obtain the preprocessed die state reference image and the preprocessed current die state image, includes: performing non-local mean filtering on the die state reference image and the current die state image and then performing global contrast normalization to obtain the preprocessed die state reference image and the preprocessed current die state image. It should be understood that in the visual inspection scenario of metal stamping dies, there are often problems such as oil film residue, processing texture, and uneven illumination in the production environment on the die surface. These interferences will cause abnormal gray-scale distribution and significant high-frequency noise in the image, directly affecting the accuracy of subsequent geometric alignment and differential operations. Traditional preprocessing methods (such as Gaussian filtering or histogram equalization) are difficult to balance noise suppression and detail preservation, and are prone to edge blurring or contrast distortion of small defects. Therefore, in the technical solution of the present application, the die state reference image and the current die state image are preprocessed to obtain the preprocessed die state reference image and the preprocessed current die state image. In particular, in a specific example of the present application, a cascaded preprocessing strategy of non-local mean filtering and global contrast normalization can be adopted: First, through non-local mean filtering, the similarity of non-adjacent regions in the image is used for weighted denoising. While eliminating high-frequency interferences such as oil stain reflection and random noise, the local structural features of small defects such as micro-cracks and pitting on the die surface are retained; Subsequently, through global contrast normalization, the illumination robustness of the filtered image is enhanced, and the gray-scale offset caused by environmental light fluctuations or camera exposure differences is eliminated, making the brightness distributions of the reference image and the current image tend to be consistent. The core purpose of this preprocessing process is to reduce the interference of complex backgrounds and dynamic environments on defect detection and provide clearly structured and illumination-robust input data for subsequent precise registration and difference analysis. In this way, the signal-to-noise ratio of the die state reference image and the current die state image can be improved, enabling the uniform artifacts of normal wear and the gradient mutation characteristics of abnormal defects to be separable in the subsequent die state change significance map, thus laying a foundation for defect detection of metal stamping dies.
[0036] Specifically, in step S142, geometric alignment is performed on the preprocessed die state reference image and the preprocessed current die state image to align the preprocessed current die state image to the coordinate system of the preprocessed die state reference image, so as to obtain an aligned current die state image. It should be understood that during the intelligent detection process of the metal stamping die, each time the die is replaced and installed at the detection station, there may be pose offsets due to mechanical positioning tolerances or manual operation deviations, resulting in geometric deformations such as rotation, translation, or scale scaling between the current state image and the reference image. If direct differential operations are performed, such spatial misalignments will introduce a large amount of pixel-level difference noise unrelated to defects (such as die edge misalignment artifacts and texture region phase offsets), seriously interfering with the extraction of significant features of micro defects. Therefore, in this solution, the preprocessed reference image and the current image are forced to be unified in the spatial coordinate system through geometric alignment. The core purpose is to eliminate the non-defective geometric differences caused by the physical pose changes of the die, ensuring that subsequent differential operations only reflect the real deformations or damage evolutions on the die surface. In specific implementation, a strategy of joint optimization of feature point robust matching and perspective transformation is adopted: First, SIFT feature points of the key structures of the die (such as positioning holes and intersection points of ridge lines) are extracted, and the mismatched point pairs are removed through the RANSAC algorithm. Then, based on the optimal matching point set, the perspective transformation matrix is solved to accurately map the current image to the coordinate system of the reference image. This process effectively overcomes the interference of repeated textures on the die surface (such as grid-like machining marks) to feature matching, and at the same time is compatible with the slight deformations caused by long-term use of the die. After geometric alignment, the two images reach sub-pixel-level alignment accuracy in the spatial dimension, making the uniformly distributed features of normal wear and the locally mutated features of abnormal defects significantly separable in the difference map.
[0037] Specifically, in step S143, calculate the pixel-level difference between the current state image of the alignment mold and the preprocessed mold state reference image to obtain the mold state change significance map. It should be understood that since the visual features of normal wear and abnormal defects of the mold during the production process are highly similar in the image, it is difficult for traditional methods to directly distinguish between the two from a single shot. Therefore, by further calculating the pixel-level difference between the current state image of the geometrically aligned mold and the preprocessed reference image, it is possible to effectively strip static background information such as the inherent texture and processing marks of the mold, and lock the detection focus on the local dynamic changes caused by the generation or expansion of defects. The purpose of this operation is to establish a "relative detection" logic for defects through a reference comparison mechanism, using the historical state of the mold itself as a reference system, and converting sub-pixel-level anomalies (such as the extension of microcracks and the spread of pitting corrosion) originally submerged in the complex background into significant difference regions, thereby overcoming the inherent defect of insufficient sensitivity of single-frame detection to small targets. Judging from the execution effect, the mold state change significance map generated by this step not only amplifies the signal-to-noise ratio difference between the defect features and the background noise by highlighting the abnormal change amount in the spatial domain, but also provides input data with clear physical meaning for subsequent multi-scale feature decoupling. Especially after geometric alignment and preprocessing to eliminate pose offset and lighting interference, the differential operation can accurately capture real defect features such as surface structure deformation and material loss of the mold, rather than artifacts caused by environmental variables. This enables the subsequent feature extraction network to more efficiently separate the low-frequency pseudo-changes of normal wear from the high-frequency edge responses unique to defects, and finally achieve reliable identification of subtle anomalies in the defect germination stage.
[0038] In an embodiment of the present application, the step S150 of extracting local visual difference features from the mold state change saliency map to obtain a local visual semantic coding feature map of the mold state difference includes: passing the mold state change saliency map through a local visual difference feature extractor based on a pyramid network to obtain the local visual semantic coding feature map of the mold state difference. It should be understood that although the mold state change saliency map generated through differential operations can highlight the change regions on the mold surface, it is essentially a pixel-level gray-scale difference distribution, mixed with multi-scale noise and real defect features at the same time. Since the morphology of minute defects (such as micro-cracks and pitting) on the mold surface usually presents cross-scale feature coupling - for example, the extension direction of a crack shows a linear trend at the macroscopic scale, while the sub-pixel-level fracture details at its edge need to be captured at high resolution, it is difficult for traditional single-scale convolutional neural networks to capture such cross-level spatial correlations simultaneously. In addition, low-contrast artifacts caused by normal wear are prone to be confused with the high-frequency responses of real defects at a single scale, and physical meaning separation needs to be achieved through multi-level feature decoupling. Therefore, in a specific example of the present application, a local visual difference feature extractor based on a pyramid network is used to extract local visual difference features from the mold state change saliency map to obtain a local visual semantic coding feature map of the mold state difference. The purpose of this step is to utilize the unique multi-scale feature fusion mechanism of the pyramid network, and through hierarchical downsampling and cross-layer feature aggregation, construct a hierarchical semantic expression from local subtle differences to global structural associations. The pyramid network gradually extracts underlying features such as gradients, edges, and textures in the saliency map through receptive fields of different scales, and integrates the spatial distribution rules at the high-level semantic level, thereby converting pixel-level gray-scale fluctuations into defect feature descriptions with engineering interpretability (such as the topological continuity of cracks and the discrete distribution pattern of pitting). Judging from the execution effect, this multi-scale coding strategy can effectively distinguish the essential differences between normal wear and abnormal defects: at the low-resolution feature layer, the network focuses on the overall trend of changes on the mold surface and suppresses isolated noise caused by oil residue or uneven illumination; at the high-resolution feature layer, it enhances the sensitivity to microscopic features of defects such as sub-pixel-level edge fractures and material loss. Through the cross-scale feature interaction of the pyramid structure, the system can adaptively balance the relationship between local details and global context, for example, recognizing the sharp edge response presented by a micro-crack locally and its consistency in the extension direction at the macroscopic scale. This hierarchical coding mechanism provides a feature basis with spatial semantic associations for subsequent feature saliency, enabling the attention mechanism based on receptive field anchoring to accurately locate the energy aggregation regions of defect features, thereby improving the classification confidence of minute defects and reducing the false change misjudgment rate caused by normal wear to an industrially acceptable level.
[0039] Figure 4It is a flowchart for local visual semantic feature saliency of the local visual semantic coding feature map of the die state difference in the intelligent visual inspection method for hardware stamping die production according to an embodiment of the present application to obtain a local visual semantic enhanced coding feature map of the die state difference. As Figure 4As shown, in the embodiment of the present application, step S160, which performs local visual semantic feature saliency on the local visual semantic encoding feature map of the mold state difference to obtain a local visual semantic enhanced encoding feature map of the mold state difference, includes: S161, performing feature dissociation on the local visual semantic encoding feature map of the mold state difference and then performing information compression processing to obtain a local visual semantic feature vector of the mold state difference to be enhanced by distillation; S162, based on the spatial distribution of the local visual semantic feature vector of the mold state difference to be enhanced by distillation, screening out a set of pixel-level local visual semantic feature vectors of the mold state difference within the local receptive field from the set of pixel-level local visual semantic feature vectors of the mold state difference; S163, based on the set of pixel-level local visual semantic feature vectors of the mold state difference within the local receptive field, enhancing the saliency of the local visual semantic feature vector of the mold state difference to be enhanced to obtain an enhanced local visual semantic feature vector of the mold state difference; S164, aggregating information of multiple enhanced local visual semantic feature vectors of the mold state difference to obtain a local visual semantic enhanced encoding feature map of the mold state difference as the enhanced local visual semantic feature of the mold state difference. It should be understood that although the pyramid network has initially separated the response patterns of normal wear and real defects through multi-scale feature encoding, the local semantic features of minute defects (such as sub-millimeter cracks) on the mold surface may still be diluted by interference from complex background noise or pseudo-changes in adjacent regions. Since conventional convolution operations use a static receptive field of a fixed size, they cannot adaptively adjust the feature aggregation range according to the dynamic changes in the defect morphology (such as the discrete distribution of pitting or the topological extension of cracks), resulting in ambiguity in the semantic expression of key defect features in the spatial dimension. In addition, the residual oil stain on the mold surface and the non-uniform distribution of machining textures in local areas will introduce high-frequency noise unrelated to defects. If a classification decision is directly made based on the original feature map, it is likely to cause weakening of the defect edge response or activation of false features. Therefore, in the technical solution of the present application, further local visual semantic feature saliency is performed on the local visual semantic encoding feature map of the mold state difference to obtain a local visual semantic enhanced encoding feature map of the mold state difference. The process of local visual semantic feature saliency processing aims to break through the rigid constraint of the receptive field of the traditional network through an effective anchoring mechanism of the feature receptive field, dynamically adjust the aggregation scale of the context information of the mold state difference according to the feature distribution characteristics at each pixel position, so as to construct a semantic enhanced expression that matches the physical characteristics of the defect in the feature space. By decoupling the local visual semantic encoding feature map of the mold state difference into pixel-level initial feature vectors and implementing information distillation, the network first extracts the core feature representation at each position, and then analyzes the spatial structure characteristics of the compressed feature vectors to dynamically predict the most suitable local receptive field size.This adaptive mechanism enables the network to expand the receptive field to capture continuous edge features according to the linear extension characteristics of microcracks, while shrinking the receptive field for discrete pitting to focus on the local mutations of microscopic material loss. After screening the feature vectors within the local receptive field, by introducing conformal representational constraints (such as the commutation optimization of the volume space and boundary representation), the network further strengthens the spatial consistency expression of defect features, ensuring that the pixel-level feature vectors in the enhanced local visual semantic enhanced encoding feature map of the mold state difference not only retain the high-frequency sharpness of the defect edge but also integrate the context association of its macroscopic topological structure. For example, in the oil stain interference area, the network suppresses the cross-region noise propagation by shrinking the receptive field; in the real defect area, it enhances the semantic continuity of the crack extension direction by expanding the receptive field. For microcracks of 0.1 mm level, the dynamic expansion of the receptive field can associate the discrete pixel points at the fracture edge to form a coherent linear feature response; for low-contrast artifacts caused by normal wear, the contraction of the local receptive field effectively blocks its false association with the surrounding noise, thus achieving the precise separation of the two in the feature space.
[0040] In the embodiment of the present application, in step S161, performing feature dissociation on the local visual semantic encoding feature map of the mold state difference and then performing information compression processing to obtain the local visual semantic feature vector of the mold state difference to be enhanced by distillation, includes: S1611, performing feature dissociation on the local visual semantic encoding feature map of the mold state difference along the channel dimension to obtain a set of local visual semantic pixel-level feature vectors of the mold state difference; S1612, extracting the local visual semantic pixel-level feature vector at the (i, j) pixel position in the local visual semantic encoding feature map of the mold state difference from the set of local visual semantic pixel-level feature vectors of the mold state difference as the local visual semantic feature vector of the mold state difference to be enhanced; S1613, performing information compression on the local visual semantic feature vector of the mold state difference to be enhanced to obtain the local visual semantic feature vector of the mold state difference to be enhanced by distillation.
[0041] Specifically, step S1611, performing feature dissociation on the local visual semantic encoding feature map of the mold state difference along the channel dimension to obtain a set of local visual semantic pixel-level feature vectors of the mold state difference, is represented by the mold state feature dissociation formula as:
[0042]
[0043] ;
[0044] Among them, is the local visual semantic encoding feature map of the mold state difference, the set of real numbers, are respectively The height, width and number of channels, For Perform feature dissociation, and are the first and second local visual semantic pixel-level feature vectors of mold state difference. and The pixel-level feature vector of the mold state difference at the pixel location represents the local visual semantics. It is understandable that because the physical characteristics of mold surface defects exhibit non-independent distributions across different channel representations, traditional parallel channel processing methods can easily dilute critical defect signals with redundant information from adjacent channels. By decoupling along the channel dimension, the multidimensional feature vector at each pixel location can be deconstructed into independent single-channel feature subsets, eliminating the strong correlation constraints between channels caused by the shared convolution kernel mechanism. This approach aims to construct pixel-level feature units with independent semantic expression capabilities, laying the foundation for subsequent refined feature processing. Decoupling reconstructs the topology of the feature space, ensuring that the feature vector at each pixel location retains only its inherent properties, avoiding interference from inter-channel statistical dependencies on defect feature determination. By decoupling the channels, subsequent information compression and receptive field anchoring mechanisms can more accurately focus on localized physical change patterns, such as the gradient abrupt changes in microcracks or the local contrast differences in pitting. This decoupling strategy is essentially a process of disentangling the feature space, releasing previously masked defect-specific signals by separating redundant correlations between channels. In this way, feature decoupling significantly improves the distinguishability of defect features. Furthermore, by eliminating inter-channel interference, the network can more effectively capture the local topological features of defects, such as the linear extension direction of cracks or the discrete distribution pattern of pitting. Finally, the decoupled collection of pixel-level feature vectors of local visual semantics of mold state differences provides a clearly structured input for subsequent information compression, enabling feature enhancement based on the commutativity of volumetric space-boundary representations. This channel decoupling mechanism essentially constructs a semantic isolation zone for defect features, providing a high-purity feature substrate for multi-scale feature fusion and dynamic receptive field adjustment.
[0045] Specifically, in step S1612, the mold state difference local visual semantic pixel-level feature vector at the (i, j)th pixel position in the mold state difference local visual semantic encoding feature map is extracted from the set of the mold state difference local visual semantic pixel-level feature vectors as the mold state difference local visual semantic feature vector to be enhanced. The mold state difference local visual semantic feature extraction formula to be enhanced is expressed as:
[0046]
[0047] in, It is the local visual semantic feature vector of the mold state difference to be enhanced. It should be understood that since the germination and expansion of defects such as microcracks and pitting on the mold surface usually exhibit localized characteristics, the key discriminant information often focuses on the gradient mutation or texture distortion area in the neighborhood of specific pixels. Through the local feature anchoring mechanism centered on the pixel (i,j), the network can accurately lock the analysis focus on the defect energy aggregation area, avoiding the dilution of defect signals by noise in non-related areas during the global feature fusion process. This essentially constructs a feature response unit centered on the local topological structure of the defect, enabling the network to independently analyze the spatial heterogeneity of the deformation of the mold surface pixel by pixel. Through the localized mapping of the feature space, refined modeling of defect semantics can be achieved. Among them, each local visual semantic pixel-level feature vector of the mold state difference serves as an independent decision-making unit, carrying key geometric information such as the gradient direction and contrast change within its neighborhood. This design enables the network to dynamically perceive the local morphological characteristics of the defect, such as the linear edge response of a 0.1mm-level microcrack or the discrete distribution pattern of pitting. Through local operations anchored by the central pixel, the network further establishes an association channel between pixel-level features and the global context. For example, it associates discrete pixel points on the fracture edge of the crack through the receptive field dynamic adjustment mechanism to form a coherent defect semantic expression. In this way, the sensitivity of the network to local anomalies is enhanced, and by independently analyzing the physical change patterns in each pixel neighborhood, it effectively distinguishes the uniform artifacts of normal wear from the high-frequency edge responses unique to defects.
[0048] Specifically, in step S1613, the local visual semantic feature vector of the mold state difference to be enhanced is compressed to obtain the local visual semantic feature vector of the mold state difference of the to-be-enhanced distilled mold, which is represented by the to-be-enhanced distilled mold state feature information compression formula:
[0049]
[0050] Among them, is the first norm of the vector, It is the local visual semantic feature vector of the state difference of the die to be enhanced by distillation. It should be understood that since the multi-scale feature encoding process of the pyramid network will introduce semantic responses at different levels, the features at a single pixel position contain cross-scale and cross-channel repetitive information. This redundancy will not only increase the complexity of subsequent calculations but also obscure the local specificity of defect features. Through information compression operations, the network can refine the features, removing redundant components irrelevant to defect discrimination while retaining the core features representing the essential attributes of defects. Through the compression process, the network forces the features to discard redundant statistical correlations and instead focus on the physical essential features of defects, such as the consistency of the gradient direction of microcracks or the local contrast mutation of pitting corrosion. This compression is essentially a process of mapping the high-dimensional feature manifold to a low-dimensional manifold, enabling the feature vector to have stronger generalization ability and anti-interference ability while retaining key discriminant information, providing an input base with clear physical meaning for the subsequent receptive field adaptation mechanism. In this way, information compression significantly reduces the computational load of the model, improves the real-time detection efficiency by removing redundant features, and forms a more compact defect response pattern in the spatial dimension, enabling the local receptive field adjustment mechanism to more accurately capture the topological structure features of defects.
[0051] In the embodiment of the present application, in step S162, based on the spatial distribution of the local visual semantic feature vector of the state difference of the die to be enhanced by distillation, a set of pixel-level local visual semantic feature vectors of the state difference of the die within the local receptive field is selected from the set of pixel-level local visual semantic feature vectors of the state difference of the die, including: S1621, based on the characteristic distribution spatial structure characteristics of the local visual semantic feature vector of the state difference of the die to be enhanced by distillation, determining the size of the characteristic receptive field of the local visual semantic feature vector of the state difference of the die to be enhanced; S1622, based on the size of the characteristic receptive field, selecting the set of pixel-level local visual semantic feature vectors of the state difference of the die within the local receptive field from the set of pixel-level local visual semantic feature vectors of the state difference of the die.
[0052] Specifically, in step S1621, based on the characteristic distribution spatial structure characteristics of the local visual semantic feature vector of the state difference of the die to be enhanced by distillation, the size of the characteristic receptive field of the local visual semantic feature vector of the state difference of the die to be enhanced is determined, which is represented by the characteristic receptive field size determination formula:
[0053]
[0054] where is the value of the logarithmic function with base 2, is The size of the characteristic receptive field. It should be understood that due to the diverse spatial topological structures presented by the defect morphologies such as the extension direction of microcracks and the discrete distribution of pitting corrosion in the local area, it is difficult for a fixed-size receptive field to simultaneously meet the requirements of capturing the global continuity and local mutability of defects. By dynamically adjusting the receptive field size, the network can adaptively match the optimal context perception range according to the spatial structure characteristics such as the distribution density and gradient change rate of defect features. By constructing a feature interpretation mechanism with semantic adaptability, the spatial structure characteristics of the feature distribution (such as gradient direction consistency, contrast mutation intensity, texture continuity, etc.) are mapped into the regulation parameters of the receptive field size, enabling the network to dynamically optimize the feature aggregation range according to the physical characteristics of the defects. For example, for microcracks presenting linear extension features, the network expands the receptive field to associate the discrete pixel points at the fracture edges; while for discretely distributed pitting corrosion defects, it contracts the receptive field to suppress the interference of background noise. This dynamic adjustment mechanism breaks the prior constraint of the receptive field size in traditional networks, enabling the feature extraction process to be deeply coupled with the physical characteristics of the defects. In this way, the spatial analysis ability of defect features is significantly improved, enabling the network to adaptively select the optimal context information range according to the defect morphology. In particular, in the technical solution of this application, the receptive field size is no longer a preset prior knowledge, but is adaptively determined by the feature content, making the selection of the receptive field more semantic and targeted, more effectively utilizing the context information, and improving the accuracy and robustness of saliency detection.
[0055] Specifically, in step S1622, based on the size of the characteristic receptive field, the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field is screened out from the set of pixel-level die state difference local visual semantic feature vectors of the image die state difference. It is expressed by the pixel-level die state difference feature screening formula within the local receptive field as:
[0056]
[0057] Wherein, is the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field, 、 、 and are respectively the 、 、 and Local visual semantic feature vectors of pixel-level die state differences within the local receptive field of pixel positions. It should be understood that through a dynamic screening mechanism based on the receptive field size, the network can accurately delimit the context-aware range matching the defect morphology, ensuring that only the neighborhood features closely related to defect discrimination are incorporated into the enhancement calculation. By binding the receptive field size to the local structural characteristics of the defect, the network can dynamically adjust the coverage range of feature retrieval: for defects presenting long-range continuous features (such as cracks), the screening range is extended to multi-scale neighborhoods to capture edge continuity; for discretely distributed defects (such as pitting), the screening range is contracted to the local core area to suppress background noise interference. This mechanism enables the feature aggregation process to be deeply coupled with the physical characteristics of the defect, breaking through the limitations of the static setting of receptive fields in traditional networks. In this way, the spatial resolution accuracy of defect features is significantly improved, and through dynamic screening, it is ensured that the feature vectors participating in the enhancement precisely cover the morphological key areas of the defect. This dynamic screening mechanism essentially constructs a context-aware filter for defect features, enabling the network to autonomously optimize the spatial resolution and semantic purity of feature enhancement without manual intervention.
[0058] In an embodiment of the present application, in step S163, based on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field, significant enhancement is performed on the local visual semantic feature vector to be enhanced to obtain an enhanced local visual semantic feature vector of die state differences, including: S1631, performing significant weighted fusion on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field to obtain a local visual semantic fusion feature vector of pixel-level die state differences within the local receptive field; S1632, based on the local visual semantic fusion feature vector of pixel-level die state differences within the local receptive field, performing feature enhancement on the local visual semantic feature vector to be enhanced to obtain the enhanced local visual semantic feature vector of die state differences, where the enhanced local visual semantic feature vector of die state differences is the channel feature vector at the (i,j) pixel position of the local visual semantic enhanced encoding feature map of die state differences.
[0059] Specifically, in step S163, based on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field, significant enhancement is performed on the local visual semantic feature vector to be enhanced to obtain an enhanced local visual semantic feature vector of die state differences, which is expressed by the formula:
[0060]
[0061]
[0062] where and is a trainable weighting hyperparameter, is a significantly enhanced weight factor of is a function, is vector multiplication, is the enhanced die state difference local visual semantic feature vector at the th pixel position in the corresponding die state difference local visual semantic enhanced coding feature map. It should be understood that by constructing a non-linear weighted fusion mechanism based on the spatial structure characteristics of the local receptive field, the network can dynamically adjust the contribution weights of different feature vectors in the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field, making the energy concentration of defect features significantly higher than background noise. Through significant weighted fusion, the network not only aggregates underlying features such as gradient direction and contrast change within the local receptive field, but also semantically aligns the macroscopic topological structure (such as crack extension direction) of the defect with the microscopic deformation features (such as pitting edge fracture) through volume space-boundary representation. This fusion mechanism enables the enhanced die state difference local visual semantic feature vector to retain both the local specificity of the defect and incorporate multi-scale context information, providing a more discriminative feature basis for subsequent classification decisions. This significant enhancement mechanism essentially constructs a cognitive enhancer for defect features, enabling the network to autonomously complete the cross-modal mapping from raw pixels to defect semantics without the need for manual annotation.
[0063] Specifically, first, step S1631 is executed, that is:
[0064]
[0065] wherein, is the pixel-level die state difference local visual semantic fusion feature vector within the local receptive field.
[0066] After that, the numerical values of the trainable weighting hyperparameters of and are determined. Specifically as follows: In particular, first, the ranges of and are limited. For the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field corresponding to the die state difference local visual semantic feature vector to be enhanced by distillation in order to enhance the discriminative ability of the target feature and suppress the interference of irrelevant background noise, it is desired that the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field It can have morphological consistency, that is, it is expected that there can be a high morphological consistency between its three-dimensional volume representation and two-dimensional contour representation.
[0067] Therefore, first determine the three-dimensional volume representation vector as:
[0068] ;
[0069] The two-dimensional contour representation vector is:
[0070]
[0071] Then, through the trainable weighting hyperparameters and , to make the two-dimensional contour surface - three-dimensional volume space tensor have the regularity that satisfies the commutation symmetry, that is, to make the space two-norm representation of the exchange difference vector between the three-dimensional volume representation vector and the two-dimensional contour representation vector tend to the product of the coefficients and :
[0072]
[0073] where is the equal scaling coefficient.
[0074] In this way, by reasonably setting the two-dimensional contour constraint conditions and utilizing the spatial morphological consistency characteristics of the three-dimensional volume, the requirements of the regularity normalization criterion are met, so that the pixel-level feature vectors of the local visual semantics of the mold state differences in the local receptive field are highly unified in structure, significantly enhancing the context correlation feature saliency performance of the local visual semantic feature vectors of the mold state differences. After that, since the sum value of the trainable weighting hyperparameters of and is one during weighting, thus, the numerical values of the trainable weighting hyperparameters of and can be calculated.
[0075] Finally, based on the calculated numerical values of the trainable weighting hyperparameters of and , perform step S1632, that is, based on the pixel-level mold state difference local visual semantic fusion feature vector in the local receptive field, perform feature enhancement on the local visual semantic feature vector of the mold state difference to be enhanced to obtain the enhanced local visual semantic feature vector of the mold state difference.
[0076] Specifically, in step S164, multiple local visual semantic feature vectors of the enhanced die state differences are aggregated to obtain a local visual semantic enhanced coding feature map of the die state differences as the local visual semantic features of the enhanced die state differences. It should be understood that information aggregation not only helps to improve the ability to capture defect features, but also can effectively separate the subtle differences between normal wear and real defects, thus providing solid data support for the final classification decision. Considering that multiple local visual semantic feature vectors of the enhanced die state differences each represent information segments in different regions of the die surface, there may be overlapping or complementary relationships between them. Directly analyzing based on multiple local visual semantic feature vectors of the enhanced die state differences often makes it difficult to comprehensively reflect the overall state changes of the die, especially when it comes to complex structures across scales (such as the continuity of cracks and their extension directions). Therefore, by aggregating information, multiple local visual semantic feature vectors of the enhanced die state differences can be integrated to construct a more comprehensive and coherent feature representation. Specifically, the process of information aggregation aims to fuse the most representative feature information in each local region while suppressing the noise interference that may cause misjudgment. To achieve this goal, weighted average or other forms of linear combination methods are usually used to realize the fusion between multiple local visual semantic feature vectors of the enhanced die state differences. In this process, the weights should be dynamically adjusted according to the amount of information carried by each local visual semantic feature vector of the enhanced die state differences and its correlation with the surrounding environment. For example, in the area with obvious defect signals, the corresponding local visual semantic feature vector of the enhanced die state differences should be assigned a higher weight value to highlight the changes there; while in the relatively stable area or the area that only shows normal wear, its weight should be appropriately reduced to avoid introducing unnecessary artifacts that affect the overall judgment. In addition, the conformal representational constraint conditions can be combined to further optimize the weight allocation strategy to ensure that the local visual semantic enhanced coding feature map of the aggregated die state differences not only retains the important details in the original feature vectors but also can reflect the consistency of the global context.
[0077] In an embodiment of the present application, step S170, determining whether there are defects in the to-be-detected metal stamping die based on the die state difference local visual semantic enhanced coding feature map, includes: passing the die state difference local visual semantic enhanced coding feature map through a defect detector based on a classifier to obtain a defect detection result, where the defect detection result is used to indicate whether there are defects in the to-be-detected metal stamping die. It should be understood that compared with the original image data, the die state difference local visual semantic enhanced coding feature map can more directly reflect the actual condition of the die surface and is more easily understood and processed by machine learning algorithms. In this step, the classifier automatically determines which category the corresponding sample belongs to (for example, defective or non-defective) according to the input die state difference local visual semantic enhanced coding feature map. Although traditional classification methods such as Support Vector Machine (SVM), Decision Tree, etc. perform well in some specific scenarios, they are often unable to cope with the complex and changeable die surface conditions. In recent years, with the development of deep learning technology, especially the successful application of Convolutional Neural Network (CNN) and its derivative models in the field of image recognition, more and more research has begun to turn to using more advanced deep learning frameworks to construct defect detectors. Such models, relying on their powerful non-linear fitting ability and efficient feature learning mechanism, can directly extract the most discriminative information from the original data without the need for manual feature design and make accurate classification decisions based on this. Specifically, when the die state difference local visual semantic enhanced coding feature map is input into the defect detector based on the classifier, it actually allows the model to comprehensively scan and deeply analyze the die state difference local visual semantic enhanced coding feature map. In this process, the classifier first parses the input data layer by layer, and each layer is responsible for capturing feature information at different levels. For example, in the initial several layers, the model may focus on some low-level elements such as edges and textures; as the number of layers increases, it will gradually integrate these primary features to form more abstract concepts such as shapes and structures. Finally, at the top layer, the classifier will synthesize all the obtained information and output a prediction probability distribution for the sample category. This distribution reflects the likelihood of the current sample belonging to each category (i.e., defective and non-defective).
[0078] In summary, an intelligent vision inspection method for the production of hardware stamping dies based on the embodiments of the present application is elucidated. It acquires the current state image of the hardware stamping die to be inspected and the defect-free die state reference image and performs geometric alignment. Then, it calculates the difference map to obtain the saliency distribution of potential defect regions, and uses a pyramid network for hierarchical encoding of multi-scale difference features to separate normal wear artifacts and real defects. Further, a feature receptive field anchoring mechanism is introduced to dynamically adjust the attention weights according to the die geometry topology. Finally, a lightweight classifier is used to make a decision on the enhanced difference features to achieve sub-millimeter-level anomaly recognition. In this way, it can effectively improve the sensitivity to minor changes in hardware stamping dies and help detect defects at an early stage.
[0079] Figure 5 FIG. is a system block diagram of an intelligent vision inspection system for the production of hardware stamping dies according to an embodiment of the present application. As Figure 5 shown, the intelligent vision inspection system 100 for the production of hardware stamping dies according to an embodiment of the present application includes: a to-be-inspected hardware stamping die placement module 110 for placing the to-be-inspected hardware stamping die at the inspection station; a current state image acquisition module 120 of the die for acquiring the current state image of the to-be-inspected hardware stamping die through a camera; a defect-free die state reference image extraction module 130 for extracting the defect-free die state reference image from the background database; a significant coding module 140 for die state change, which geometrically aligns the die state reference image and the current state image of the die and then calculates the pixel-level difference between the images to obtain a die state change saliency map; a local visual difference feature extraction module 150 for extracting local visual difference features from the die state change saliency map to obtain a local visual semantic coding feature map of die state differences; a local visual semantic feature saliency module 160 for saliency of local visual semantic features of the local visual semantic coding feature map of die state differences to obtain a local visual semantic enhanced coding feature map of die state differences; a determination module 170 for the existence of defects in the to-be-inspected hardware stamping die, which determines whether there are defects in the to-be-inspected hardware stamping die based on the local visual semantic enhanced coding feature map of die state differences.
[0080] Here, those skilled in the art can understand that the specific operations of each step in the above intelligent vision inspection system for the production of hardware stamping dies have been introduced in detail in the description of the intelligent vision inspection method for the production of hardware stamping dies above, and therefore, the repeated description thereof will be omitted. Figures 1 to 4
Claims
1. An intelligent vision inspection method for the production of hardware stamping dies, characterized in that, Including: Placing the hardware stamping die to be detected at the detection station; Collecting the current state image of the hardware stamping die to be detected through a camera; Extracting the die state reference image marked as defect-free from the background database; Performing geometric alignment on the die state reference image and the current state image of the die, and then calculating the pixel-level difference between the images to obtain the die state change significance map; Extracting local visual difference features from the die state change significance map to obtain the die state difference local visual semantic coding feature map; Performing local visual semantic feature saliency on the die state difference local visual semantic coding feature map to obtain the die state difference local visual semantic enhanced coding feature map; Based on the die state difference local visual semantic enhanced coding feature map, determining whether there are defects in the hardware stamping die to be detected.
2. The intelligent vision inspection method for the production of hardware stamping dies according to claim 1, wherein Performing geometric alignment on the die state reference image and the current state image of the die, and then calculating the pixel-level difference between the images to obtain the die state change significance map, including: Performing image preprocessing on the die state reference image and the current state image of the die to obtain the preprocessed die state reference image and the preprocessed current state image of the die; Performing geometric alignment on the preprocessed die state reference image and the preprocessed current state image of the die to align the preprocessed current state image of the die to the coordinate system of the preprocessed die state reference image to obtain the aligned current state image of the die; Calculating the pixel-level difference between the aligned current state image of the die and the preprocessed die state reference image to obtain the die state change significance map.
3. The intelligent vision inspection method for the production of hardware stamping dies according to claim 2, wherein Performing image preprocessing on the die state reference image and the current state image of the die to obtain the preprocessed die state reference image and the preprocessed current state image of the die, including: performing non-local mean filtering on the die state reference image and the current state image of the die, and then performing global contrast normalization to obtain the preprocessed die state reference image and the preprocessed current state image of the die.
4. The intelligent vision inspection method for the production of hardware stamping dies according to claim 3, characterized in that, Extracting local visual difference features from the die state change significance map to obtain the die state difference local visual semantic coding feature map, including: passing the die state change significance map through a local visual difference feature extractor based on a pyramid network to obtain the die state difference local visual semantic coding feature map.
5. The intelligent vision inspection method for the production of hardware stamping dies according to claim 4, characterized in that, Performing local visual semantic feature saliency on the die state difference local visual semantic coding feature map to obtain the die state difference local visual semantic enhanced coding feature map, including: First performing feature dissociation on the die state difference local visual semantic coding feature map and then performing information compression processing to obtain the die state difference local visual semantic feature vector to be enhanced and distilled; Based on the spatial distribution of the die state difference local visual semantic feature vector to be enhanced and distilled, screening out the set of pixel-level die state difference local visual semantic feature vectors within the local receptive field from the set of pixel-level die state difference local visual semantic feature vectors; Based on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field, perform saliency enhancement on the local visual semantic feature vectors of die state differences to be enhanced to obtain enhanced local visual semantic feature vectors of die state differences; Aggregate the information of multiple enhanced local visual semantic feature vectors of die state differences to obtain a local visual semantic enhanced coding feature map of die state differences as the enhanced local visual semantic feature of die state differences.
6. The intelligent vision inspection method for the production of hardware stamping dies according to claim 5, wherein Perform feature dissociation on the local visual semantic coding feature map of die state differences and then perform information compression processing to obtain local visual semantic feature vectors of die state differences to be enhanced by distillation, including: Perform feature dissociation on the local visual semantic coding feature map of die state differences along the channel dimension to obtain a set of local visual semantic pixel-level feature vectors of die state differences; Extract the local visual semantic pixel-level feature vector of die state differences at the (i,j) pixel position in the local visual semantic coding feature map of die state differences from the set of local visual semantic pixel-level feature vectors of die state differences as the local visual semantic feature vector of die state differences to be enhanced; Perform information compression on the local visual semantic feature vector of die state differences to be enhanced to obtain the local visual semantic feature vector of die state differences to be enhanced by distillation.
7. The intelligent vision inspection method for the production of hardware stamping dies according to claim 6, characterized in that, Based on the spatial distribution of the local visual semantic feature vectors of die state differences to be enhanced by distillation, screen out a set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field from the set of local visual semantic pixel-level feature vectors of die state differences, including: Determine the size of the feature receptive field of the local visual semantic feature vector of die state differences to be enhanced based on the feature distribution spatial structure characteristics of the local visual semantic feature vector of die state differences to be enhanced by distillation; Based on the size of the feature receptive field, screen out the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field from the set of local visual semantic pixel-level feature vectors of die state differences.
8. The intelligent vision inspection method for the production of hardware stamping dies according to claim 7, characterized in that Based on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field, perform saliency enhancement on the local visual semantic feature vectors of die state differences to be enhanced to obtain enhanced local visual semantic feature vectors of die state differences, including: Perform saliency weighted fusion on the set of local visual semantic feature vectors of pixel-level die state differences within the local receptive field to obtain a local visual semantic fusion feature vector of pixel-level die state differences within the local receptive field; Based on the local visual semantic fusion feature vector of pixel-level die state differences within the local receptive field, perform feature enhancement on the local visual semantic feature vectors of die state differences to be enhanced to obtain the enhanced local visual semantic feature vectors of die state differences, where the enhanced local visual semantic feature vectors of die state differences are the channel feature vectors at the (i,j) pixel position of the local visual semantic enhanced coding feature map of die state differences.
9. The intelligent vision inspection method for the production of hardware stamping dies according to claim 8, wherein, Based on the locally visually semantic enhanced encoded feature map of the die state difference, determining whether there are defects in the hardware stamping die to be detected, including: passing the locally visually semantic enhanced encoded feature map of the die state difference through a defect detector based on a classifier to obtain a defect detection result, and the defect detection result is used to indicate whether there are defects in the hardware stamping die to be detected.
10. An intelligent vision inspection system for the production of hardware stamping dies, characterized in that, Including: A module for placing the hardware stamping die to be detected, which is used to place the hardware stamping die to be detected at the detection station; A module for collecting the current die state image, which is used to collect the current die state image of the hardware stamping die to be detected through a camera; A module for extracting the reference image of the die state without defects, which is used to extract the reference image of the die state marked as defect-free from the background database; A module for significantly encoding the die state change, which is used to geometrically align the reference image of the die state and the current die state image and then calculate the pixel-level difference between the images to obtain a die state change significance map; A module for extracting local visual difference features, which is used to extract local visual difference features from the die state change significance map to obtain a locally visually semantic encoded feature map of the die state difference; A module for significantly enhancing local visual semantic features, which is used to significantly enhance the local visual semantic features of the locally visually semantic encoded feature map of the die state difference to obtain a locally visually semantic enhanced encoded feature map of the die state difference; A module for determining whether there are defects in the hardware stamping die to be detected, which is used to determine whether there are defects in the hardware stamping die to be detected based on the locally visually semantic enhanced encoded feature map of the die state difference.
Citation Information
Cited By
Injection mold production quality supervision method and system
CN120807516A
A method and system for quality supervision in injection mold production
CN120807516B