Multi-dimensional coupling spatiotemporal anomaly positioning method, device, equipment and medium for die castings
By acquiring and fusing visible light, X-ray, and infrared images of die-cast parts, and combining them with multi-dimensional information fusion methods, the problem of insufficient positioning accuracy of die-cast parts in existing technologies has been solved, and high-precision multi-dimensional anomaly positioning has been achieved.
Patent Information
- Application Number
- CN202610456864.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-03
AI Technical Summary
In existing technologies, abnormal positioning of die-cast parts often relies on comparison of a single mode or static features, failing to effectively integrate process sensitivity, spatial structure, and modal reliability. This makes it difficult to adapt to structural differences and modal reliability changes in different regions, resulting in insufficient positioning accuracy and robustness.
Visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of die-cast parts are acquired. By using a multi-dimensional coupled spatiotemporal anomaly localization method, spatial and temporal information of visible light, X-ray, and infrared modes are fused, and modal confidence and inconsistency scores of each pixel are calculated to achieve multi-dimensional information fusion localization.
It improves the pixel-level precision positioning capability of internal defects in die-cast parts, enhances the ability to capture minute defects, and improves positioning accuracy and robustness.
Smart Images

Figure CN122335787A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of anomaly detection technology, and in particular to a method, apparatus, equipment and medium for locating multi-dimensional coupled spatiotemporal anomalies in die castings. Background Technology
[0002] Internal defects (porosity, shrinkage) in integrated die-cast parts for new energy vehicles require pixel-level precision to ensure maintenance and quality traceability. In related technologies, anomaly localization of die-cast parts often relies on single-modal or static feature comparisons, failing to integrate multi-dimensional information such as process sensitivity, spatial structure, and modal reliability. This makes it difficult to adapt to structural differences and modal reliability variations in different regions (thin-walled / thick-walled) of the die-cast part, and the positioning accuracy and robustness cannot meet industrial requirements. Summary of the Invention
[0003] To address the aforementioned technical problems, this application provides a method, apparatus, electronic device, storage medium, and computer program product for locating multi-dimensional coupled spatiotemporal anomalies in die-cast parts.
[0004] According to a first aspect of this application, a multi-dimensional coupled spatiotemporal anomaly localization method for die-casting parts is provided, comprising:
[0005] Acquire visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of the die-cast parts;
[0006] The visible light modal confidence of each pixel is determined based on the visible light surface image, the X-ray modal confidence of each pixel is determined based on the X-ray transmission image, and the infrared modal confidence of each pixel is determined based on the infrared thermal imaging image sequence.
[0007] Extract external spatial feature maps from the visible light images, extract internal spatial feature maps from the X-ray transmission images, and extract local temporal feature maps from the infrared thermal imaging image sequence;
[0008] Based on the external spatial feature map and the internal spatial feature map, the spatial inconsistency score of each pixel is determined, and based on the local temporal feature map, the temporal inconsistency score of each pixel is determined.
[0009] Based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, spatial inconsistency score and temporal inconsistency score of each pixel, the visible light mode weight, X-ray mode weight and infrared mode weight of each pixel are determined.
[0010] The initial inconsistency score of each pixel is determined by adding the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score, and the product of the infrared mode weights and the temporal inconsistency score.
[0011] The final inconsistency score of each pixel is determined based on the initial inconsistency score of each pixel, and the abnormal region is determined based on the final inconsistency score of each pixel.
[0012] Optionally, determining the final inconsistency score of each pixel based on the initial inconsistency score of each pixel includes:
[0013] Based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image and process sensitivity weight of each pixel, the multidimensional coupling weight of each pixel is determined;
[0014] The final inconsistency score of each pixel is determined by multiplying the initial inconsistency score of each pixel with the multidimensional coupling weight.
[0015] Optionally, determining the final inconsistency score of each pixel based on the initial inconsistency score of each pixel includes:
[0016] The external spatial feature map, the internal spatial feature map, and the local temporal feature map are fused locally to obtain a three-modal local fusion feature.
[0017] The external spatial feature map, the internal spatial feature map, and the local temporal feature map are fused globally to obtain a three-modal global fused feature.
[0018] Calculate the similarity between the three-modal local fusion features and the three-modal global fusion features, and determine the inconsistency score of the fusion features for each pixel based on the similarity.
[0019] The initial inconsistency score of each pixel and the inconsistency score of the fused feature of each pixel are weighted and averaged to obtain the final inconsistency score of each pixel.
[0020] Optionally, determining the visible light mode weight, X-ray mode weight, and infrared mode weight of each pixel based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, spatial inconsistency score, and temporal inconsistency score of each pixel includes:
[0021] Sure Local mean of spatial inconsistency score Local mean of time series inconsistency score ;
[0022] According to the following formula:
[0023] ,Sure Visible mode weights of position ;
[0024] According to the following formula:
[0025] ,Sure Location-based X-ray modal weights ;
[0026] According to the formula: , Infrared mode weights of position ;
[0027] in, express Confidence of location in visible light modes. express X-ray modal confidence of location express Infrared modal confidence of location.
[0028] Optionally, determining the multidimensional coupling weight of each pixel based on its visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image, and process sensitivity weight includes:
[0029] The visible light surface image and the X-ray transmission image are fused to obtain a fused image, and the local grayscale gradient of each pixel extracted from the fused image is used as the spatial weight of each pixel.
[0030] The multimodal weights of each pixel are determined based on the visible light mode confidence, X-ray mode confidence, and infrared mode confidence of each pixel.
[0031] The spatial weight, process sensitivity weight, and multimodal weight are weighted and averaged, and the weighted average is processed using the Sigmoid function to obtain the multidimensional coupling weight of each pixel.
[0032] Optionally, determining the spatial inconsistency score of each pixel based on the external spatial feature map and the internal spatial feature map includes:
[0033] The external spatial feature map and the internal spatial feature map are respectively subjected to local feature fusion to obtain two-modal local fusion features;
[0034] The external spatial feature map and the internal spatial feature map are respectively subjected to global feature fusion to obtain two-modal global fused features;
[0035] Calculate the similarity between the two-modal local fusion features and the two-modal global fusion features, and determine the spatial inconsistency score of each pixel based on the similarity.
[0036] Optionally, determining the temporal inconsistency score of each pixel based on the local temporal feature map includes:
[0037] Temporal similarity is extracted from the local temporal feature map using a dynamic time warping algorithm;
[0038] Temporal inconsistency scores are determined based on temporal similarity.
[0039] According to a second aspect of this application, a multi-dimensional coupled spatiotemporal anomaly localization device for die-cast parts is provided, comprising:
[0040] The image acquisition module is used to acquire visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of the die-cast parts;
[0041] The modal confidence determination module is used to determine the visible light modal confidence of each pixel based on the visible light surface image, the X-ray modal confidence of each pixel based on the X-ray transmission image, and the infrared modal confidence of each pixel based on the infrared thermal imaging image sequence.
[0042] The feature extraction module is used to extract external spatial feature maps from the visible light image, internal spatial feature maps from the X-ray transmission image, and local temporal feature maps from the infrared thermal imaging image sequence.
[0043] The spatial inconsistency score determination module is used to determine the spatial inconsistency score of each pixel based on the external spatial feature map and the internal spatial feature map.
[0044] The temporal inconsistency score determination module is used to determine the temporal inconsistency score of each pixel based on the local temporal feature map.
[0045] The modal weight determination module is used to determine the visible light modal weight, X-ray modal weight, and infrared modal weight of each pixel based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, spatial inconsistency score, and temporal inconsistency score of each pixel.
[0046] The initial inconsistency score determination module is used to add the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score to the product of the infrared mode weights and the temporal inconsistency score to determine the initial inconsistency score of each pixel.
[0047] The final inconsistency score determination module is used to determine the final inconsistency score of each pixel based on the initial inconsistency score of each pixel;
[0048] The abnormal region determination module is used to determine abnormal regions based on the final inconsistency score of each pixel.
[0049] Optionally, the final inconsistency score determination module is specifically used to determine the multidimensional coupling weight of each pixel based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image and process sensitivity weight; and to determine the final inconsistency score of each pixel by multiplying the initial inconsistency score of each pixel with the multidimensional coupling weight.
[0050] Optionally, the final inconsistency score determination module is specifically used to perform local feature fusion of the external spatial feature map, the internal spatial feature map, and the local temporal feature map to obtain trimodal local fusion features; to perform global feature fusion of the external spatial feature map, the internal spatial feature map, and the local temporal feature map to obtain trimodal global fusion features; to calculate the similarity between the trimodal local fusion features and the trimodal global fusion features, and to determine the fusion feature inconsistency score of each pixel based on the similarity; and to perform weighted averaging of the initial inconsistency score of each pixel and the fusion feature inconsistency score of each pixel to obtain the final inconsistency score of each pixel.
[0051] Optionally, the modal weight determination module is specifically used to determine... Local mean of spatial inconsistency score Local mean of time series inconsistency score ;as well as,
[0052] According to the following formula:
[0053] ,Sure Visible mode weights of position ;
[0054] According to the following formula: ,Sure Location-based X-ray modal weights ;
[0055] According to the formula: , Infrared mode weights of position ;
[0056] in, express Confidence of location in visible light modes. express X-ray modal confidence of location express Infrared modal confidence of location.
[0057] Optionally, the final inconsistency score determination module is specifically used to determine the multidimensional coupling weights of each pixel through the following steps:
[0058] The visible light surface image and the X-ray transmission image are fused to obtain a fused image, and the local grayscale gradient of each pixel extracted from the fused image is used as the spatial weight of each pixel.
[0059] The multimodal weights of each pixel are determined based on the visible light mode confidence, X-ray mode confidence, and infrared mode confidence of each pixel.
[0060] The spatial weight, process sensitivity weight, and multimodal weight are weighted and averaged, and the weighted average is processed using the Sigmoid function to obtain the multidimensional coupling weight of each pixel.
[0061] Optionally, the spatial inconsistency score determination module is specifically used to perform local feature fusion on the external spatial feature map and the internal spatial feature map respectively to obtain two-modal local fusion features; perform global feature fusion on the external spatial feature map and the internal spatial feature map respectively to obtain two-modal global fusion features; calculate the similarity between the two-modal local fusion features and the two-modal global fusion features, and determine the spatial inconsistency score of each pixel based on the similarity.
[0062] Optionally, the timing inconsistency score determination module is specifically used to extract timing similarity from the local timing feature map using a dynamic time warping algorithm; and to determine the timing inconsistency score based on the timing similarity.
[0063] According to a third aspect of this application, an electronic device is provided, comprising: a processor configured to execute a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the method described in the first aspect.
[0064] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0065] According to a fifth aspect of this application, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to perform the method described in the first aspect.
[0066] The technical solution provided in this application has the following advantages compared with the prior art:
[0067] Cross-modal anomaly localization is achieved by acquiring images in three modalities: visible light, X-ray, and infrared. Specifically, spatial information from the visible light and X-ray modalities is used to capture spatial anomalies caused by defects, obtaining spatial inconsistency scores for each pixel. Temporal information from the infrared modal is used to capture temporal (thermal conduction) anomalies caused by defects, obtaining temporal inconsistency scores for each pixel, thus improving the ability to detect minute defects. By calculating the modal confidence scores of the three modalities, modal reliability is determined. Then, multi-dimensional information including spatial, temporal, and modal reliability is fused for anomaly localization, thereby improving the accuracy of pixel-level localization. Attached Figure Description
[0068] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0069] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a flowchart of a multi-dimensional coupled spatiotemporal anomaly localization method for die-cast parts in an embodiment of this application;
[0071] Figure 2 This is a flowchart illustrating how to determine the final inconsistency score of each pixel based on the initial inconsistency score of each pixel in an embodiment of this application.
[0072] Figure 3 This is another flowchart illustrating how the final inconsistency score of each pixel is determined based on the initial inconsistency score of each pixel in the embodiments of this application.
[0073] Figure 4 This is a schematic diagram of a multi-dimensional coupled spatiotemporal anomaly positioning device for die-cast parts in an embodiment of this application;
[0074] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of this application. Detailed Implementation
[0075] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0076] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0077] See Figure 1 , Figure 1 This is a flowchart of a multi-dimensional coupled spatiotemporal anomaly localization method for die-cast parts in this application embodiment, which may include the following steps:
[0078] Step S102: Acquire visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of the die-cast part.
[0079] Visible light surface images X-ray transmission images reflect the surface texture, scratches, and other appearance characteristics of die-cast parts. Infrared thermal imaging image sequence revealing internal structural defects It can capture thermal conduction anomalies, where N represents the number of infrared thermal imaging image sequences. It can simultaneously acquire multimodal data of die-cast parts, triggered by TTL (Transistor-Transistor Logic) hardware with a time alignment error of <10ms. Spatial alignment is achieved through a calibration board to ensure pixel-level spatial matching of the three modal images.
[0080] Step S104: Determine the visible light mode confidence of each pixel based on the visible light surface image, determine the X-ray mode confidence of each pixel based on the X-ray transmission image, and determine the infrared mode confidence of each pixel based on the infrared thermal imaging image sequence.
[0081] Visible light mode confidence According to the formula: Sure, For visible light images in The local noise intensity at a location (calculated using Gaussian filtered residuals) indicates that the lower the noise, the more reliable the visible light mode (high confidence in areas with unobstructed surfaces and clear textures). This represents the maximum noise intensity of the visible light image.
[0082] X-ray modal confidence According to the formula: Sure, For X-ray images in The local contrast of a location (calculated by the gray-level difference of neighboring pixels) indicates that the higher the contrast, the clearer the internal structural features and the stronger the reliability of the X-ray modality. This represents the maximum contrast of an X-ray transmission image.
[0083] Infrared modal confidence According to the formula: Sure, for Local temperature standard deviation at each location. For each spatial location in the infrared thermal imaging image sequence. Extract its in Intra-frame temperature change sequence The degree of temperature fluctuation at that location is calculated using the standard deviation formula. This reflects the temperature stability at that location. Defective areas (such as pores) experience significantly greater temperature fluctuations than normal areas due to abnormal heat conduction (normal metal areas have uniform heat conduction and small temperature fluctuations). It can be directly used as a reverse indicator of the "reliability" of infrared modes (the smaller the fluctuation, the more reliable the infrared mode). This represents the maximum value of the temperature standard deviation of an infrared thermal imaging image sequence.
[0084] Step S106: Extract external spatial feature maps from visible light images, extract internal spatial feature maps from X-ray transmission images, and extract local temporal feature maps from infrared thermal imaging image sequences.
[0085] In this embodiment, a visible light encoder can be used to extract surface texture and structural features from a visible light image, outputting an external spatial feature map. The visible light encoder can employ a lightweight convolutional neural network (such as MobileNetV3), and the external spatial feature map can be represented as follows: d=128. An X-ray encoder is used to penetrate and capture internal structural features, outputting an internal spatial feature map. The X-ray encoder can employ a convolutional network or a lightweight U-Net variant. The internal spatial feature map can be represented as... Local temporal feature maps are extracted from infrared thermal imaging image sequences using an infrared encoder. The infrared encoder can employ I3D (Inflated 3D ConvNet) + temporal attention to focus on the temporal changes in heat conduction and output local temporal feature maps. Each spatial location Each corresponds to a d-dimensional temporal feature vector.
[0086] Step S108: Based on the external spatial feature map and the internal spatial feature map, determine the spatial inconsistency score of each pixel, and based on the local temporal feature map, determine the temporal inconsistency score of each pixel.
[0087] Both the external and internal spatial feature maps reflect the spatial structure of the die-cast part; therefore, they can be combined to determine the spatial inconsistency score. Optionally, the external and internal spatial feature maps can be fused locally to obtain two-modal local fusion features. The external and internal spatial feature maps can also be fused globally to obtain two-modal global fusion features. The similarity between the two-modal local fusion features and the two-modal global fusion features is calculated, and the spatial inconsistency score for each pixel is determined based on the similarity. This can be expressed as the following formula:
[0088] ;
[0089] in, express Spatial inconsistency score of location express External spatial characteristics of location With interior space features The fusion feature (element-level addition). Representation of external spatial feature map Internal space feature map The global fusion feature Similarity is used to measure the consistency between the local fusion features of two modalities and the global fusion features of two modalities. It can be expressed as cosine similarity. The maximum similarity value is the value of normal die-cast parts without anomalies, used to normalize the similarity to... .
[0090] Infrared thermal imaging image sequences reflect the temporal characteristics of die-cast parts. The DTW (Dynamic Time Warping) algorithm can be used to extract temporal similarity from local temporal feature maps; based on temporal similarity, a temporal inconsistency score is determined. This can be expressed as the following formula:
[0091] ;
[0092] in, The temporal inconsistency score is represented by the position (𝑖,𝑗). This represents a sequence of local infrared thermal images at location (x, y). This represents a global infrared thermal imaging image sequence. This refers to the dynamic time warping algorithm, used to capture differences in time series trends. This represents the maximum DTW value for normal die-cast parts without any anomalies. Introducing a timing inconsistency score can capture dynamic physical anomalies caused by defects, improving the sensitivity of detecting minute defects.
[0093] Step S110: Based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, spatial inconsistency score and temporal inconsistency score of each pixel, determine the visible light mode weight, X-ray mode weight and infrared mode weight of each pixel.
[0094] Visible light mode weights and X-ray mode weights are used as weights for spatial inconsistency scores, while infrared mode weights are used as weights for temporal inconsistency scores to calculate the initial inconsistency score. For each mode, the higher the mode confidence, the higher the corresponding mode weight. In this embodiment, the modal weights can also be corrected using the intensity of defect-related anomalous signals (i.e., spatial inconsistency scores and temporal inconsistency scores) to improve the accuracy of the initial inconsistency score calculation. The higher the spatial inconsistency score, the higher the visible light mode weights and X-ray mode weights; the higher the temporal inconsistency score, the higher the infrared mode weight.
[0095] Optionally, it can be determined Local mean of spatial inconsistency score Local mean of time series inconsistency score That is, you can select... Spatial inconsistency scores are calculated for multiple locations near the given location, and the average is obtained to obtain the local mean of the spatial inconsistency scores. Select The temporal inconsistency scores of multiple locations near the given location are calculated, and the average is obtained to obtain the local mean of the temporal inconsistency scores. Then, according to the following formula:
[0096] ,Sure Visible mode weights of position ;
[0097] According to the following formula:
[0098] ,Sure Location-based X-ray modal weights ;
[0099] According to the formula: , Infrared mode weights of position ;in, express Confidence of location in visible light modes. express X-ray modal confidence of location express Infrared modal confidence of location.
[0100] It should be noted that the calculation method for each modal weight is not limited to this. For example, simple modifications to the above formula are all within the scope of protection of this application.
[0101] Step S112: Add the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score to the product of the infrared mode weights and the temporal inconsistency score to determine the initial inconsistency score of each pixel.
[0102] Visible light and X-ray modes correspond to spatial information, while infrared modes correspond to temporal information. Therefore... Initial inconsistency score of position .
[0103] Step S114: Determine the final inconsistency score of each pixel based on the initial inconsistency score of each pixel, and determine the abnormal region based on the final inconsistency score of each pixel.
[0104] In this embodiment, the initial inconsistency score of each pixel can be directly used as the final inconsistency score of each pixel, or the initial inconsistency score of each pixel can be combined with other dimensions to determine the final inconsistency score. For each pixel, if the final inconsistency score of the pixel is greater than a preset threshold, then the pixel is determined to be an abnormal region.
[0105] See Figure 2 , Figure 2 A flowchart illustrating how to determine the final inconsistency score of each pixel based on the initial inconsistency score in an embodiment of this application includes the following steps:
[0106] Step S1141: Based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, visible light surface image, X-ray transmission image and process sensitivity weight of each pixel, determine the multidimensional coupling weight of each pixel.
[0107] Anomaly localization is related to multi-dimensional information such as process sensitivity, spatial structure, and modal reliability. Therefore, a three-factor coupling modulation mechanism of "process sensitivity-spatial structure-modal reliability" can be designed to dynamically adapt to the multi-dimensional differences in different regions. The process sensitivity of each location in a die-cast part is different, which can be obtained through output from an external process modeling module or pre-training with historical data. For example, mainstream industrial simulation and quality control software (such as ANSYS and Siemens Teamcenter) have built-in "process sensitivity analysis modules," which can automatically output process parameter sensitivity matrices by importing historical data or mechanistic models.
[0108] The calculation method for multidimensional coupling weights can be as follows:
[0109] A fused image is obtained by fusing a visible light surface image and an X-ray transmission image. The fused image can be represented as follows: , This is a pixel-level multiplication that integrates surface and internal structural complexity. The local grayscale gradients of each pixel extracted from the fused image are used as the spatial weights of each pixel. The spatial weight of a location can be expressed as , This represents gradient operation.
[0110] Based on the visible light mode confidence, X-ray mode confidence, and infrared mode confidence of each pixel, the multimodal weights of each pixel are determined. For example, the multimodal weights can be... .
[0111] The spatial weight, process sensitivity weight, and multimodal weight are weighted and averaged, and then the sigmoid function is used to process the weighted average to obtain the multidimensional coupling weight of each pixel. This can be expressed as the following formula:
[0112] Multidimensional Coupling Weights ;
[0113] in, Sigmoid function , , For example, the balance coefficient. , , , express Location-based process sensitivity weighting.
[0114] Step S1143: The product of the initial inconsistency score of each pixel and the multidimensional coupling weight is determined as the final inconsistency score of each pixel.
[0115] Final Inconsistency Score .
[0116] See Figure 3 , Figure 3 Another flowchart illustrating the determination of the final inconsistency score of each pixel based on the initial inconsistency score of each pixel in this application embodiment includes the following steps:
[0117] Step S1142: Perform local feature fusion on the external spatial feature map, internal spatial feature map and local temporal feature map to obtain trimodal local fusion features.
[0118] For each position The three-modal features at this location can be fused to obtain a three-modal local fused feature. This can be represented as:
[0119] .
[0120] Conv1d's role is to fuse and unify the dimensions of local multimodal features. It compresses the concatenated high-dimensional sequence into a single, unified form. Local fusion features of dimensions This approach preserves local feature information from all three modalities while achieving dimensionality uniformity, laying the foundation for subsequent similarity comparison with global features. The kernel size is set to... (Matching feature splicing length for three modalities), output channel number set to (Consistent with the feature dimension of a single modality).
[0121] Step S1144: Global feature fusion is performed on the external spatial feature map, the internal spatial feature map, and the local temporal feature map to obtain the three-modal global fusion feature.
[0122] First, the external spatial feature map, internal spatial feature map, and local temporal feature map are subjected to global average pooling to obtain global average pooled features of three modalities. Then, AttentionFusion is used to adaptively fuse the global multimodal features, as shown below:
[0123]
[0124] This represents the global fusion feature of the three modes. This represents the global average pooling characteristics of the visible light modes. This represents the global average pooling feature of the X-ray modality. This represents the global average pooling characteristics of the infrared modality.
[0125] During feature fusion, the weights of the global average pooling features for each modality can be calculated using an attention mechanism. For example, the global average pooling features for the X-ray modality are more sensitive to internal defects and have higher weights; the global average pooling features for the visible light modality have lower weights when affected by surface noise. The calculation of attention weights depends on the confidence levels of the three modalities (…). , , As a priori, this ensures that the fusion weights are strongly correlated with modal reliability.
[0126] Step S1146: Calculate the similarity between the three-modal local fusion features and the three-modal global fusion features, and determine the inconsistency score of the fusion features of each pixel based on the similarity.
[0127] The inconsistency score of location fusion features can be represented as: .
[0128] In step S1148, the initial inconsistency score of each pixel and the fusion feature inconsistency score of each pixel are weighted and averaged to obtain the final inconsistency score of each pixel.
[0129] .
[0130] , For example, weighting coefficients. , .
[0131] In this embodiment of the application, it can also be combined with Figure 2 and Figure 3 The method shown determines the final inconsistency score. It can be represented as follows:
[0132]
[0133] .
[0134] By combining information from more dimensions to determine the final inconsistency score, the accuracy of the final inconsistency score determination can be improved, thereby improving the accuracy of anomaly localization.
[0135] The multi-dimensional coupled spatiotemporal anomaly localization method for die-cast parts in this application dynamically adapts to multi-dimensional differences in different regions by fusing process sensitivity, spatial structure, and modal reliability through a coupled modulation mechanism. It captures dynamic physical anomalies of defects through a spatiotemporal dual-dimensional verification method (i.e., DTW time-series analysis and multi-modal spatial feature consistency verification). Through local-global multi-modal fusion, it achieves collaborative verification of visible light, X-ray, and infrared light modal features. This multi-dimensional collaborative verification improves pixel-level positioning accuracy.
[0136] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0137] This application also provides a multi-dimensional coupled spatiotemporal anomaly localization device for die-cast parts. See [link to relevant documentation]. Figure 4 The die-casting multi-dimensional coupled spatiotemporal anomaly positioning device 400 includes:
[0138] Image acquisition module 402 is used to acquire visible light surface images, X-ray transmission images and infrared thermal imaging image sequences of die-cast parts;
[0139] The modal confidence determination module 404 is used to determine the visible light modal confidence of each pixel based on the visible light surface image, the X-ray modal confidence of each pixel based on the X-ray transmission image, and the infrared modal confidence of each pixel based on the infrared thermal imaging image sequence.
[0140] The feature extraction module 406 is used to extract external spatial feature maps from visible light images, internal spatial feature maps from X-ray transmission images, and local temporal feature maps from infrared thermal imaging image sequences.
[0141] The spatial inconsistency score determination module 408 is used to determine the spatial inconsistency score of each pixel based on the external spatial feature map and the internal spatial feature map.
[0142] The temporal inconsistency score determination module 410 is used to determine the temporal inconsistency score of each pixel based on the local temporal feature map;
[0143] The modal weight determination module 412 is used to determine the visible light modal weight, X-ray modal weight and infrared modal weight of each pixel based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, spatial inconsistency score and temporal inconsistency score of each pixel.
[0144] The initial inconsistency score determination module 414 is used to add the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score to the product of the infrared mode weights and the temporal inconsistency score to determine the initial inconsistency score of each pixel.
[0145] The final inconsistency score determination module 416 is used to determine the final inconsistency score of each pixel based on the initial inconsistency score of each pixel;
[0146] The abnormal region determination module 418 is used to determine abnormal regions based on the final inconsistency score of each pixel.
[0147] Optionally, the final inconsistency score determination module 416 is specifically used to determine the multidimensional coupling weight of each pixel based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image and process sensitivity weight; and to determine the final inconsistency score of each pixel by multiplying the initial inconsistency score of each pixel with the multidimensional coupling weight.
[0148] Optionally, the final inconsistency score determination module 416 is specifically used to perform local feature fusion of the external spatial feature map, the internal spatial feature map, and the local temporal feature map to obtain trimodal local fusion features; to perform global feature fusion of the external spatial feature map, the internal spatial feature map, and the local temporal feature map to obtain trimodal global fusion features; to calculate the similarity between the trimodal local fusion features and the trimodal global fusion features, and to determine the fusion feature inconsistency score of each pixel based on the similarity; and to perform weighted average processing on the initial inconsistency score of each pixel and the fusion feature inconsistency score of each pixel to obtain the final inconsistency score of each pixel.
[0149] Optionally, the modal weight determination module 412 is specifically used to determine... Local mean of spatial inconsistency score Local mean of time series inconsistency score ;as well as,
[0150] According to the following formula:
[0151] ;
[0152] Sure Visible mode weights of position ;
[0153] According to the following formula:
[0154] ;
[0155] Sure Location-based X-ray modal weights ;
[0156] According to the formula: ; Infrared mode weights of position ;
[0157] in, express Confidence of location in visible light modes. express X-ray modal confidence of location express Infrared modal confidence of location.
[0158] Optionally, the final inconsistency score determination module 416 is specifically used to determine the multidimensional coupling weights of each pixel through the following steps:
[0159] The visible light surface image and the X-ray transmission image are fused to obtain a fused image, and the local gray-level gradient of each pixel extracted from the fused image is used as the spatial weight of each pixel.
[0160] The multimodal weights of each pixel are determined based on the visible light mode confidence, X-ray mode confidence, and infrared mode confidence of each pixel.
[0161] The spatial weight, process sensitivity weight, and multimodal weight are weighted and averaged, and the sigmoid function is used to process the weighted average to obtain the multidimensional coupling weight of each pixel.
[0162] Optionally, the spatial inconsistency score determination module 408 is specifically used to perform local feature fusion on the external spatial feature map and the internal spatial feature map respectively to obtain two-modal local fusion features; to perform global feature fusion on the external spatial feature map and the internal spatial feature map respectively to obtain two-modal global fusion features; to calculate the similarity between the two-modal local fusion features and the two-modal global fusion features, and to determine the spatial inconsistency score of each pixel based on the similarity.
[0163] Optionally, the timing inconsistency score determination module 410 is specifically used to extract timing similarity from the local timing feature map using a dynamic time warping algorithm; and to determine the timing inconsistency score based on the timing similarity.
[0164] The specific details of each module or unit in the above-mentioned device have been described in detail in the corresponding methods, so they will not be repeated here.
[0165] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0166] This application also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the above-described method for locating multi-dimensional coupled spatiotemporal anomalies in die-cast parts.
[0167] Reference Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.
[0168] like Figure 5 As shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.
[0169] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0170] Communication interface 504 is used to communicate with other electronic devices or servers.
[0171] The processor 502 is used to execute program 510, specifically the relevant steps in the above method embodiments.
[0172] Specifically, program 510 may include program code that includes computer operation instructions.
[0173] Processor 502 may be a central processing unit, a specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0174] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0175] Specifically, program 510 can be used to cause processor 502 to execute the steps in the above embodiment of the multi-dimensional coupling spatiotemporal anomaly localization method for die-cast parts.
[0176] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0177] In this embodiment of the application, a computer-readable storage medium is also provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the above-mentioned method for multi-dimensional coupled spatiotemporal anomaly localization of die-cast parts is implemented.
[0178] It should be noted that the computer-readable storage medium shown in this application can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency, etc., or any suitable combination thereof.
[0179] In this embodiment of the application, a computer program product is also provided, which, when run on a computer, causes the computer to execute the above-described method for locating multi-dimensional coupled spatiotemporal anomalies in die-cast parts.
[0180] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0181] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for positioning multi-dimensional coupling space-time anomalies of die castings, characterized in that, include: Acquire visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of the die-cast parts; The visible light modal confidence of each pixel is determined based on the visible light surface image, the X-ray modal confidence of each pixel is determined based on the X-ray transmission image, and the infrared modal confidence of each pixel is determined based on the infrared thermal imaging image sequence. Extract external spatial feature maps from the visible light images, extract internal spatial feature maps from the X-ray transmission images, and extract local temporal feature maps from the infrared thermal imaging image sequence; Based on the external spatial feature map and the internal spatial feature map, the spatial inconsistency score of each pixel is determined, and based on the local temporal feature map, the temporal inconsistency score of each pixel is determined. Based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, spatial inconsistency score and temporal inconsistency score of each pixel, the visible light mode weight, X-ray mode weight and infrared mode weight of each pixel are determined. The initial inconsistency score of each pixel is determined by adding the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score, and the product of the infrared mode weights and the temporal inconsistency score. The final inconsistency score of each pixel is determined based on the initial inconsistency score of each pixel, and the abnormal region is determined based on the final inconsistency score of each pixel.
2. The method of claim 1, wherein, The determination of the final inconsistency score for each pixel based on the initial inconsistency score of each pixel includes: Based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image and process sensitivity weight of each pixel, the multidimensional coupling weight of each pixel is determined; The final inconsistency score of each pixel is determined by multiplying the initial inconsistency score of each pixel with the multidimensional coupling weight.
3. The method of claim 1, wherein, The determination of the final inconsistency score for each pixel based on the initial inconsistency score of each pixel includes: The external spatial feature map, the internal spatial feature map, and the local temporal feature map are fused locally to obtain a three-modal local fusion feature. The external spatial feature map, the internal spatial feature map, and the local temporal feature map are fused globally to obtain a three-modal global fused feature. Calculate the similarity between the three-modal local fusion features and the three-modal global fusion features, and determine the inconsistency score of the fusion features for each pixel based on the similarity. The initial inconsistency score of each pixel and the inconsistency score of the fused feature of each pixel are weighted and averaged to obtain the final inconsistency score of each pixel.
4. The method of claim 1, wherein, The determination of the visible light mode weight, X-ray mode weight, and infrared mode weight for each pixel based on the visible light mode confidence, X-ray mode confidence, infrared mode confidence, spatial inconsistency score, and temporal inconsistency score includes: determining spatial inconsistency score local mean of positions and timing inconsistency score local mean ; According to the following formula: ,Sure Visible mode weights of position ; According to the following formula: ,Sure Location-based X-ray modal weights ; According to the formula: , Infrared mode weights of position ; in, express Confidence of location in visible light modes. express X-ray modal confidence of location express Infrared modal confidence of location.
5. The method according to claim 2, characterized in that, The multidimensional coupling weights of each pixel are determined based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, visible light surface image, X-ray transmission image, and process sensitivity weights, including: The visible light surface image and the X-ray transmission image are fused to obtain a fused image, and the local grayscale gradient of each pixel extracted from the fused image is used as the spatial weight of each pixel. The multimodal weights of each pixel are determined based on the visible light mode confidence, X-ray mode confidence, and infrared mode confidence of each pixel. The spatial weight, process sensitivity weight, and multimodal weight are weighted and averaged, and the weighted average is processed using the Sigmoid function to obtain the multidimensional coupling weight of each pixel.
6. The method according to claim 1, characterized in that, The determination of spatial inconsistency scores for each pixel based on the external spatial feature map and the internal spatial feature map includes: The external spatial feature map and the internal spatial feature map are respectively subjected to local feature fusion to obtain two-modal local fusion features; The external spatial feature map and the internal spatial feature map are respectively subjected to global feature fusion to obtain two-modal global fused features; Calculate the similarity between the two-modal local fusion features and the two-modal global fusion features, and determine the spatial inconsistency score of each pixel based on the similarity.
7. The method according to claim 1, characterized in that, The step of determining the temporal inconsistency score of each pixel based on the local temporal feature map includes: extracting temporal similarity from the local temporal feature map using a dynamic time warping algorithm; and determining the temporal inconsistency score based on the temporal similarity.
8. A multi-dimensional coupled spatiotemporal anomaly positioning device for die-cast parts, characterized in that, include: The image acquisition module is used to acquire visible light surface images, X-ray transmission images, and infrared thermal imaging image sequences of the die-cast parts; The modal confidence determination module is used to determine the visible light modal confidence of each pixel based on the visible light surface image, the X-ray modal confidence of each pixel based on the X-ray transmission image, and the infrared modal confidence of each pixel based on the infrared thermal imaging image sequence. The feature extraction module is used to extract external spatial feature maps from the visible light image, internal spatial feature maps from the X-ray transmission image, and local temporal feature maps from the infrared thermal imaging image sequence. The spatial inconsistency score determination module is used to determine the spatial inconsistency score of each pixel based on the external spatial feature map and the internal spatial feature map. The temporal inconsistency score determination module is used to determine the temporal inconsistency score of each pixel based on the local temporal feature map. The modal weight determination module is used to determine the visible light modal weight, X-ray modal weight, and infrared modal weight of each pixel based on the visible light modal confidence, X-ray modal confidence, infrared modal confidence, spatial inconsistency score, and temporal inconsistency score of each pixel. The initial inconsistency score determination module is used to add the product of the sum of the visible light mode weights and X-ray mode weights of each pixel and the spatial inconsistency score to the product of the infrared mode weights and the temporal inconsistency score to determine the initial inconsistency score of each pixel. The final inconsistency score determination module is used to determine the final inconsistency score of each pixel based on the initial inconsistency score of each pixel; The abnormal region determination module is used to determine abnormal regions based on the final inconsistency score of each pixel.
9. An electronic device, characterized in that, include: A processor for executing a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the method of any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.