Part anomaly detection method, system, equipment and medium
By acquiring processing videos and process parameter features, using 3D convolutional neural networks and twin comparison technology, combined with an attention classifier, the problem of insufficient accuracy in identifying abnormal parts in existing technologies is solved, achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510825272.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies only perform anomaly identification based on production videos, without considering other factors in the parts production process, resulting in poor accuracy in part anomaly identification.
By acquiring processing videos and process parameter features, the spatiotemporal features are extracted using a 3D convolutional neural network, and anomaly detection is performed by combining twin contrast and attention classifiers, taking into account the process parameters of production equipment and the spatiotemporal characteristics of parts.
It improves the accuracy of identifying abnormal parts, reduces the identification differences caused by the difference between the video and the actual object, adapts to the process fluctuations of different batches of production, and reduces the false detection rate and missed detection rate.
Smart Images

Figure CN120689326A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of component detection, and in particular to a component abnormality detection method, system, device and medium. Background Art
[0002] AI visual inspection involves using AI (artificial intelligence) to analyze video data of a monitored object after visual inspection technology has been used to identify it. For example, in one specific application scenario, AI can be used to analyze production videos of automotive parts for anomalies and generate corresponding anomaly data.
[0003] The prior art discloses a method and system for monitoring abnormalities in automotive parts based on AI visual inspection, comprising: obtaining an original part production video of a target automotive part, and, based on corresponding production links, segmenting the original part production video to output multiple segmented part production sub-videos corresponding to the target automotive part; performing video frame optimization on each segmented part production sub-video to output multiple optimized part production sub-videos corresponding to the target automotive part; and utilizing a target neural network that has undergone an updated operation to perform abnormality identification on each optimized part production sub-video to output multiple part abnormality data corresponding to the target automotive part.
[0004] However, there are often differences between videos and actual objects. This technology only identifies anomalies based on production videos, without considering other factors in the parts production process. It is unable to reduce the identification differences caused by the differences between videos and actual objects, resulting in poor accuracy in identifying parts anomalies. Summary of the Invention
[0005] In order to overcome the problem that the existing technology only performs anomaly identification based on production videos, does not take into account other factors in the component production process, and cannot reduce the identification differences caused by the differences between the video and the actual object, resulting in poor accuracy in component anomaly identification, the present application provides a component anomaly detection method, system, device and medium.
[0006] In a first aspect, in order to solve the above technical problems, the present application provides a component anomaly detection method, comprising: obtaining a processing video and process parameter features for a target component, where the process parameter features are features formed by combining process parameters of a production equipment that processes the target component; Extract features from the processing video to obtain the spatiotemporal characteristics of the target parts; Perform twin comparison on the target component to obtain the twin comparison features of the target component; Using the preset attention classifier, anomaly detection is performed based on twin contrast features, spatiotemporal features, and process parameter features to obtain the target anomaly type of the target component.
[0007] Furthermore, feature extraction is performed on the processing video to obtain the spatiotemporal features of the target parts, including: Use the preset 3D convolutional neural network to extract the original spatiotemporal features of the processed video; Extract key features from the original spatiotemporal features, including key processing area features and key processing period features of the target parts; Use the preset spatiotemporal attention network to enhance key features and obtain enhanced features; Based on the enhanced features and other features in the original spatiotemporal features except the key features, the spatiotemporal features of the target parts are formed.
[0008] Furthermore, twin comparison is performed on the target component to obtain the twin comparison features of the target component, including: Obtain the standard size parameters corresponding to the target parts; A standard digital twin model is constructed based on standard size parameters; Based on the standard digital twin model, twin comparison is performed on the target components to obtain the twin comparison features of the target components.
[0009] Furthermore, twin comparison is performed on the target component based on the standard digital twin model to obtain the twin comparison features of the target component, including: Generate standard 3D point clouds and standard texture images corresponding to target parts based on standard digital twin models; Perform three-dimensional measurement on the target parts to obtain the measured 3D point cloud, and obtain the geometric deviation based on the measured 3D point cloud and the standard 3D point cloud. Photograph the target component to obtain a measured texture image, and calculate the texture similarity based on the measured texture image and the standard texture image; The twin comparison features of the target parts are formed based on texture similarity and geometric deviation.
[0010] Furthermore, using the preset attention classifier, anomaly detection is performed based on twin contrast features, spatiotemporal features, and process parameter features to obtain the target anomaly type of the target component, including: In the preset attention classifier, the twin contrast features, spatiotemporal features, and process parameter features are spliced to obtain fusion features; Process the fused features to obtain the abnormal probability and abnormal type weight of the target component; Using the preset dynamic threshold algorithm, the target abnormality type of the target component is determined based on the abnormality probability and abnormality type weight.
[0011] Furthermore, the attention classifier includes a multi-head attention network, a feedforward neural network, and an activation network; The fused features are processed to obtain the abnormal probability and abnormal type weight of the target component, including: The fused features are processed using a multi-head attention network to obtain the initial attention features and abnormal type attention scores; The initial attention features are nonlinearly transformed using a feedforward neural network to obtain the target attention features; The activation network is used to calculate and process the target attention features to obtain the abnormal probability of the target component; The abnormal type attention score is normalized to obtain the abnormal type weight of the target component.
[0012] Furthermore, a preset dynamic threshold algorithm is used to determine the target abnormality type of the target component based on the abnormality probability and abnormality type weight, including: Obtain multiple historical abnormality determination threshold values and multiple standard deviations of the target component; wherein the historical abnormality determination threshold value mean is the historical abnormality threshold value mean corresponding to the abnormality type in the abnormal range included in the target component, and the standard deviation is the historical standard deviation corresponding to the abnormality type in the abnormal range included in the target component; Based on the mean of the historical anomaly determination threshold and the corresponding standard deviation, the corresponding dynamic anomaly determination threshold is calculated; Determine the corresponding abnormal situation based on the dynamic abnormality judgment threshold and abnormality probability; When the abnormal situation is that an abnormality exists and the corresponding abnormality type weight is greater than or equal to the preset weight, the abnormality type that matches the abnormality type weight is determined as the target abnormality type of the target component.
[0013] In a second aspect, the present application also provides a component anomaly detection system, comprising: An acquisition module is used to obtain a processing video and process parameter characteristics of a target component, where the process parameter characteristics are a combination of process parameters of a production equipment used to process the target component; Feature extraction module, used to extract features from processing videos and obtain the spatiotemporal features of target parts; The twin comparison module is used to perform twin comparison on the target component to obtain the twin comparison features of the target component; The anomaly detection module is used to use a preset attention classifier to perform anomaly detection based on twin contrast features, spatiotemporal features, and process parameter features to obtain the target anomaly type of the target component.
[0014] In a third aspect, the present application also provides a computing device comprising a memory, a processor, and a program stored in the memory and running on the processor, wherein when the processor executes the program, the steps of a component abnormality detection method as described above are implemented.
[0015] In a fourth aspect, the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device executes the steps of a component abnormality detection method.
[0016] The beneficial effects of the present application are: by extracting features from the processing video of the target component to obtain spatiotemporal features, performing twin comparison on the target component to obtain twin comparison features, and using a preset attention classifier to perform anomaly detection based on the twin comparison features, spatiotemporal features and process parameter features for the target component to obtain the target anomaly type of the target component. In this way, when identifying the target anomaly type of the target component, not only the processing video of the target component is considered, but also the process parameter features composed of the process parameters of the production equipment that processes the target component, and the spatiotemporal features of the target component during processing are considered, so that the influencing factors of multiple sources can be taken into account in the target component identification process, thereby reducing the recognition differences caused by only identifying the processing video, and thus improving the accuracy of component anomaly identification, so as to improve the accuracy of the obtained target anomaly type. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a component abnormality detection method according to an exemplary embodiment of the present application; Figure 2 This is a schematic diagram of the structure of a spatiotemporal attention network in an exemplary embodiment of the present application; Figure 3 This is a flowchart of anomaly detection using an attention classifier in an exemplary embodiment of the present application; Figure 4 This is a schematic diagram of a structural system architecture of an exemplary embodiment of the present application, using the provided component anomaly detection method; Figure 5 The figure is a schematic structural diagram of a component abnormality detection system according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0018] The following examples are provided to further explain and supplement the present application and do not constitute any limitation to the present application.
[0019] Existing technologies rely solely on image connected domain analysis of 2D video frames, failing to integrate 3D spatial features with the temporal dynamics of the production process. This leads to missed detection issues for complex geometric components (such as curved surfaces and 3D assembly structures). For example, the detection accuracy of 3D dimensional anomalies, such as the edge curvature deviation of automotive wheels and the coaxiality of bearing inner and outer rings, is insufficient, with a missed detection rate as high as 15% to 20%. Furthermore, relying on discrete calculations of connected domains in images fails to effectively capture the temporal correlations between production steps (such as dynamic parameter changes between stamping and polishing processes), resulting in weak detection of intermittent anomalies (such as surface microcracks caused by transient vibration). Furthermore, existing solutions fail to integrate production equipment sensor data (such as pressure, temperature, and vibration parameters) with prior knowledge of component design (such as CAD models). Relying solely on visual data for anomaly detection, they struggle to cope with noise interference in industrial environments (such as lighting fluctuations and equipment vibration). For example, when the light intensity in the workshop fluctuates by more than 20%, the false detection rate rises to 12%. In addition, the model threshold is fixed and cannot be dynamically adjusted according to the process parameters of different batches of production. It has poor adaptability to the flexible detection of multiple models of parts and requires frequent manual retraining of the model, which is inefficient.
[0020] In order to solve the above problems, the embodiments of the present application provide a component abnormality detection method, system, device and medium, and these embodiments will be described in detail below.
[0021] The component anomaly detection method provided in the embodiments of the present application can be specifically executed by a server. It should be noted that the server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, and this is not a limitation.
[0022] See also Figure 1 , Figure 1 A component abnormality detection method is shown as an exemplary embodiment of the present application. Figure 1 As shown, the present application provides a component abnormality detection method, comprising: S11, obtaining a processing video and process parameter features for a target component, where the process parameter features are a combination of process parameters of a production equipment used to process the target component; S12, extracting features from the processing video to obtain the spatiotemporal features of the target parts; S13, performing twin comparison on the target component to obtain a twin comparison feature of the target component; S14, using a preset attention classifier, performs anomaly detection based on twin contrast features, spatiotemporal features, and process parameter features to obtain the target anomaly type of the target component.
[0023] The component anomaly detection method of the embodiment provided in the present application extracts features from the processing video of the target component to obtain spatiotemporal features, performs twin comparison on the target component to obtain twin comparison features, and uses a preset attention classifier to perform anomaly detection based on the twin comparison features, spatiotemporal features, and process parameter features for the target component to obtain the target anomaly type of the target component. In this way, when identifying the target anomaly type of the target component, not only the processing video of the target component is considered, but also the process parameter features composed of the process parameters of the production equipment that processes the target component, and the spatiotemporal features of the target component during processing are considered, so that the influencing factors of multiple sources can be taken into account in the target component identification process, thereby reducing the recognition differences caused by only identifying the processing video, and thus improving the accuracy of component anomaly identification, so as to improve the accuracy of the obtained target anomaly type.
[0024] In this embodiment, processing videos and process parameter features are acquired through a 3D vision unit, which includes a 3D camera, an industrial area scan camera, and a production data interface. This unit is used to collect visual data (processing videos and subsequent measured 3D point clouds) and process parameters of components. The industrial area scan camera uses a Baslerace 25-megapixel camera paired with a Computar 8mm industrial lens. It supports dynamic ROI configuration to reduce data redundancy and is used to synchronize with the 3D camera to capture target components during processing, obtaining depth and RGB images of the target components. The RGB images are then enhanced using adaptive histogram equalization (CLAHE) to enhance the edge contrast of RGB images in low-light environments. A processing video consisting of the depth image and the enhanced RGB image of the target component is obtained. The production data interface connects to a Siemens PLC (Programmable Logic Controller) device via the OPCUA (Open Platform Communications Unified Architecture) protocol, collecting real-time process parameters such as press pressure, injection molding machine temperature, and conveyor speed from the production equipment processing the target components, with a data synchronization error of ≤10ms.
[0025] Optionally, feature extraction is performed on the processing video to obtain the spatiotemporal features of the target parts, including: Use the preset 3D convolutional neural network to extract the original spatiotemporal features of the processed video; Extract key features from the original spatiotemporal features, including key processing area features and key processing period features of the target parts; Use the preset spatiotemporal attention network to enhance key features and obtain enhanced features; Based on the enhanced features and other features in the original spatiotemporal features except the key features, the spatiotemporal features of the target parts are formed.
[0026] In the embodiment provided by the present application, a preset 3D convolutional neural network is used to extract the original spatiotemporal features of the processed video, and the key features in the original spatiotemporal features are extracted and enhanced using a preset spatiotemporal attention network to obtain enhanced features, thereby forming the spatiotemporal features of the target components. This facilitates subsequent abnormality identification, and allows the enhanced features in the spatiotemporal features to be identified, thereby improving the accuracy of abnormality identification.
[0027] In this embodiment, the key area corresponding to the key processing area feature refers to the structural part of the component that is prone to abnormalities in the spatial dimension. The key area includes areas with complex geometric shapes, assembly interfaces or areas prone to surface defects. For example, the area with complex geometric shapes includes the edge of the automobile hub: as a high curvature area, it is prone to dents, protrusions or dimensional deviations. The spatiotemporal attention network is used to enhance the point cloud curvature features and image edge texture of the area (the edge of the automobile hub) through the spatial attention network it includes: through W s F 3D Matrix operations highlight spatial distribution anomalies in edge point clouds, identifying key processing area features. For example, the assembly interface, including the contact surface between the inner and outer rings of an engine bearing, is prone to coaxial deviation during assembly. The spatial attention network focuses on the 3D point cloud alignment accuracy in this area (such as the positional residual after ICP registration) and the seal ring installation texture in the RGB image to identify key processing area features.
[0028] The key periods corresponding to the key processing period features refer to the process stages in the production process that are prone to abnormalities. The key periods include dynamic process periods such as stamping, welding, pressing, and conveyor belts. For example, the peak period of the stamping process: when the pressure sensor detects an instantaneous pressure fluctuation (such as exceeding the rated value by 10%), the spatiotemporal attention network captures the video frame sequence of this period through the temporal attention network, identifies surface microcracks caused by mold wear (such as a sudden drop in SSIM texture similarity), and obtains the corresponding key processing period features. For another example, the acceleration / deceleration stage of the conveyor belt: due to vibration, the displacement of parts may occur. The temporal attention network focuses on the temporal changes of the point cloud of 16 consecutive frames (T≥16) in this period to detect whether there is a geometric deviation caused by loose assembly, and obtains the corresponding pipe processing period features.
[0029] In an exemplary embodiment provided by the present application, first, the processed video input to the 3D convolutional neural network (3D-CNN) is a continuous video sequence, and the original spatiotemporal features output by the 3D convolutional neural network are the spatiotemporal feature tensor F 3D ∈R T ×H×W×C , supports capturing dynamic features of the time dimension T≥16 frames.
[0030] The 3D convolutional neural network uses an 8-layer 3D convolutional neural network with the following parameters: Input layer: spatiotemporal tensor F of processed video 3D ∈R 16×224×224×3 (16 frames of continuous video, single frame resolution 224×224, 3-channel RGB).
[0031] Convolutional layer: The first four layers use 3×3×3 convolution kernels (stride size 1×1×1), and the last four layers use 2×2×2 pooling kernels (stride size 2×2×2). Each layer is followed by BatchNorm and ReLU activation functions.
[0032] Output layer: spatiotemporal feature tensor F of original spatiotemporal features st ∈R 2×28×28×256 , capturing the spatiotemporal dynamic features of a 16-frame sequence.
[0033] Second, see Figure 2 , Figure 2 This is a schematic diagram of the structure of the spatiotemporal attention network in an exemplary embodiment of this application, as shown in FIG. Figure 2 As shown in Figure 2, the spatiotemporal attention network includes a spatial attention network and a temporal attention network. Spatial attention network: 3D spatiotemporal feature tensor F of the original spatiotemporal features output by the 3D convolutional neural network (3D-CNN) 3D Perform global average pooling and maximum pooling, and then generate a spatial attention map A after splicing through 1×1 convolution and Sigmoid activation (S-shaped growth curve) s ∈R H×W×1 , the calculation formula of the spatial attention map is as follows: A s =σ(W s F 3D +b s ); Among them, F 3D The 3D spatiotemporal feature tensor of the original spatiotemporal features extracted by the 3D convolutional neural network contains the feature information of automobile parts in the spatial and temporal dimensions. s is a learnable weight matrix used to 3D Perform linear transformation so that the model can learn the importance of different features based on the data. σ is the Sigmoid activation function, and b sis the bias vector used to adjust the result of the linear transformation.
[0034] Temporal Attention Network: 3D spatiotemporal feature tensor F 3D Perform cross-frame feature association, extract time series dependencies through bidirectional LSTM (Long Short-Term Memory, long short-term memory network), and output the time attention map A after Softmax normalization (normalized exponential function) t ∈R T×1×1 , to focus on the key frames where the anomaly occurs. The calculation formula of the temporal attention map is: in, F 3D The transpose of W t is a learnable weight matrix, Perform linear transformation, b t is the bias vector.
[0035] For the spatial attention map A s and temporal attention map A t Perform feature weighting to obtain enhanced feature F weighted , same as F_weighted.
[0036] Optionally, twin comparison is performed on the target component to obtain twin comparison features of the target component, including: Obtain the standard size parameters corresponding to the target parts; A standard digital twin model is constructed based on standard size parameters; Based on the standard digital twin model, twin comparison is performed on the target components to obtain the twin comparison features of the target components.
[0037] In the embodiment provided by the present application, a standard digital twin model is constructed based on the standard size parameters corresponding to the target component, and a twin comparison is performed on the target component to obtain a twin comparison feature that includes the production differences of the target component. This facilitates the subsequent abnormality identification based on the twin comparison feature, while being able to take into account the abnormal impact caused by the production differences of the target component, thereby improving the accuracy of abnormality identification.
[0038] In this embodiment, the standard size parameters corresponding to the target component are obtained from the CAD design drawing of the target component. The standard size parameters include standard geometric dimensions and standard textures. The standard digital twin model is constructed based on the standard size parameters. The specific steps are as follows: CAD modeling: Use CATIAV5 (Computer-Aided Three-Dimensional Interactive Application) to build 3D models of standard parts according to standard geometric dimensions and export them to STL format files with an accuracy of ≤0.01mm; Texture mapping: Use Blender (open source 3D modeling software) to map standard textures to the corresponding positions of the 3D model of standard parts according to their corresponding standard coordinates. Generate standard textures on the 3D model of standard parts to obtain a standard digital twin model to support multi-channel texture comparison such as diffuse reflection and roughness.
[0039] Optionally, a twin comparison is performed on the target component based on the standard digital twin model to obtain the twin comparison features of the target component, including: Generate standard 3D point clouds and standard texture images corresponding to target parts based on standard digital twin models; Perform three-dimensional measurement on the target parts to obtain the measured 3D point cloud, and obtain the geometric deviation based on the measured 3D point cloud and the standard 3D point cloud. Photograph the target component to obtain a measured texture image, and calculate the texture similarity based on the measured texture image and the standard texture image; Twin comparison features of target parts are formed based on texture similarity and geometric deviation.
[0040] In the embodiment provided by the present application, the geometric deviation is calculated based on the standard 3D point cloud corresponding to the target component generated based on the standard digital twin model and the measured 3D point cloud, and the texture similarity is calculated based on the standard texture image corresponding to the target component generated based on the standard digital twin model and the measured texture image obtained by shooting, so that the formed twin comparison features include two types of production differences, namely geometric size difference and texture difference, which is convenient for subsequent abnormality identification based on the twin comparison features. The influence of geometric size abnormalities and texture abnormalities caused by the production differences of the target components can be taken into account, thereby improving the accuracy of abnormality identification.
[0041] In an exemplary embodiment provided by this application, first, a standard 3D point cloud and a standard texture image corresponding to the target component are generated based on the standard digital twin model. The specific steps are: the surface triangular mesh of the standard digital twin model is parsed by the Python OCC (Open Cascade Community Edition, an open source CAD software development platform based on Python) library to generate a standard 3D point cloud containing the standard geometric dimensions of the target component (the density can be 100 points / mm). 2), and a standard texture image containing the standard texture of the target component.
[0042] Perform three-dimensional measurement on the target parts to obtain the measured 3D point cloud. The specific steps are as follows: Use the laser scanner set on the 3D camera and the rotating worktable to realize the full surface point cloud collection of the parts, with a sampling frequency of 100Hz and a point cloud density of ≥50 points / mm 2 , obtaining a measured 3D point cloud. The measured 3D point cloud is preprocessed using PCL (PointCloudLibrary) before subsequent geometric deviation calculations. This preprocessing includes outlier removal (radius filtering, radius 0.2mm), voxel grid downsampling (voxel size 0.5mm), and bilateral filtering for noise reduction (spatial standard deviation 0.3mm, range standard deviation 0.15).
[0043] The geometric deviation obtained based on the measured 3D point cloud and the standard 3D point cloud is the spatial position deviation. The calculation process of the spatial position deviation is as follows: The point-to-plane ICP algorithm (50 iterations, convergence threshold 0.01 mm) is used to calculate the rotation matrix R and translation vector t between the measured 3D point cloud and the standard 3D point cloud (R, t). The calculation formula is as follows: Among them, P real,i is the i-th point of the measured 3D point cloud, P std,i is the i-th point of the standard 3D point cloud, R is the rotation matrix, and t is the translation vector; the goal of this formula is to find the optimal rotation matrix R and translation vector t so that the measured 3D point cloud P real , after rotation and translation, it is compared with the standard 3D point cloud P std , the sum of squared errors between them is minimized. This optimization problem is solved by the iterative closest point (ICP) algorithm to achieve registration between the measured point cloud and the theoretical point cloud; Convert (R, t) to the x / y / z axis deviation value (unit: mm) in the Cartesian coordinate system to form the 3D spatial position deviation [ΔP x , ΔPy, ΔP z ]. The calculation accuracy of spatial position deviation can be better than 0.1mm.
[0044] In addition, geometric deviation can also be shape deviation (such as flatness, roundness), direction deviation (such as perpendicularity, coaxiality), etc.
[0045] Secondly, the target component is photographed to obtain a measured texture image. This measured texture image is then enhanced for subsequent texture similarity calculations. Image enhancement uses adaptive histogram equalization (CLAHE) to improve edge contrast in low-light environments. Video frame synchronization is achieved through OpenCV (OpenSource Computer Vision Library), aligning each frame of the measured texture image with the processed video based on timestamps (time error ≤ 5ms).
[0046] The texture similarity is calculated based on the measured texture image and the standard texture image. The specific steps are as follows: The measured texture image T is calculated based on the SSIM (Structural Similarity Index) algorithm. real With the standard texture image T std The structural similarity between them is output in the range of [0,1], forming a 1-dimensional feature vector [S].
[0047] The formula for calculating texture similarity using the SSIM algorithm is as follows: Among them, SSIM(T real , T std ) represents texture similarity, μ real Represents the measured texture image T real The pixel mean, μ std Represents the standard texture image T std The pixel mean, Represents the measured texture image T real The pixel variance of represents the pixel variance of the standard texture image, σ real,std Represents the pixel covariance between the measured texture image and the standard texture image, C1=(k1L) 2 、C2=(k2L) 2 , used to avoid the denominator being zero (L is the dynamic range of pixel values, such as 255; k1≈0.01, k2≈0.03, k1 and k2 are both empirical coefficients). The detection resolution of texture similarity can be 0.05mm 2 .
[0048] Then, twin contrast features of the target parts are formed based on texture similarity and geometric deviation. The specific steps are as follows: the two types of features, texture similarity and geometric deviation, are spliced into a 4-dimensional vector [ΔP x ,ΔPy,ΔP z ,S], as the twin comparison feature F of the target component twist .
[0049] Optionally, a preset attention classifier is used to perform anomaly detection based on twin contrast features, spatiotemporal features, and process parameter features to obtain target anomaly types of target parts, including: In the preset attention classifier, the twin contrast features, spatiotemporal features, and process parameter features are spliced to obtain fusion features; The formula for this feature splicing is: F fusion =Concat(F st ,F param ,F twist )·W fusion ; Among them, F fusion is the fusion feature F twist is the twin contrast feature, F st is the spatiotemporal feature, F param is the process parameter characteristic, W fusion To fuse the weight matrix, the concatenated feature vector is linearly transformed to map the multimodal features into a unified feature space for subsequent abnormality recognition model (attention classifier) processing; Concat is a concatenation operation to convert the spatiotemporal features F st , process parameter characteristics F param Compare the twin features F twist Splice together along a certain dimension to form a new feature vector; Process the fused features to obtain the abnormal probability and abnormal type weight of the target component; Using the preset dynamic threshold algorithm, the target abnormality type of the target component is determined based on the abnormality probability and abnormality type weight.
[0050] In the embodiment provided by the present application, in the preset attention classifier, the fusion features obtained by splicing the twin contrast features, spatiotemporal features and process parameter features are processed to obtain the abnormal probability and abnormal type weight of the target component, and the target abnormal type of the target component is determined based on the abnormal probability and abnormal type weight using the preset dynamic threshold algorithm. In this way, when identifying the target abnormal type of the target component, not only the processing video of the target component is considered, but also the process parameter features composed of the process parameters of the production equipment that processes the target component, and the spatiotemporal features of the target component during processing are considered. The abnormal probability and abnormal type weight of the three types of features are also considered, so that the recognition influence and common recognition influence of multiple source features can be taken into account in the target component identification process, thereby further reducing the recognition difference caused by only identifying the processing video, and further improving the accuracy of component abnormality identification. Among them, the preset attention classifier can be a Transformer classifier.
[0051] Optionally, the attention classifier includes a multi-head attention network, a feedforward neural network, and an activation network; The fused features are processed to obtain the abnormal probability and abnormal type weight of the target component, including: The fused features are processed using a multi-head attention network to obtain the initial attention features and abnormal type attention scores; The initial attention features are nonlinearly transformed using a feedforward neural network to obtain the target attention features; The activation network is used to calculate and process the target attention features to obtain the abnormal probability of the target component; The calculation formula for abnormal probability is as follows: p(anomaly)=Softmax(F tranformer W class ); Among them, p (anomaly) is the probability of anomaly, ranging from [0,1]. The larger the value, the higher the possibility of anomaly. transformer represents the target attention feature, W class represents the anomaly weight matrix; The abnormal type attention score is normalized to obtain the abnormal type weight of the target component.
[0052] In the embodiment provided by the present application, first, the fusion features are processed using a multi-head attention network to obtain initial attention features and abnormal type attention scores, and the initial attention features are nonlinearly transformed using a feedforward neural network to obtain target attention features, thereby realizing nonlinear fusion of cross-modal features. Then, the target attention features are calculated and processed using an activation network to obtain the abnormal probability of the target component, and the abnormal type attention scores are normalized to obtain the abnormal type weight of the target component, so that when the target abnormal type of the target component is determined based on the abnormal probability and abnormal type weight, the recognition influence of the multi-source features and the common recognition influence can be taken into account, thereby further reducing the recognition differences caused by only recognizing the processed video, and further improving the accuracy of component abnormality recognition.
[0053] In this embodiment, a multi-head attention network is used to process the fused features to obtain anomaly type attention scores. The specific steps are: pre-acquiring the abnormal range of the target component, which includes multiple abnormality types; measuring the similarity / association strength between each abnormality type and the fused features in the multi-head attention network; and using this similarity / association strength as the anomaly type attention score. In other words, there are as many possible anomaly types as there are corresponding anomaly type attention scores.
[0054] See also Figure 3 , Figure 3 This is a flow chart of an exemplary embodiment of the present application, in which anomaly detection is performed using an attention classifier. Figure 3 As shown in the figure, the attention classifier includes a preprocessing network, a feature splicing layer, a feature fusion core network, a fully connected layer, and a Softmax classification layer (activation network). The preprocessing network includes a pooling layer, a fully connected layer, and a dimension expansion layer. The feature fusion core network includes a multi-head attention network (multi-head self-attention mechanism), a feedforward neural network, and a normalization and residual connection layer. In the preprocessing layer, the pooling layer is used to perform global average pooling on the spatiotemporal features and then extract the time series features. The fully connected dimensionality reduction layer is used to perform a fully connected transformation on the process parameter features. The dimension expansion layer is used to expand the dimension of the twin contrast features. In the feature splicing layer, the three types of features output by the preprocessing layer are spliced to obtain the fused features. In the feature fusion core network, the multi-head attention network is used to process the fused features to obtain the initial attention features and the abnormal type attention scores; the feedforward neural network is used to perform nonlinear transformation on the initial attention features to obtain the target attention features; the normalization and residual connection layer is used to normalize and perform residual connection processing on the target attention features to obtain the processed target attention features; In the fully connected layer, the processed target attention features are reduced in dimension to obtain the reduced-dimensional target attention features; in the Softmax classification layer (activation network), the reduced-dimensional target attention features are calculated and processed to obtain the abnormal probability of the target parts.
[0055] Optionally, a preset dynamic threshold algorithm is used to determine a target abnormality type of a target component based on the abnormality probability and the abnormality type weight, including: Obtain multiple historical abnormality determination threshold values and multiple standard deviations of the target component; wherein the historical abnormality determination threshold value mean is the historical abnormality threshold value mean corresponding to the abnormality type in the abnormal range included in the target component, and the standard deviation is the historical standard deviation corresponding to the abnormality type in the abnormal range included in the target component; Based on the mean of the historical anomaly determination threshold and the corresponding standard deviation, the corresponding dynamic anomaly determination threshold is calculated; The calculation formula for the dynamic anomaly judgment threshold is as follows: θ(t)=μ(t)+kσ(t); Where θ(t) is the dynamic anomaly determination threshold, μ(t) is the mean of the historical anomaly determination threshold, σ(t) is the standard deviation, and k is the confidence coefficient (k ≥ 1.5) (e.g., k = 2 corresponds to a 95% confidence interval). Determine the corresponding abnormal situation based on the dynamic abnormality judgment threshold and abnormality probability; When the abnormal situation is that an abnormality exists and the corresponding abnormality type weight is greater than or equal to the preset weight, the abnormality type that matches the abnormality type weight is determined as the target abnormality type of the target component.
[0056] In the embodiment provided by the present application, a preset dynamic threshold algorithm is used, and the dynamic abnormality judgment threshold and abnormal probability calculated based on the historical abnormality judgment threshold mean and standard deviation are used to determine the abnormal situation of the target component. When the abnormal situation is that there is an abnormality, and the corresponding abnormal type weight is greater than or equal to the preset weight, it indicates that the probability of the target abnormal type being greater than or equal to the abnormal type weight corresponding to the preset weight is the largest, and the abnormal type that matches the abnormal type weight is determined as the target abnormal type of the target component, so that the recognition influence (abnormal type weight) and common recognition influence (abnormal situation) of the multi-source features can be taken into account in the target component identification process, thereby further reducing the recognition difference caused by only identifying the processing video, and thus further improving the accuracy of component abnormality identification. At the same time, after determining the target abnormal type, it can be traced back and located directly based on the target abnormal type, and the corresponding abnormal features (abnormal spatiotemporal features (mainly abnormal enhancement features), and / or, abnormal twin contrast features, and / or, abnormal process parameter features) can be determined, so as to improve the efficiency of component abnormal location. Among them, the dynamic abnormality judgment threshold can be 0.95.
[0057] In this embodiment, the dynamic abnormality determination threshold can be used to update the model parameters in the attention classifier in real time, so that the attention classifier can adapt to the process fluctuations of parts produced in different batches, and the false detection rate is ≤5%.
[0058] See also Figure 4 , Figure 4 This is a schematic diagram of the structural system architecture of the component abnormality detection method provided in an exemplary embodiment of the present application, as shown in FIG. Figure 4 As shown in the figure, the structural system architecture includes data acquisition layer, data processing layer, feature extraction layer, model fusion layer and application layer.
[0059] The data acquisition layer includes 3D laser scanners (3D cameras), RGB industrial cameras (industrial area array cameras), and PLC / sensor data access (accessed through the production data interface). The data processing layer includes high-cluster point removal / voxel downsampling, CLAHE / frame synchronization, and timestamp alignment / normalization. The feature extraction layer includes 3D convolutional neural networks, ST-Attention (spatiotemporal attention network) focusing on key frames / regions, ICP registration / SSIM texture comparison (twin comparison). The model fusion layer includes Transformer classifiers and RMSprop adaptive optimization. The application layer is used to output the target abnormality type of the target parts: surface defects / dimensional deviations / assembly misalignments, and control the MES system / equipment (manufacturing execution system / equipment) based on the target abnormality type to reduce abnormalities in parts produced subsequently.
[0060] In an exemplary embodiment provided by this application, the specific steps of applying the provided component abnormality detection method are as follows: 1. Feature preprocessing: Spatiotemporal characteristics F st : 256-dimensional vector F after global average pooling after output by 3D convolutional neural network and spatiotemporal attention network st ∈R 256 .
[0061] Process parameter characteristics F param : The feature map composed of the process parameters (such as pressure, temperature, etc.) of the production equipment detected by the sensor is mapped into a 256-dimensional vector through the fully connected layer.
[0062] Twin contrast feature F twist : A 4-dimensional vector composed of geometric deviation (3D: x / y / z axis deviation value) + texture similarity (1D).
[0063] 2. Cross-modal splicing: The above three types of features are concatenated into a 516-dimensional vector F through the Concat operation fusion :F fusion =[F st , F param , F twist ], and then through the fusion weight matrix W fusion Map to a unified feature space (such as reducing the dimension to 256 dimensions).
[0064] 3. The attention classifier can be a Transformer classifier (8-layer encoder, 8 multi-head attention heads). Using the preset attention classifier, anomaly detection is performed based on twin contrast features, spatiotemporal features, and process parameter features to obtain the target anomaly type of the target component. The specific steps are as follows: Multi-head attention network: Capture cross-modal correlations (such as the correlation between geometric deviations and pressure parameters) through 8 attention heads, The formula is: MultiHead(Q,K,V)=Concat(head1,...,head h )W O ; Among them, Q, K, and V are the query matrix, key matrix, and value matrix, which are linear transformations of the fused features.
[0065] Feedforward Neural Network (FFN): Perform a nonlinear transformation on the output x of the multi-head attention network. The formula is: FFN(x) = σ(xW1+b1)W2+b2, where σ is the activation function, W1 is the first-layer weight matrix, b1 is the bias parameter corresponding to the weight matrix W1, W2 is the second-layer weight matrix, and b2 is the bias parameter corresponding to the weight matrix W2.
[0066] Activate the network: Calculate the anomaly probability through the activation network (Softmax layer), the calculation formula is: p(anomaly)=Softmax(F transformer W class ), the output range is [0,1], and the larger the value, the higher the possibility of abnormality.
[0067] 4. Using the preset dynamic threshold algorithm, based on the abnormality probability and abnormality type weight, determine the target abnormality type of the target component. The specific steps are as follows: Threshold calculation: Based on the real-time statistics of the historical anomaly judgment threshold mean μ(t) and standard deviation σ(t) of the last 1000 qualified samples, the dynamic anomaly judgment threshold is: θ(t) = μ(t) + kσ(t), where k ≥ 1.5 is the confidence coefficient (e.g., k = 2 corresponds to a 95% confidence interval).
[0068] Abnormality judgment rules: ①Threshold comparison: If the abnormal probability p≥θ(t), it means there is an abnormality, and the abnormality alarm is triggered; if p<θ(t), it means there is no abnormality, and it is judged as qualified.
[0069] ② Abnormal type positioning: through the output layer weight matrix W of the Transformer classifier class Associated with specific abnormality categories (such as "surface cracks", "assembly misalignment", and "size deviation"), each category corresponds to a dimension of the Softmax output vector.
[0070] Combined with feature importance analysis (such as attention weight visualization), locate the source of the anomaly: If the weight of the abnormality type of the geometric deviation is greater than the preset value, the target abnormality type is determined to include size or assembly abnormality; If the anomaly type weight of the texture similarity is greater than a preset value, the target anomaly type is determined to include surface defects; If the abnormality type weight of the process parameter feature is greater than a preset value, it is determined that the target abnormality type includes equipment operation abnormality (such as pressure exceeding the limit).
[0071] 5. Dynamic Update Mechanism: After every 1,000 samples are tested, μ(t) and σ(t) are recalculated to adapt to process fluctuations (such as drift in the characteristics of qualified samples due to mold wear) and ensure that the dynamic anomaly determination threshold always reflects the current production status. The RMSprop (Root Mean Square Propagation) optimizer is used to update the model parameters of the attention classifier (Transformer classifier). The learning rate starts at 0.001 and decays by 0.95 every 5,000 iterations to accommodate process drift during production.
[0072] In another exemplary embodiment provided by this application, the target component is a wheel hub, and the specific steps of applying the provided component abnormality detection method are as follows: When the wheel hub passes through the inspection station at a speed of 0.5m / s on the conveyor belt, the 3D camera collects 100 frames of measured 3D point cloud and measured texture images of the wheel hub per second. The industrial area array camera simultaneously collects 25 frames of processing video of the wheel hub in RGB image format. The process parameter characteristics of the stamping machine (production equipment), including pressure data (1000bar±5%), are accessed in real time.
[0073] Feature processing: The spatiotemporal attention network focuses on the video sequence of the stamping process in the processing video, detects surface microcracks (length ≥ 1mm) corresponding to instantaneous pressure fluctuations (> 10%), and forms spatiotemporal features.
[0074] The measured 3D point cloud and texture image were compared with the corresponding standard digital twin model to obtain the twin comparison features. The feature of the target anomaly type located in this experiment is: rim edge curvature deviation, and the corresponding dynamic anomaly judgment threshold is ±0.3mm (exceeding this threshold is considered an edge depression / convexity).
[0075] Anomaly identification: When the attention classifier processes the fusion features of a frame image of the processing video, which consists of spatiotemporal features, twin contrast features, and process parameter features, the obtained anomaly probability p (anomaly) = 0.92 (> threshold 0.85) and the pressure curve shows a sudden drop (15%), then the target anomaly type is determined to be an edge defect caused by mold wear, triggering a real-time production line shutdown warning.
[0076] Detection performance: ① Dimension detection accuracy: rim diameter deviation ≤ 0.1mm, bolt hole position deviation ≤ 0.08mm; ② Defect recognition speed: single hub detection time ≤ 200ms, meeting the production line's detection cycle of 50 pieces / minute.
[0077] In another exemplary embodiment provided by this application, the target component is an engine bearing, and the specific steps of applying the component abnormality detection method provided are as follows: Obtain engine bearing machining videos, process parameter characteristics, measured 3D point clouds, and measured texture images.
[0078] Feature extraction is performed on the processing video to obtain the spatiotemporal characteristics of the engine bearing.
[0079] Digital Twin Model: A standard 3D digital twin model of the inner and outer rings of an engine bearing was constructed, with a coaxiality tolerance of ±0.05mm and a radial runout tolerance of ±0.03mm. The measured 3D point cloud and texture image were compared with the standard digital twin model to obtain twin comparison features. This experiment identified the target anomaly types as angular deviation in the inner and outer ring assembly and seal ring misalignment.
[0080] Then, when the attention classifier processes the fusion features of a frame of the processing video, which are composed of spatiotemporal features, twin contrast features, and process parameter features, the target anomaly type obtained by anomaly identification is bearing assembly anomaly, including support for identification of bearing tilt (angle deviation > 0.1°), missing seals (texture similarity < 0.9), and abnormal press pressure (< 70kN or > 130kN). Among them, the pressure curve of the bearing press station (normal range 80-120kN) is also combined during anomaly identification to eliminate misjudgments caused by vibration and noise.
[0081] From the above experiments, we can see that the accuracy of assembly anomaly detection is 99.2%, which is 5 times higher than manual detection efficiency, and the missed detection rate is reduced from 8% to 0.8%.
[0082] The component anomaly detection method of this application, through the deep integration of 3D vision and digital twin technology, completes the combination of 3D-CNN and spatiotemporal attention, enabling the system to capture the spatial geometric features and temporal dynamic changes of components, improve the accuracy of anomaly detection of complex components, realize closed-loop verification of theoretical design and actual production, significantly reduce the dependence on large amounts of labeled data, improve the efficiency of model training in small sample scenarios, and solve the accuracy bottleneck of existing technologies for complex structure detection. At the same time, this method supports the joint modeling of production process parameters and visual features through the designed multimodal fusion network and dynamic threshold algorithm, and dynamically adjusts the anomaly judgment threshold through the statistical characteristics of real-time data, so that the system reduces the false detection rate in complex environments such as lighting changes and equipment vibration. At the same time, it supports adaptive detection of multiple models of components, and can switch detection objects without retraining the model, shortening the production line switching time and improving the flexibility of industrial detection. In this way, through the standardization of hardware selection, the specification of algorithm parameters, and the instantiation of scenario applications, the full process from data collection to anomaly recognition is implemented. Compared with existing technologies, it has significant advantages in complex parts detection accuracy, multimodal fusion efficiency, production line adaptability, etc., providing a replicable AI visual inspection solution for the intelligent manufacturing of automotive parts.
[0083] See also Figure 5 , Figure 5 A component abnormality detection system is shown as an exemplary embodiment of the present application. Figure 5 As shown, the present application provides a component anomaly detection system 500, comprising: An acquisition module 501 is used to acquire a processing video and process parameter characteristics for a target component, where the process parameter characteristics are a combination of process parameters of a production equipment used to process the target component; Feature extraction module 502, used to extract features from the processing video to obtain the spatiotemporal features of the target parts; The twin comparison module 503 is used to perform twin comparison on the target component to obtain the twin comparison features of the target component; the anomaly detection module 504 is used to use a preset attention classifier to perform anomaly detection based on the twin comparison features, spatiotemporal features and process parameter features to obtain the target anomaly type of the target component.
[0084] The component anomaly detection system 500 of the embodiment provided by the present application uses a feature extraction module 502 to extract features from the processing video of the target component obtained by the acquisition module 501 to obtain spatiotemporal features, uses a twin comparison module 503 to perform twin comparison on the target component to obtain twin comparison features, and uses a preset attention classifier in the anomaly detection module 504 to perform anomaly detection based on the twin comparison features, spatiotemporal features and the process parameter features of the target component obtained by the acquisition module 501 to obtain the target anomaly type of the target component. In this way, when identifying the target anomaly type of the target component, not only the processing video of the target component is considered, but also the process parameter features composed of the process parameters of the production equipment that processes the target component, and the spatiotemporal features of the target component during processing are considered, so that the influencing factors of multiple sources can be taken into account in the target component identification process, thereby reducing the recognition differences caused by only identifying the processing video, and thus improving the accuracy of component anomaly identification, so as to improve the accuracy of the obtained target anomaly type.
[0085] Optionally, the feature extraction module 502 is specifically configured to: Use the preset 3D convolutional neural network to extract the original spatiotemporal features of the processed video; Extract key features from the original spatiotemporal features, including key processing area features and key processing period features of the target parts; Use the preset spatiotemporal attention network to enhance key features and obtain enhanced features; Based on the enhanced features and other features in the original spatiotemporal features except the key features, the spatiotemporal features of the target parts are formed.
[0086] Optionally, the twin comparison module 503 is specifically configured to: Obtain the standard size parameters corresponding to the target parts; A standard digital twin model is constructed based on standard size parameters; Based on the standard digital twin model, twin comparison is performed on the target components to obtain the twin comparison features of the target components.
[0087] Optionally, the twin comparison module 503 is specifically configured to: Generate standard 3D point clouds and standard texture images corresponding to target parts based on standard digital twin models; Perform three-dimensional measurement on the target parts to obtain the measured 3D point cloud, and obtain the geometric deviation based on the measured 3D point cloud and the standard 3D point cloud. Photograph the target component to obtain a measured texture image, and calculate the texture similarity based on the measured texture image and the standard texture image; The twin comparison features of the target parts are formed based on texture similarity and geometric deviation.
[0088] Optionally, the anomaly detection module 504 is specifically configured to: In the preset attention classifier, the twin contrast features, spatiotemporal features, and process parameter features are spliced to obtain fusion features; Process the fused features to obtain the abnormal probability and abnormal type weight of the target component; Using the preset dynamic threshold algorithm, the target abnormality type of the target component is determined based on the abnormality probability and abnormality type weight.
[0089] Optionally, the attention classifier includes a multi-head attention network, a feedforward neural network, and an activation network; The anomaly detection module 504 is specifically configured to: The fused features are processed using a multi-head attention network to obtain the initial attention features and abnormal type attention scores; The initial attention features are nonlinearly transformed using a feedforward neural network to obtain the target attention features; The activation network is used to calculate and process the target attention features to obtain the abnormal probability of the target component; The abnormal type attention score is normalized to obtain the abnormal type weight of the target component.
[0090] Optionally, the anomaly detection module 504 is specifically configured to: Obtain multiple historical abnormality determination threshold values and multiple standard deviations of the target component; wherein the historical abnormality determination threshold value mean is the historical abnormality threshold value mean corresponding to the abnormality type in the abnormal range included in the target component, and the standard deviation is the historical standard deviation corresponding to the abnormality type in the abnormal range included in the target component; Based on the mean of the historical anomaly determination threshold and the corresponding standard deviation, the corresponding dynamic anomaly determination threshold is calculated; Determine the corresponding abnormal situation based on the dynamic abnormality judgment threshold and abnormality probability; When the abnormal situation is that an abnormality exists and the corresponding abnormality type weight is greater than or equal to the preset weight, the abnormality type that matches the abnormality type weight is determined as the target abnormality type of the target component.
[0091] It should be noted that the component anomaly detection system provided in the above-mentioned embodiment and the component anomaly detection method provided in the above-mentioned embodiment are based on the same concept. The specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual applications, the component anomaly detection system provided in the above-mentioned embodiment can, as needed, allocate the above-mentioned functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0092] A computing device according to an embodiment of the present application includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, some or all of the steps of the above-mentioned component abnormality detection method are implemented.
[0093] Among them, the computing device can be a computer, and correspondingly, its program is computer software. The above-mentioned parameters and steps in a computing device of the present application can refer to the parameters and steps in the embodiment of a component abnormality detection method above, and will not be repeated here.
[0094] In an embodiment of the present application, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are executed, the steps of the above-mentioned component abnormality detection method are executed.
[0095] The computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0096] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned computer-readable storage medium can be a non-transitory computer-readable storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient computer-readable storage medium.
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0098] Those skilled in the art will appreciate that the present application may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "module" or "system." Furthermore, in some embodiments, the present application may also be implemented in the form of a computer program product in one or more computer-readable media, the computer-readable medium containing a computer-readable program code. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof.
[0099] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0100] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A component abnormality detection method, characterized in that: include: Obtaining a processing video and process parameter characteristics for a target component, wherein the process parameter characteristics are characteristics formed by combining process parameters of a production equipment used to process the target component; Performing feature extraction on the processing video to obtain the spatiotemporal features of the target component; Performing twin comparison on the target component to obtain a twin comparison feature of the target component; Using a preset attention classifier, anomaly detection is performed based on the twin contrast features, the spatiotemporal features, and the process parameter features to obtain a target anomaly type of the target component.
2. The method according to claim 1, characterized in that The extracting features of the processing video to obtain the spatiotemporal features of the target component includes: Extracting original spatiotemporal features of the processed video using a preset 3D convolutional neural network; Extracting key features from the original spatiotemporal features, wherein the key features include key processing area features and key processing period features of the target component; Using a preset spatiotemporal attention network to enhance the key features to obtain enhanced features; The spatiotemporal features of the target component are formed based on the enhanced features and other features of the original spatiotemporal features except the key features.
3. The method according to claim 1, characterized in that The performing twin comparison on the target component to obtain the twin comparison features of the target component includes: Obtaining standard size parameters corresponding to the target component; A standard digital twin model is constructed based on the standard size parameters; A twin comparison is performed on the target component based on the standard digital twin model to obtain a twin comparison feature of the target component.
4. The method according to claim 3, characterized in that The performing twin comparison on the target component based on the standard digital twin model to obtain the twin comparison features of the target component includes: Generating a standard 3D point cloud and a standard texture image corresponding to the target component based on the standard digital twin model; Performing three-dimensional measurement on the target component to obtain a measured 3D point cloud, and obtaining a geometric deviation based on the measured 3D point cloud and the standard 3D point cloud; Photographing the target component to obtain a measured texture image, and calculating texture similarity based on the measured texture image and the standard texture image; A twin comparison feature of the target component is formed based on the texture similarity and the geometric deviation.
5. The method according to any one of claims 1 to 4, characterized in that The method of using a preset attention classifier to perform anomaly detection based on the twin contrast feature, the spatiotemporal feature, and the process parameter feature to obtain a target anomaly type of the target component includes: In a preset attention classifier, feature splicing is performed on the twin contrast feature, the spatiotemporal feature, and the process parameter feature to obtain a fusion feature; Processing the fused features to obtain an abnormality probability and an abnormality type weight of the target component; A preset dynamic threshold algorithm is used to determine the target abnormality type of the target component based on the abnormality probability and the abnormality type weight.
6. The method according to claim 5, characterized in that The attention classifier includes a multi-head attention network, a feedforward neural network and an activation network; The processing of the fusion features to obtain the abnormality probability and abnormality type weight of the target component includes: Processing the fused features using the multi-head attention network to obtain initial attention features and abnormal type attention scores; Performing a nonlinear transformation on the initial attention feature using the feedforward neural network to obtain a target attention feature; The target attention feature is calculated and processed using the activation network to obtain the abnormal probability of the target component; the abnormal type attention score is normalized to obtain the abnormal type weight of the target component.
7. The method according to claim 5, characterized in that The method of determining the target abnormality type of the target component based on the abnormality probability and the abnormality type weight by using a preset dynamic threshold algorithm includes: Acquire multiple historical abnormality determination threshold values and multiple standard deviations of the target component; wherein the historical abnormality determination threshold value mean is the historical abnormality threshold value mean corresponding to the abnormality type in the abnormal range included in the target component, and the standard deviation is the historical standard deviation corresponding to the abnormality type in the abnormal range included in the target component; Based on the historical anomaly determination threshold mean and the corresponding standard deviation, a corresponding dynamic anomaly determination threshold is calculated; Determine a corresponding abnormal situation based on the dynamic abnormality determination threshold and the abnormality probability; When the abnormal situation is that an abnormality exists and the corresponding abnormality type weight is greater than or equal to a preset weight, the abnormality type that matches the abnormality type weight is determined as the target abnormality type of the target component.
8. A component anomaly detection system, characterized in that: include: An acquisition module, configured to acquire a processing video and process parameter characteristics for a target component, wherein the process parameter characteristics are characteristics formed by combining process parameters of a production device for processing the target component; A feature extraction module is used to extract features from the processing video to obtain the spatiotemporal features of the target component; A twin comparison module, configured to perform twin comparison on the target component to obtain a twin comparison feature of the target component; An anomaly detection module is used to use a preset attention classifier to perform anomaly detection based on the twin contrast features, the spatiotemporal features, and the process parameter features to obtain the target anomaly type of the target component.
9. A computing device comprising a memory, a processor, and a program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the component abnormality detection method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the steps of a component abnormality detection method according to any one of claims 1 to 7.