Online anomaly detection method, device and equipment for traditional Chinese medicine pelleting machine and storage medium
By using multi-channel time-series data synchronization alignment and feature fusion technology, the problems of accuracy in abnormal detection and identification of unknown abnormalities in traditional Chinese medicine pill-making machines have been solved, enabling accurate identification of abnormality types and severity scoring, thereby improving the real-time performance and stability of the production process.
Patent Information
- Application Number
- CN202511334841.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing methods for detecting anomalies in traditional Chinese medicine pill-making machines are insufficient in terms of detection dimensions, data utilization, and response speed. They cannot fully reflect the complex dynamic changes in the pill-making process and lack cross-modal feature fusion and dynamic weight adjustment, resulting in inaccurate detection results and an inability to identify unknown anomaly types.
Image features are extracted by multi-channel temporal data synchronization alignment, visual transformer model and multi-layer self-attention mechanism, combined with bidirectional long short-term memory network to extract sensor features, and joint feature vectors are generated by feature fusion layer and cross-modal attention fusion. Open set recognition and extreme value theory calibration are introduced to achieve the discrimination of anomaly type and severity scoring.
It improves the accuracy and robustness of anomaly detection, can identify unknown anomalies, ensure the real-time nature and traceability of the production process, achieve precise root cause localization and multi-level threshold control, and enhance the stability of the production line and the consistency of product quality.
Smart Images

Figure CN120850172A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment anomaly detection technology, and in particular to an online anomaly detection method, device, equipment, and storage medium for a traditional Chinese medicine pill-making machine. Background Art
[0002] Currently, traditional Chinese medicine pill-making machines generally rely on manual periodic inspections or simple threshold monitoring methods based on single sensor signals to determine the equipment's operating status during production. These methods have significant shortcomings in terms of detection dimensions, data utilization, and response speed. On the one hand, existing technologies mostly use time-series signals from single or limited channels, such as vibration, current, and temperature, for anomaly detection, lacking the synchronous acquisition and comprehensive utilization of image information and multi-channel sensor data during production, thus failing to fully reflect the complex dynamic changes in the pill-making process. On the other hand, although some solutions introduce feature extraction methods based on deep learning, they are mostly limited to single-modal analysis, lacking cross-modal feature fusion and dynamic weight adjustment mechanisms, which can easily lead to inaccurate detection results when sensor data is missing, noise interference occurs, or single-modal anomalies occur.
[0003] Existing anomaly detection systems are mostly based on closed-category recognition models, which are insufficient for recognizing unknown anomaly types that may occur in actual production. They also lack mechanisms for open-set recognition and anomaly severity measurement, and cannot provide graded response strategies for anomalies of different levels.
[0004] Existing technologies have significant shortcomings in multimodal data fusion, unknown anomaly identification, and anomaly classification response, making it difficult to meet the high requirements of real-time performance, accuracy, and traceability in the production process of traditional Chinese medicine pill making machines.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] In order to overcome the shortcomings of the prior art, the present invention aims to provide an online anomaly detection method, device, equipment and storage medium for traditional Chinese medicine pill making machine, which aims to solve the problems of inaccurate anomaly detection results and inability to identify unknown anomaly types in the prior art.
[0007] The first aspect of this invention provides an online anomaly detection method for a traditional Chinese medicine pill-making machine, comprising the steps of: collecting multi-channel time-series data of the traditional Chinese medicine pill-making machine during the production process, assigning a unified timestamp and completing synchronization alignment, wherein the multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals, and temperature signals; preprocessing the image sequences in the multi-channel time-series data, extracting features from the preprocessed image sequences using a visual transformer model, and calculating and generating visual embedding vectors from the extracted features using a multi-layer self-attention mechanism; and filtering, normalizing, and sequence constructing the vibration signals, sound signals, current signals, and temperature signals in the multi-channel time-series data. The system constructs a system that extracts temporal dynamic feature representations via a bidirectional long short-term memory network, calculates sensor health scores, and generates dynamic confidence coefficients. Visual embedding vectors, temporal dynamic feature representations, and dynamic confidence coefficients are input into a feature fusion layer. First, gated weighted fusion is performed, followed by cross-membrane attention fusion. During the training phase, cross-membrane consistency contrast constraints and energy constraints are introduced to output a joint feature vector. This joint feature vector is then input into a discrimination module containing a classification head and an open-set recognition head. Anomaly scores are calibrated using extreme value theory to obtain anomaly severity scores. Finally, in a parallel process knowledge constraint inference layer, process compliance is determined, and anomaly type discrimination results and anomaly severity scores are output.
[0008] Optionally, in a first implementation of the first aspect of the present invention, the preprocessing of the image sequence in the multi-channel time-series data, the feature extraction of the preprocessed image sequence through a visual transformer model, and the calculation of the extracted features to generate a visual embedding vector through a multi-layer self-attention mechanism include the following steps: denoising the image sequence of the multi-channel time-series data to obtain a denoised image sequence, with each frame corresponding to a collection timestamp; normalizing the denoised image sequence by linearly scaling the pixel values according to a preset minimum and maximum pixel value range to obtain a normalized image sequence; dividing the normalized image sequence into blocks according to a fixed pixel size to obtain multiple numbered image block sets; inputting the image block sets into a visual transformer model to convert each image block into an image block embedding vector through an embedding mapping process; adding a position-corresponding encoding vector to each image block embedding vector to obtain an image block embedding vector set containing position information; and inputting the image block embedding vector set containing position information into a multi-layer self-attention mechanism for calculation to obtain a visual embedding vector, wherein the visual embedding vector represents the image spatial features under the corresponding timestamp.
[0009] Optionally, in a second implementation of the first aspect of the present invention, the set of image patch embedding vectors containing location information is input into a multi-layer self-attention mechanism for calculation to obtain a visual embedding vector. This includes the following steps: inputting the set of image patch embedding vectors containing location information into the first layer of the multi-layer self-attention mechanism, and obtaining a weighted feature representation through correlation calculation between features; sequentially passing the feature representation output from the first layer to the calculation layer of the multi-layer self-attention mechanism, continuously updating the weight distribution between features in each layer, and gradually enhancing the modeling ability of global spatial dependencies; and after completing the calculation of all multi-layer self-attention mechanisms, outputting the final visual embedding vector, wherein the visual embedding vector integrates global context information and represents the image spatial features under the corresponding timestamp.
[0010] Optionally, in a third implementation of the first aspect of the present invention, the vibration signal, sound signal, current signal, and temperature signal in the multi-channel time-series data are filtered, normalized, and sequenced. Temporal dynamic feature representations are extracted via a bidirectional long short-term memory network, and a sensor health score is calculated and a dynamic confidence coefficient is generated. This includes the following steps: filtering, normalizing, and sequenced the vibration signal, sound signal, current signal, and temperature signal in the multi-channel time-series data to form a channel sequence arranged by timestamps; inputting the channel sequence into a bidirectional long short-term memory network to obtain temporal dynamic feature representations corresponding to the timestamps; statistically analyzing the amplitude interval dwell ratio, spectral centroid change rate, and envelope mutation count within a sliding time window to generate a stability sub-score; and using adjacent similar channels and redundant... Using the remaining channels as a reference, the consistency of sorting, peak synchronization, and event overlap rate of the time-series dynamic feature representation are compared to generate a consensus sub-score. The time-frequency fingerprint database of normal and fault conditions established during the debugging phase is invoked to perform fingerprint matching and grading on the current time window, generating a fingerprint matching sub-score. Data integrity verification is performed to detect missing samples, communication delays, saturation shearing, and range out-of-bounds errors, outputting an integrity flag. The stability sub-score, consensus sub-score, fingerprint matching sub-score, and integrity flag are sorted and fused by voting to obtain the sensor health score. The sensor health score is then graded, mapped, and time-smoothed to generate a dynamic confidence coefficient, which is then linked to the time-series dynamic feature representation in a one-to-one correspondence. The time series of the sensor health score and the dynamic confidence coefficient are recorded.
[0011] Optionally, in the fourth implementation of the first aspect of the present invention, the visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficient are input into the feature fusion layer. First, gated weighted fusion is performed, then cross-membrane attention fusion is executed. During the training phase, cross-membrane consistency contrast constraints and energy constraints are introduced to output a joint feature vector. The steps include: inputting the visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficient into the feature fusion layer; using the dynamic confidence coefficient to perform gated weighted processing on the temporal dynamic feature representation, adjusting the temporal dynamic feature representation of each channel according to the corresponding dynamic confidence coefficient, generating a weighted temporal dynamic feature representation; performing a time alignment operation on the visual embedding vector and the weighted temporal dynamic feature representation within a sliding time window, the alignment strategy including time delay compensation and missing segment interpolation, generating an aligned temporal dynamic feature representation. The system employs a dynamic feature representation approach. Based on the visual embedding vector and the aligned temporal dynamic feature representation, a cross-modal similarity matrix is calculated, and a cross-modal consistency mask is generated according to a similarity threshold. Cross-modal attention fusion is performed on the visual embedding vector and the aligned temporal dynamic feature representation in both the global and local time window paths, resulting in fused output vectors for the two paths. The fused output vectors from the two paths are concatenated and merged, and the feature contribution is reduced in the low-value region of the cross-modal consistency mask. Simultaneously, the fusion result is re-weighted based on the dynamic confidence coefficient to obtain candidate fusion vectors. During the training phase, cross-modal consistency contrast constraints and energy constraints are applied to the candidate fusion vectors. The constrained candidate fusion vectors are output as joint feature vectors, and the time series of the joint feature vectors and the cross-modal consistency mask are stored together.
[0012] Optionally, in the fifth implementation of the first aspect of the present invention, the joint feature vector is input into a discrimination module containing a classification head and an open set identification head. Anomaly scores are calibrated using extreme value theory to obtain an anomaly severity score. In a parallel process knowledge constraint inference layer, process compliance is determined, and anomaly type discrimination results and anomaly severity scores are output. The steps include: inputting the joint feature vector into the discrimination module, which consists of a classification head, an open set identification head, an extreme value theory calibration unit, a process knowledge constraint inference layer, and a conflict coordination unit; performing category discrimination on the joint feature vector using the classification head, outputting candidate classification results for known categories and corresponding confidence levels; performing open set discrimination on the joint feature vector using the open set identification head, generating anomaly scores reflecting the degree to which samples deviate from normal categories, and forming an anomaly score time series arranged in timestamp order; and selecting the anomaly score time series within a sliding time window using the extreme value theory calibration unit. For high-value samples at the tail end, a tail distribution model is constructed. Based on the tail distribution model, a quantile mapping relationship is generated, and the anomaly score corresponding to each time stamp is converted into an anomaly severity score within a fixed interval. Alarm-level and maintenance-level thresholds are adaptively generated according to the tail distribution model. A process knowledge constraint inference layer matches process data with preset process rules to generate process compliance judgments and compliance markers for the corresponding time stamps. The process data refers to parameters reflecting the process operation status, such as temperature range, current ripple, and pellet geometry and shape indicators collected during the traditional Chinese medicine pellet-making process. A conflict coordination unit performs consistency comparisons on the classification candidate results, anomaly scores, anomaly severity scores, and process compliance judgments. When a conflict between multiple judgment results is detected, a decision is made according to preset priority and time continuity rules, outputting the anomaly type discrimination result for the corresponding time stamp. The anomaly type discrimination result and the anomaly severity score are then output.
[0013] Optionally, in the sixth implementation of the first aspect of the present invention, after obtaining the anomaly type discrimination result and the anomaly severity score, the method further includes the following steps: constructing an equipment-process-quality heterogeneity diagram based on the anomaly type discrimination result and the anomaly severity score, obtaining the root cause localization result through a graph neural network, calculating the risk function score, triggering linkage control commands according to multi-level thresholds and visualizing the results; storing and alarming the root cause localization result, linkage control commands, multi-channel time series data, and joint feature vectors, comparing the current and historical feature distribution changes to determine whether a release drift has occurred, and collecting and organizing cross-membrane state scarce fault sample data in the offline stage.
[0014] The second aspect of this invention provides an online anomaly detection device for a traditional Chinese medicine pill-making machine, comprising: an alignment module for collecting multi-channel time-series data of the traditional Chinese medicine pill-making machine during the production process, assigning a unified timestamp, and completing synchronous alignment, wherein the multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals, and temperature signals; a visual embedding vector generation module for preprocessing the image sequences in the multi-channel time-series data, extracting features from the preprocessed image sequences using a visual transformer model, and calculating and generating visual embedding vectors from the extracted features using a multi-layer self-attention mechanism; and a dynamic feature extraction module for filtering the vibration signals, sound signals, current signals, and temperature signals in the multi-channel time-series data. Normalization and sequence construction are performed, and temporal dynamic feature representations are extracted via a bidirectional long short-term memory network to calculate sensor health scores and generate dynamic confidence coefficients. A fusion module inputs the visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficients into the feature fusion layer. First, gated weighted fusion is performed, followed by cross-membrane attention fusion. Cross-membrane consistency contrast constraints and energy constraints are introduced during the training phase to output a joint feature vector. An anomaly output module inputs the joint feature vector into a discrimination module containing a classification head and an open-set recognition head. Anomaly scores are calibrated using extreme value theory to obtain anomaly severity scores. Finally, in a parallel process knowledge constraint inference layer, process compliance is determined, and anomaly type discrimination results and anomaly severity scores are output.
[0015] A third aspect of the present invention provides an electronic device comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the computer-readable instructions in the memory to cause the electronic device to perform the various steps of the online anomaly detection method for a traditional Chinese medicine pill-making machine as described above.
[0016] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-readable instructions that, when executed on a computer, cause the computer to perform the steps of the online anomaly detection method for a traditional Chinese medicine pill-making machine as described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: This invention introduces multi-channel sensors and industrial cameras into a traditional Chinese medicine pill-making production line, achieving unified timestamp identification and synchronous alignment of image sequences and multi-source time-series data such as vibration, sound, current, and temperature, ensuring data integrity and timeliness. Combining visual transformers for image spatial feature extraction and bidirectional long short-term memory networks for modeling temporal dynamic features, and incorporating innovative mechanisms such as dynamic confidence coefficient weighting, cross-modal consistency masks, and global and local dual-path attention fusion in the feature fusion stage, the system can effectively capture abnormal features in multimodal data, improving the robustness and accuracy of anomaly detection. Through open set recognition and extreme value theory calibration, it can not only identify known anomalies but also quantify the severity of unknown anomalies. Furthermore, in the process knowledge constraint inference layer, it combines real-time process data for compliance judgment, ensuring the comprehensiveness and reliability of anomaly assessment. Based on a heterogeneous graph structure of equipment-process-quality and a dual-path message passing algorithm, this invention enables precise root cause localization and triggers multi-level threshold control strategies based on risk assessment results, reducing unnecessary downtime while ensuring production safety. Compared with existing technologies, this invention improves the real-time monitoring capability, unknown anomaly identification capability, anomaly handling precision level and fault tracing efficiency of the pelleting machine production process, effectively improving the stability of the production line and the consistency of product quality. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart of an online anomaly detection method for a traditional Chinese medicine pill-making machine provided in an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of the online anomaly detection device for a traditional Chinese medicine pill-making machine provided by the present invention.
[0022] Figure 3 This is a schematic diagram of the electronic device structure provided by the present invention. Detailed Implementation
[0023] This invention provides an online anomaly detection method, apparatus, device, and storage medium for a traditional Chinese medicine pill-making machine. The terms "first," "second," "third," "fourth," etc. (if applicable) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0024] Please see Figure 1 , Figure 1 A flowchart of an online anomaly detection method for a traditional Chinese medicine pill-making machine provided by the present invention is shown in the figure, which includes the following steps: S10. Collect multi-channel time-series data of the Chinese medicine pill-making machine during the production process, assign a unified timestamp mark and complete synchronization alignment. The multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals and temperature signals. Specifically, in this embodiment, the following equipment can be deployed in the traditional Chinese medicine pelleting machine production line: at least one high-definition camera (e.g., 1920×1080 resolution, 25fps frame rate) for acquiring image sequences, for example, one high-definition camera is installed at the pellet forming outlet to capture the roundness and surface smoothness of the pellets; one high-definition camera is installed at the feeding screw to capture the uniformity of material conveying; at least one vibration sensor for acquiring vibration signals, for example, one vibration sensor is installed in the pelleting machine motor bearing housing to monitor mechanical vibration and determine bearing wear; one vibration sensor is installed on the roller shaft to monitor cutting stability; at least one sound sensor for acquiring sound signals, for example, one sound sensor is installed in the pellet forming area to collect collision sounds during forming and identify abnormal noises of material jamming; at least one current sensor for acquiring current signals, for example, one current sensor is connected in series in the heating mold power supply circuit to monitor current fluctuations and reflect heating uniformity; at least one temperature sensor for acquiring temperature signals, for example, one temperature sensor is attached to the mold surface to monitor the pellet forming temperature, etc.
[0025] After acquiring multi-channel time-series data, the multi-channel time-series data acquired by different acquisition devices are marked according to a unified time base. Time axis correction and interpolation processing are then used to ensure that the multi-channel time-series data have a corresponding relationship at the same time point. Specifically, firstly, the clock of the production line PLC controller (with an accuracy of 1ms) is used as a unified base. Both the camera and sensors are synchronized to this clock via the Ethernet NTP protocol. Each frame of the camera image is marked with the acquisition time (e.g., "2025-08-21 08:00:00.000", "2025-08-21 08:00:00.040"). The vibration sensor generates one data point every 10ms, marked with the time (e.g., "2025-08-21 08:00:00.000", "2025-08-21 08:00:00.010"). Interpolation processing is then performed on data with different sampling rates. For example, align the image sequence of 0.04 seconds / frame from the camera with the signal of 0.01 seconds / vibration from the vibration sensor: within the interval from "2025-08-21 08:00:00.000" to "2025-08-21 08:00:00.040", take 4 data points (0.000s, 0.010s, 0.020s, 0.030s) from the vibration sensor and associate them with the image of that frame. Supplement the virtual data at intermediate moments through linear interpolation to ensure that the same time point (such as 0.020s) has both image features and vibration features.
[0026] S20. Preprocess the image sequence in the multi-channel time series data, extract features from the preprocessed image sequence through the visual transformer model, and calculate and generate a visual embedding vector from the extracted features through a multi-layer self-attention mechanism. Specifically, the image sequence of multi-channel time-series data is denoised to obtain a denoised image sequence, with each frame corresponding to a timestamp of acquisition. In the production line of traditional Chinese medicine pill-making machines, images captured by industrial cameras at the pill-forming exit often exhibit salt-and-pepper noise (randomly appearing black and white noise) and Gaussian noise (overall blur) due to workshop dust and equipment vibration. For salt-and-pepper noise, a 3×3 median filter can be used. For example, if a pure white noise point (pixel value 255) exists on the surface of a pill in a frame, and the grayscale values of the surrounding 8 pixels are 150-160, the median filter takes the median value of these 9 pixels (155) and replaces the original noise value, eliminating the isolated noise point. For Gaussian noise, a 2D Gaussian filter (standard deviation σ=1.0, window 5×5) can be used. For example, if the pill edges are blurred due to slight lens shake, a Gaussian filter can smooth high-frequency noise, making the edge contours clearer. Each frame of the denoised image is marked with its acquisition time (e.g., "2025-08-21 09:15:00.040") to ensure time synchronization with sensor data. This embodiment eliminates noise interference through denoising, preventing noise points from being misidentified as pellet defects during subsequent feature extraction (e.g., mistaking salt-and-pepper noise for surface cracks); while preserving key features such as pellet edges and surface textures, laying the foundation for subsequent spatial feature extraction; timestamp association ensures temporal consistency between image data and other sensor data, supporting cross-modal fusion.
[0027] Next, the denoised image sequence is normalized by linearly scaling the pixel values according to a preset minimum and maximum pixel value range, resulting in a normalized image sequence. The pixel value range of the denoised image is 0-1024 (10-bit camera output), while the visual transformer model typically requires the input pixel value to be in the range of 0-255 (8-bit standard). The preset minimum pixel value = 0, maximum pixel value = 1024, and the linear scaling formula is: Normalized pixel value = (original pixel value - minimum pixel value) / (maximum pixel value - minimum pixel value) × 255. For example: the pixel value of the normal area on the surface of the pellet is 512 → after normalization, it is (512-0) / 1024×255=127; the pixel value of the surface defect area (darker) is 100 → after normalization, it is (100 / 1024)×255≈25; the pixel value of the background area is 1024 → after normalization, it is 255. This embodiment uses normalization to unify the pixel value scale, avoiding feature deviations caused by differences in camera hardware (such as different dynamic ranges of different batches of cameras); at the same time, it reduces the impact of excessively large or small pixel values on model training (such as avoiding large pixel values from dominating feature calculation).
[0028] The normalized image sequence is then divided into blocks according to a fixed pixel size, resulting in multiple numbered image block sets. For example, a normalized 512×512 pixel image is divided into 16×16 pixel blocks, resulting in (512 / 16)×(512 / 16)=32×32=1024 image blocks, numbered 0-1023 (row-major). Each block records its position in the original image; for example, image block 0 corresponds to the top left corner (0,0)-(15,15), and image block 1 corresponds to (16,0)-(31,15). This embodiment decomposes the high-resolution image into small blocks through block processing, reducing the computational complexity of the visual transformer (avoiding direct processing of 512×512 high-dimensional data). Each block participates in feature extraction as an independent unit, facilitating the capture of local details (such as tiny cracks on the surface of a pellet that may only occupy 1-2 blocks). The correspondence between numbering and position provides a basis for subsequent addition of position encoding, ensuring that the model understands the spatial layout of the blocks.
[0029] Next, the image patch set is input into the visual transformer model, and each image patch is converted into an image patch embedding vector through an embedding mapping process. Specifically, the preprocessed image patch set is input into the input end of the visual transformer model in sequence, and the number of channels and size of the input data are matched with the structure of the visual transformer model. In the input embedding layer of the visual transformer model, a fixed-dimensional vectorization encoding operation is performed on each image patch to map the two-dimensional pixel information into a one-dimensional feature vector.
[0030] As an example, in this step, each 16×16 image patch needs to be converted into a fixed-dimensional vector (e.g., 768-dimensional) to adapt to the self-attention mechanism of the visual transformer model. The embedding mapping process is as follows: the 16×16 image patch is flattened into a 256-dimensional vector (16×16=256); the 256-dimensional vector is then mapped to a 768-dimensional embedding vector using a linear projection matrix (size 256×768). For example, after mapping, a patch containing normal surface texture (e.g., patch 512) has a higher dimension related to smoothness (e.g., 0.8); a patch containing cracks (e.g., patch 513) has a higher dimension related to edge discontinuity (e.g., 0.9). Converting the pixel information of the image patches into high-dimensional feature vectors facilitates the model's learning of abstract semantics (e.g., high-level features such as cracks and deformations); the fixed-dimensional embedding vectors ensure that the features of different image patches are comparable and fused, providing a unified input for subsequent self-attention calculations.
[0031] Next, a location-corresponding encoding vector is added to each image patch embedding vector, resulting in a set of image patch embedding vectors containing location information. Specifically, sine-cosine position encoding is used to add location information to the embedding vector of each image patch (ensuring the model knows the spatial location of the patch in the original image). The position encoding formula (for the i-th dimension of the embedding vector) is as follows: If i is even: PE (pos, i) = sin (pos / 10000^(2i / d_model)) If i is odd: PE (pos, i) = cos (pos / 10000^(2i / d_model)) where pos is the image patch number (0-1023) and d_model = 768 (embedding vector dimension).
[0032] For example, the positional encoding of image patch 0 (top left corner) has the following dimensions: 0-th dimension: sin(0 / 10000^(0 / 768)) = 0, 1-th dimension: cos(0 / 10000^(0 / 768)) = 1; the positional encoding of image patch 512 (center) has the following dimensions: 0-th dimension: sin(512 / 10000^(0 / 768)) ≈ sin(512) ≈ 0.98, 1-th dimension: cos(512) ≈ -0.19. Finally, the embedding vector for each image patch = image patch embedding vector + positional encoding vector. This embodiment compensates for the inherent neglect of spatial location by adding a position-corresponding encoding vector to each image patch embedding vector (the self-attention mechanism only focuses on feature relevance and does not directly encode location), enabling the model to distinguish image patches with the same content but different locations (such as the pellet in the top left corner and the pellet in the bottom right corner); ensuring that the global spatial structure can be captured (such as the overall features of the pellet distribution density and arrangement direction in the image).
[0033] Finally, the set of image patch embedding vectors containing location information is input into a multi-layer self-attention mechanism for calculation to obtain a visual embedding vector, which represents the image spatial features at the corresponding timestamp. Specifically, this step involves: inputting the set of image patch embedding vectors containing location information into the first layer of the multi-layer self-attention mechanism, and obtaining a weighted feature representation through correlation calculation between features; sequentially passing the feature representation output from the first layer to the calculation layers of the multi-layer self-attention mechanism, continuously updating the weight distribution between features in each layer to gradually enhance the modeling ability of global spatial dependencies; after completing the calculation of all multi-layer self-attention mechanisms, the final visual embedding vector is output, which integrates global context information and represents the image spatial features at the corresponding timestamp.
[0034] As an example, a 12-layer self-attention mechanism (12 attention heads per layer, total dimension 768) is used to iteratively calculate 1024 embedding vectors containing location information. Self-attention calculation: Each layer calculates the association weights between blocks using a "query (Q)-key (K)-value (V)" mechanism. For example, when a surface crack is detected on a pellet, the attention weight of the block containing the crack (e.g., block 513) significantly increases with surrounding blocks (blocks 512, 514, 575, etc.) (e.g., 0.8), while the weight with the background block (e.g., block 0) is lower (e.g., 0.1). Higher-level self-attention (e.g., layer 10) strengthens global associations, such as associating cracked blocks with irregularly shaped pellets, capturing the composite features of cracks accompanied by deformation. Output visual embedding vector: After 12 layers of calculation, a 768-dimensional vector is output, containing global spatial features (e.g., 30% of the pellets have surface cracks, concentrated in the central region of the image). In this embodiment, the multi-layer self-attention mechanism can dynamically focus on key areas (such as defect areas), suppress irrelevant backgrounds, and improve the discriminativeness of features. It can also fuse local details (such as cracks in a single pellet) with global context (such as the distribution density of defects) to comprehensively characterize the spatial features of the image. Its output visual embedding vector can be directly used for cross-modal fusion and combined with the temporal features of the sensor to improve the accuracy of anomaly detection, such as associating image cracks with vibration signal anomalies to diagnose equipment failure.
[0035] This embodiment introduces denoising, normalization, and segmentation operations during the image sequence preprocessing stage to ensure the stability and consistency of input data, significantly reducing the interference of environmental noise, illumination changes, and equipment jitter on feature extraction. In the visual transformer model, by combining positional encoding and multi-layer self-attention mechanism, high-precision modeling of global and local spatial relationships in the image is achieved, improving the weighted expression ability of key region features. Through serialized feature updates and global dependency enhancement, the model has stronger spatial feature aggregation and contextual semantic understanding capabilities in multi-channel temporal scenarios, significantly improving the accuracy and robustness of anomaly detection, especially exhibiting higher generalization performance and real-time processing capabilities under the complex working conditions of traditional Chinese medicine pill-making machines.
[0036] S30. Filter, normalize and construct sequences for vibration signals, sound signals, current signals and temperature signals in multi-channel time series data. Extract time series dynamic feature representation through bidirectional long short-term memory network, calculate sensor health score and generate dynamic confidence coefficient. Specifically, this step includes: S31, filtering, normalizing, and constructing sequences for vibration, sound, current, and temperature signals in multi-channel time-series data to form a channel sequence arranged by timestamps. The channel sequence is then input into a bidirectional long short-term memory network to obtain the time-series dynamic feature representation corresponding to each timestamp. This embodiment employs different filtering methods for different signal noise characteristics. For example, for vibration signals containing 50Hz power frequency interference and high-frequency mechanical noise, a fourth-order Butterworth low-pass filter (cutoff frequency 100Hz) is used to filter high-frequency noise above 100Hz, retaining the effective frequency band (10-80Hz) of the motor bearing vibration. For instance, the amplitude of the 50Hz interference in the original vibration signal is 0.2mm / s, which is reduced to 0.05mm / s after filtering. For current signals containing 50Hz harmonics from power grid fluctuations, a notch filter (50Hz) is used to eliminate power frequency harmonic interference. For example, the original value of the heating mold current contains 50Hz fluctuations (peak-to-peak value 1.2A), which is reduced to 0.3A after filtering.
[0037] Next, the signal scale is unified through normalization. First, the normal ranges for each signal are preset: vibration: 0.1-0.5 mm / s; sound: 60-80 dB; current: 5-8 A; temperature: 55-60℃. A linear normalization formula is used: Normalized value = (Original value - Minimum value) / (Maximum value - Minimum value) (mapped to the 0-1 range). For example: vibration signal 0.3 mm / s → (0.3-0.1) / (0.5-0.1) = 0.5; temperature 58℃ → (58-55) / (60-55) = 0.6.
[0038] Then, the normalized data of the four channels are arranged according to the timestamp (one sampling point every 100ms) to form a channel sequence (format: time step × number of channels, such as 100×4, that is, a 10-second window containing 100 time steps, with 4 signal values for each time step).
[0039] The channel sequence is input into a bidirectional Long Short-Term Memory (LSTM) network (2 layers, 128 hidden units per layer), which outputs a 128-dimensional temporal dynamic feature representation. For example, when a motor bearing wears, the vibration signal continuously increases, and the trend change dimension value in the LSTM feature increases from 0.2 to 0.8, capturing long-term dependencies. The bidirectional LSTM can simultaneously capture past and future temporal dependencies (such as the causal relationship of temperature increase → current increase), generating feature representations rich in dynamic trends.
[0040] S32. Within the sliding time window, statistically analyze the amplitude interval dwell ratio, the rate of change of the spectral centroid, and the envelope mutation count to generate a stability sub-score; In this step, taking vibration signals as an example, the sliding time window is set to 10 seconds (100 sampling points), and the following indicators are calculated: Amplitude range dwell ratio: The preset normal amplitude range is [0.1, 0.3] mm / s, and the abnormal range is [0.3, 0.5] mm / s; the percentage of sampling points in the normal range within 10 seconds is counted: 80 points in the normal range → dwell = 80 / 100 = 80%. Spectrum centroid change rate: Fourier transform is performed on every 2-second sub-window to calculate the spectrum centroid (centroid on the frequency axis): Under normal operating conditions, the centroid is stable at 50Hz. In the current window, the centroid rises from 0Hz to 60Hz → change rate = (60-50) / 50 = 20%. Envelope mutation count: Extract the signal envelope (using Hilbert transform), set a mutation threshold of 0.1 mm / s (exceeding this value is considered a mutation), and within 10 seconds, the envelope suddenly increases from 0.2 mm / s to 0.4 mm / s (1 time), and suddenly decreases from 0.3 mm / s to 0.1 mm / s (1 time) → total count = 2 times. Stability sub-score calculation: Weight allocation: dwell ratio 40%, spectral change rate 30%, mutation count 30%. Scoring formula: (80%×100)×0.4 + (1-20%)×100×0.3 + (1-2 / 5)×100×0.3=32+24+24=80 points (full score 100, mutation count threshold is 5 times, exceeding which results in 0). This step evaluates signal stability from three dimensions: amplitude stability (static), spectral change (dynamic), and transient mutation (burst), comprehensively reflecting whether the sensor is in a stable working state.
[0041] S33. Using adjacent channels of the same type and redundant channels as references, compare the ordering consistency, peak synchronization and event overlap rate of the time-series dynamic feature representation to generate a consensus sub-score. In this step, the vibration signal (main channel) is compared with the vibration signal of the adjacent motor (redundant channel) and the sound signal of the same area (related channel). Sort consistency: The temporal dynamic characteristics of the two channels within 10 seconds are sorted by the degree of abnormality (from high to low). The main channel is sorted [time 5, time 3, time 8], and the redundant channel is sorted [time 5, time 8, time 3] → consistency = 2 / 3 ≈ 67% (the first two are the same). Peak synchronization: The vibration peak of the main channel occurs at time 5 (0.4 mm / s), and the peak of the redundant channel occurs at time 5.1s (0.38 mm / s) → time difference 0.1s (≤0.2s threshold), synchronization = 100%. Event overlap rate: The main channel detects 3 events with amplitude exceeding 0.3 mm / s (times 2, 5, 7), and the sound signal detects 2 high-frequency abnormal noise events (times 5, 7) → 2 overlaps → overlap rate = 2 / 3 ≈ 67%. Consensus score calculation: Weight allocation: Ranking consistency 30%, peak synchronicity 40%, event overlap rate 30%. Score: 67%×100×0.3 + 100×0.4 + 67%×100×0.3=20.1+40+20.1=80.2. This embodiment reduces misjudgments caused by single-channel failures (such as loose sensors) through cross-validation of adjacent / redundant channels; the event overlap of associated channels (such as vibration and sound) can verify the authenticity of anomalies (such as mechanical failures often accompanied by vibration and sound anomalies), improving the reliability of the score.
[0042] S34. Call the time-frequency fingerprint database of normal and fault conditions established during the debugging phase, perform fingerprint matching and classification on the current time window, and generate fingerprint matching sub-scores. The time-frequency fingerprint database is a set of time-frequency feature templates of typical normal and various fault conditions collected and processed during the debugging phase. By collecting multi-channel signal data under various known operating conditions of the equipment, wavelet transform is used to extract the joint feature distribution of the signal in the time and frequency domains, and key feature parameters are quantized, encoded and archived to finally form a standardized time-frequency feature database that can be used for real-time matching and classification.
[0043] In this step, the construction of the time-frequency fingerprint database during the debugging phase includes: Normal operating conditions: collecting vibration signals during normal motor operation, extracting time-frequency features through wavelet transform (e.g., 60% energy in the 50Hz band and 20% in 100Hz) to form a normal fingerprint; Fault operating conditions: collecting vibration signals during bearing wear, with time-frequency features of 50Hz energy decreasing to 40% and 150Hz energy increasing to 30%, forming a wear fingerprint. Current window matching: performing wavelet transform on the current 10-second vibration signal to extract time-frequency features: 55% energy in 50Hz, 25% in 100Hz, and 10% in 150Hz. Matching degree with normal fingerprints: (55% / 60% + 25% / 20% + 10% / 0%) / 3≈0.92+1.25+0) / 3≈0.72 (72%); Matching degree with worn fingerprints: (55% / 40% +25% / 0%+10% / 30%) / 3≈1.38+0+0.33) / 3≈0.57 (57%). Therefore, the fingerprint matching sub-score is: the highest matching degree of 72 points (matching normal operating conditions). This embodiment uses known operating condition characteristics (normal / faulty) during the debugging phase as a benchmark to quickly determine whether the current signal conforms to the expected pattern; time-frequency domain features are better able to distinguish subtle faults (such as high-frequency harmonics from bearing wear) than time-domain features, improving the accuracy of anomaly identification.
[0044] S35. Perform data integrity verification, detect missing samples, communication delay, saturation shearing and range out of bounds, and output integrity flags; In this step, as an example, the current signal (sampling rate 50Hz) is verified, including missing sampling detection: 500 points should be collected within 10 seconds, but only 490 are actually collected → 10 points are missing → marked as slightly missing; communication delay: the difference between the data transmission timestamp and the collection timestamp is 0.5s (threshold ≤ 0.3s) → marked as delay exceeding the standard; saturation shearing: the maximum value of the current signal is clamped at 10A (sensor range 5-10A), but the actual calculated value should reach 10.5A → marked as saturation; range exceeding the limit: the temperature signal is -5℃ (range 0-100℃) → marked as exceeding the limit; integrity marking: considering the above issues, a mark value of 60 points is output (out of 100, 10-20 points are deducted for each issue). This embodiment avoids using invalid data for feature extraction by identifying anomalies in data acquisition / transmission (such as sensor failure, communication interruption); it provides a basis for data quality dimensions for subsequent health scoring, reducing the interference of poor-quality data on decision-making.
[0045] S36. The stability sub-score, consensus sub-score, fingerprint matching sub-score and integrity marker are sorted, voted and fused to obtain the sensor health score; In this step, as an example, assume a vibration sensor has four sub-scores: stability (80 points), consensus (80 points), fingerprint matching (72 points), and integrity (60 points). The ranking and voting process is as follows: Rank the four scores: 80, 80, 72, 60; remove the highest and lowest scores, and take the average of the remaining two scores: (80+72) / 2 = 76 points; or use weighted voting (stability 30%, consensus 30%, fingerprint matching 20%, integrity 20%): 80×0.3+80×0.3+72×0.2+60×0.2=24+24+14.4+12=74.4 points → take 74 points. This embodiment avoids the bias of a single indicator by integrating multi-dimensional indicators (stability, consensus, fingerprint matching, and integrity); reduces the interference of extreme values through ranking and voting (e.g., removing extreme values), or highlights key indicators (such as stability and consensus) through weighted fusion, making the health score more reliable.
[0046] S37. Perform hierarchical mapping and time smoothing on the sensor health score to generate dynamic confidence coefficients, and establish a one-to-one correspondence between the channels and the time-series dynamic feature representation, and record the time series of sensor health scores and dynamic confidence coefficients.
[0047] In this step, as an example, the sensor health score is mapped as follows: health score ≥90 → confidence coefficient 1.0; 80-89 → 0.9; 70-79 → 0.8; 60-69 → 0.7; <60 → 0.5 (low confidence); if the vibration sensor health score is 74, then it is mapped to 0.8. Time smoothing refers to using a moving average (3 time windows). The confidence coefficients of the first two windows are 0.85 and 0.8, respectively. After smoothing the current window, (0.85+0.8+0.8) / 3≈0.82. Correlation and recording: The dynamic confidence coefficient of 0.82 is correlated with the temporal dynamic characteristics of the vibration channel. Similarly, coefficients (such as 0.85, 0.90, and 0.78) are generated for the sound, current, and temperature channels, and the time series is recorded. This step converts the health score into coefficients (0-1) that can be directly used for feature weighting through hierarchical mapping, which facilitates the dynamic adjustment of channel weights in subsequent fusion layers (reducing the feature weights of low-confidence channels); reduces instantaneous fluctuations (such as score drops caused by sudden interference) through time smoothing, making the confidence coefficient more stable and improving the robustness of cross-modal fusion; and uses one-to-one channel correspondence to ensure that the weight of each feature matches its sensor state, avoiding misleading decisions by faulty sensors.
[0048] This embodiment achieves accurate assessment and robust tracking of sensor health status by introducing a multi-dimensional feature fusion and dynamic confidence coefficient calculation mechanism, demonstrating significant innovation and application value. Multi-channel time-series data, after filtering, normalization, and sequence construction, is input into a bidirectional long short-term memory network to fully capture temporal dependencies and obtain a high-fidelity dynamic representation of time-series features. The stability sub-score integrates three complementary indicators: amplitude interval dwell ratio, spectral centroid change rate, and envelope mutation count, simultaneously reflecting the stability characteristics of amplitude, frequency, and transient changes. A consensus sub-score is introduced, significantly enhancing the robustness of sensor anomaly detection by comparing the ordering consistency, peak synchronization, and event overlap rate of adjacent and redundant channels. A standardized time-frequency fingerprint database built during the debugging phase enables real-time rapid matching and grading of normal and fault conditions, improving the accuracy of anomaly identification. Through integrity verification and multi-sub-score ranking and voting fusion, combined with grading mapping and time smoothing to generate dynamic confidence coefficients, continuous and reliable tracking of health status is achieved. This invention significantly improves the accuracy of subsequent cross-modal fusion, enabling the model to work robustly even when sensors have noise, faults, or missing data. Ultimately, it improves the accuracy and robustness of anomaly detection in traditional Chinese medicine pill-making machines, providing a reliable basis for fault diagnosis and linkage control of these machines.
[0049] S40. Input the visual embedding vector, temporal dynamic feature representation and dynamic confidence coefficient into the feature fusion layer, first perform gated weighted fusion, then perform cross-membrane attention fusion, and introduce cross-membrane consistency contrast constraint and energy constraint during the training phase to output joint feature vector; Specifically, this step includes S41, inputting the visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficient into the feature fusion layer; wherein, the feature fusion layer is a deep fusion module used to perform correlation modeling and information complementarity of data features from different modalities within a unified representation space. For example, the visual embedding vector: for example, a 768-dimensional vector from a visual transformer, representing the image spatial features of the pellet at a certain timestamp; the temporal dynamic feature representation: for example, a 128-dimensional vector from a bidirectional LSTM, fusing the temporal trends of vibration (0.3mm / s), sound (75dB), current (6.5A), and temperature (58℃) (e.g., vibration amplitude shows an upward trend, current fluctuations increase); the dynamic confidence coefficient: for example, the real-time reliability scores of four channels (vibration 0.82, sound 0.85, current 0.90, temperature 0.78), these three are associated by timestamp (e.g., "2025-08-21 10:00:00.000"), and are jointly input into the feature fusion layer. The purpose of this step is to achieve a preliminary correlation between visual and multi-sensor time-series data, providing raw input for cross-modal fusion. The introduction of dynamic confidence coefficients provides a basis for subsequent weighted processing, ensuring that the fusion process takes data reliability into account.
[0050] S42. The time-series dynamic feature representation is gated and weighted using dynamic confidence coefficients, adjusting the magnitude of each channel's time-series dynamic feature representation according to the corresponding dynamic confidence coefficient, generating a weighted time-series dynamic feature representation. Specifically, for each time-series dynamic feature representation, the dynamic confidence coefficient corresponding to the channel is retrieved according to the channel correspondence between the channel index and the dynamic confidence coefficient. At each time-series, the time-series dynamic feature vector of the corresponding channel is multiplied element-wise with the dynamic confidence coefficient, achieving feature magnitude scaling by the coefficient. Channels with higher dynamic confidence coefficients retain a larger proportion of features, and vice versa. The weighted operation results of each channel at each time-series are recombined into a weighted time-series dynamic feature representation with the same dimension as the original feature sequence, while maintaining the time order.
[0051] In this step, the weighted time-series feature = original time-series dynamic feature × dynamic confidence coefficient (corresponding to each channel). For example, the vibration channel time-series feature (32 dimensions) × 0.82 (confidence coefficient) reduces the weight of low-reliability features; the current channel time-series feature (32 dimensions) × 0.90 (high confidence) retains the dominance of high-reliability features; sound (32 dimensions) × 0.85 and temperature (32 dimensions) × 0.78 are adjusted according to the confidence level. The output is a 128-dimensional weighted time-series dynamic feature representation (each channel weight matched to sensor reliability). This step dynamically suppresses the interference of low-confidence channels (such as unreliable data from temperature sensors due to drift) on the fusion result; and enhances the effective features of high-confidence channels (such as current sensors), thereby improving the signal-to-noise ratio of the fused features.
[0052] S43. Perform time alignment operation on the visual embedding vector and the weighted temporal dynamic feature representation within the sliding time window. The alignment strategy includes time delay compensation and missing segment interpolation to generate the aligned temporal dynamic feature representation. In this step, since the visual embedding vector is generated every 0.04 seconds (25fps camera) and the weighted temporal feature is generated every 0.1 seconds (10Hz sensor sampling rate), there is a time difference between the two. Therefore, time alignment is required. The alignment strategy includes: time delay compensation, for example, if the visual data is detected to be 0.02 seconds behind the temporal data (due to image processing time), the timestamp of the visual embedding vector is shifted backward by 0.02 seconds to ensure consistency with the time reference; missing segment interpolation, for example, if visual data is missing within a 0.1-second time interval (brief camera interruption), linear interpolation of the visual embedding vectors from the preceding and following frames (0.08 seconds and 0.12 seconds) is used to generate virtual visual features at the 0.1-second time. The aligned temporal dynamic feature representation is output, maintaining a 1:1 time correspondence with the visual embedding vector. This step can solve the problem of time misalignment caused by different acquisition frequencies and processing times of multimodal data, ensuring that visual and temporal features at the same moment can be fused; by interpolation, it can make up for data loss, avoid fusion interruption caused by local disconnection, and improve system robustness.
[0053] S44. Calculate the cross-modal similarity matrix based on the visual embedding vector and the aligned temporal dynamic feature representation, and generate a cross-modal consistency mask according to the similarity threshold; the calculation of the cross-modal similarity matrix based on the visual embedding vector and the aligned temporal dynamic feature representation specifically refers to normalizing the visual embedding vector and the temporal dynamic feature representation in the feature dimension, calculating the similarity score between each visual feature and each temporal feature element by element using cosine similarity, and finally arranging all pairwise similarity scores in matrix form to form a similarity matrix reflecting the cross-modal feature association strength.
[0054] In this step, as an example, the cosine similarity formula is used to calculate the similarity between the 768-dimensional visual embedding vector and the temporal features after alignment with the 128-dimensional vector (mapped to 768-dimensional). The results show that the visual features of the pellet surface crack have a similarity of 0.85 (high correlation) with the vibration features (caused by bearing wear), and a similarity of 0.3 (low correlation) with the temperature features. A similarity threshold of 0.6 is set, with values ≥0.6 marked as 1 (high consistency) and values <0.6 marked as 0 (low consistency). The mask is [1,0] (vibration channel 1, temperature channel 0). This step provides a correlation filter for subsequent fusion by quantifying the correlation strength between the visual features and each temporal channel; the consistency mask can reduce the contribution of low-correlation features (such as temperature being unrelated to cracks) and avoid interference from invalid information.
[0055] S45. Perform cross-modal attention fusion on the visual embedding vector and the aligned temporal dynamic feature representation in the global time window path and the local time window path respectively to obtain the fused output vector of the two paths. In this step, the global time window path focuses on long-term trends. For example, within 60 seconds, visual features show a gradual increase in particle cracks, while time-series features show a continuous rise in vibration amplitude. An attention mechanism strengthens the long-term correlation between these two factors, generating a global fusion vector (896 dimensions). The local time window path focuses on instantaneous anomalies. For example, within 5 seconds, visual features capture a sudden deformation of a particle, while time-series features show a sudden surge in current. An attention mechanism focuses on the strongly correlated features at that moment, generating a local fusion vector (896 dimensions). After fusion, the output vectors of the two paths are combined (global + local, each 896 dimensions). This step captures long-term cross-modal trends (such as the gradual process of equipment degradation) through the global path and instantaneous anomalies (such as sudden failures) through the local path, taking into account anomaly patterns at different time scales. The attention mechanism automatically focuses on key correlated areas (such as the strong correlation between cracks and vibration), improving the discriminative power of the fused features.
[0056] S46. Concatenate and merge the fusion output vectors of the two paths, reduce the feature contribution in the low value region of the cross-modal consistency mask, and then re-weight the fusion result according to the dynamic confidence coefficient to obtain the candidate fusion vector. In this step, for example, the global (896-dimensional) and local (896-dimensional) fused vectors are concatenated to 1792 dimensions, and then compressed to 896 dimensions through a fully connected layer (preserving key information). Next, features in low-value regions of the cross-modal consistency mask (e.g., temperature channel, mask 0) are assigned a weight of 0.3 (reducing contribution), while high-value regions (vibration channel, mask 1) are assigned a weight of 1.0. Finally, dynamic confidence coefficients (vibration 0.82, current 0.90) are used to reweight the merged features, strengthening the influence of high-confidence channels. The final output is an 896-dimensional candidate fused vector (fusing global / local features, filtered by both masking and confidence). This step integrates long-term and short-term features through concatenated merging, filters low-correlation information through masking adjustments, strengthens high-reliability features through secondary weighting, and optimizes the candidate vectors through multiple rounds of optimization. A dynamic adjustment mechanism adapts to changes in sensor state (e.g., a sudden drop in confidence for a certain channel), ensuring the stability of the fusion result.
[0057] S47. During the training phase, apply cross-modal consistency contrast constraints to the candidate fusion vectors, apply energy constraints to the candidate fusion vectors, output the constrained candidate fusion vectors as joint feature vectors, and store the time series of the joint feature vectors and the cross-modal consistency mask together.
[0058] In this step, as an example, the cross-modal consistency contrast constraint requires that, for the same anomalous sample (such as a cutting tool malfunction), the Euclidean distance between the visual embedding vector and the temporal feature vector should be ≤0.1 (to ensure modal consistency); for different anomalous samples (such as cutting tool malfunction and uneven feeding), the distance should be ≥0.5 (to enhance discriminability). The energy constraint limits the L2 norm of the candidate fusion vector to ≤5 to avoid training instability caused by excessively large vector magnitudes. The final output is an 896-dimensional joint feature vector optimized by the constraints, while simultaneously storing the time series of this vector and the cross-modal consistency mask. This step forces the model to learn common features between modalities through consistency constraints (such as the visual and temporal representation of malfunctions needing to be consistent), improving the cross-modal interpretability of features; it prevents feature vector explosion through energy constraints, improving model training stability and generalization ability; and it provides a historical baseline for subsequent distribution drift detection by storing the time series.
[0059] This invention achieves significant innovations in multimodal feature association and information complementarity by introducing a dynamic confidence coefficient gating weighting mechanism and a cross-modal consistency mask dual-constraint fusion strategy. On one hand, the dynamic confidence coefficient reflects the health status of temporal features in each channel at different times. By dynamically adjusting the contribution of channel features through gating weighting, it suppresses the interference of low-confidence channels on the fusion result, improving the robustness of cross-modal fusion from the source. On the other hand, the cross-modal consistency mask accurately characterizes the association strength between visual features and temporal features through a cosine similarity matrix, and reduces the feature contribution in low-consistency regions, effectively avoiding the misfusion of irrelevant information between modalities. The dual-path cross-modal attention fusion using global and local time windows preserves long-term dependencies while also considering short-term key information, achieving deep alignment of features in both time and modality dimensions. Superimposing consistency contrast constraints and energy constraints on the fusion output not only improves feature discriminative power and discrimination ability but also enhances the model's generalization performance under distribution variations and noise interference, achieving higher precision and more stable joint feature representation output.
[0060] S50. Input the joint feature vector into the discrimination module containing the classification head and the open set recognition head, perform extreme value theory calibration on the anomaly score to obtain the anomaly severity score, and complete the process compliance judgment in the parallel process knowledge constraint reasoning layer, and output the anomaly type discrimination result and the anomaly severity score.
[0061] Specifically, this step includes: S51. Input the joint feature vector into the discrimination module, which consists of a classification head, an open set identification head, an extreme value theory calibration unit, a process knowledge constraint reasoning layer, and a conflict coordination unit. The following detailed examples and explanations of each step in the production scenario of a traditional Chinese medicine pelleting machine (taking abnormal pellet shape caused by roller wear as an example) are provided. First, the joint feature vector is input to the discrimination module. For example, this joint feature vector is 896-dimensional and incorporates the following information: visual features: irregular pellet edges and surface cracks (from the visual embedding vector); temporal features: increased roller shaft vibration amplitude (0.4 mm / s) and increased current fluctuation (6.5 A → 7.8 A) (from the weighted temporal dynamic features); dynamic confidence coefficients: 0.82 for the vibration channel and 0.90 for the current channel (reflecting data reliability). This joint feature vector is input to the discrimination module, where various units work collaboratively: the classification head identifies known anomalies, the open set identification head detects unknown anomalies, the extreme value theory calibration unit quantifies severity, the process knowledge inference layer verifies compliance, and the conflict coordination unit resolves contradictions.
[0062] S52. The classification head performs category discrimination on the joint feature vector and outputs the classification candidate results and corresponding confidence scores of the known categories. The classification head is a dedicated discrimination module that connects the feature representation and the final category output. It consists of a fully connected layer, a normalization layer and an activation function. Its function is to map the high-dimensional joint feature vector to the category space to complete the category discrimination process. The input joint feature vector is input to the fully connected layer and projected to the same dimension as the preset number of categories. Then, the feature distribution is stabilized by normalization operation, and the projection result is converted into the probability distribution of each category using Softmax. Finally, the classification candidate results and corresponding confidence scores of the known categories are output according to the probability distribution, realizing the mapping from deep features to category labels.
[0063] In this step, the classification head is used to identify known anomalies. For example, the preset known anomaly categories include three types: hob wear, uneven feeding, and abnormal die temperature. After inputting the joint feature vector, the classification head calculates the probability of each category: assuming hob wear: 0.85; uneven feeding: 0.10; abnormal die temperature: 0.05; then the output classification candidate result is: hob wear, corresponding to a confidence level of 0.85. This step allows for rapid location of known anomaly types using the classification head. The classification head, trained based on historical fault data, can efficiently match typical fault characteristics, providing a preliminary judgment basis for subsequent processing.
[0064] S53. The joint feature vector is subjected to open set discrimination by the open set identification head to generate an anomaly score reflecting the degree of deviation of the sample from the normal category, and an anomaly score time series is formed in the order of timestamps. The open set identification head is a discrimination module for unknown category identification and anomaly detection. It consists of a feature normalization layer, a category center vector storage unit, a distance metric calculation unit, and an anomaly scoring function. The working process is as follows: the joint feature vector is normalized and the distance is calculated with the known category center vector established in the training phase. The distance value is mapped to an anomaly score reflecting the degree of deviation of the sample from the normal category using the anomaly scoring function. The anomaly scores are organized into an anomaly scoring time series in the order of timestamps to continuously track the abnormal change trend of the sample in the time dimension.
[0065] In this step, the open set recognition head primarily determines whether a sample belongs to an unknown category based on feature distance. For example, the Euclidean distance between the joint feature vector and the centers of known categories in the training set, such as normal operating conditions and hob wear, is calculated: for instance, distance to the center of normal operating conditions: 0.72 (threshold 0.5, exceeding this is considered abnormal); distance to the center of hob wear: 0.35 (within the known category range). The anomaly scoring formula is: (distance to the nearest known category / maximum allowable distance) × 100, resulting in a score of (0.35 / 0.5) × 100 = 70 points. The anomaly scoring time series is formed by arranging the data by timestamps: [65, 68, 70, 72], corresponding to four consecutive time windows, reflecting the increasing trend of anomaly severity. This step overcomes the limitations of closed categories, generating anomaly scores even for unknown anomalies that have not been trained (such as hob installation misalignment), avoiding missed detections; the time series reflects the dynamic changes of anomalies, providing a temporal basis for severity assessment.
[0066] S54. Using the extreme value theory calibration unit, high-value samples from the tail of the anomaly scoring time series are selected within the sliding time window to construct a tail distribution model. Based on the tail distribution model, a quantile mapping relationship is generated, converting the anomaly score corresponding to each timestamp into an anomaly severity score within a fixed interval. Corresponding alarm and maintenance thresholds are adaptively generated according to the tail distribution model. Specifically, within the sliding time window, anomaly scores are sorted from high to low, and a fixed proportion or number of high-value samples are selected as the tail sample set for normalization. For the normalized tail sample set, a generalized Pareto distribution is used as the tail distribution model. The shape parameter, scale parameter, and location threshold parameter are solved using the maximum likelihood estimation method, where the location threshold is the minimum value of the tail samples. Based on the fitted tail distribution parameters, the quantile values corresponding to each target quantile are calculated, constructing a mapping relationship from anomaly score to quantile probability. The fitting accuracy of the tail distribution model is evaluated through back-substitution testing.
[0067] In this step, as an example, the sliding time window is set to 4 time steps (corresponding to 4 anomaly scores: 65, 68, 70, 72). High-value samples are selected from the tail: sorted in descending order [72, 70, 68, 65], the top 30% of high values (i.e., 72, 70) are taken as tail samples; tail distribution model construction: a generalized Pareto distribution is used for fitting, and the parameters are obtained through maximum likelihood estimation: shape parameter 0.1, scale parameter 5, and location threshold 68; quantile mapping: the anomaly scores are converted into severity scores of 0-100. Adaptive threshold generation: based on the tail distribution, the alarm-level threshold is set to 60 points (corresponding to the upper limit of the 95% confidence interval under normal operating conditions), and the maintenance-level threshold is set to 85 points (corresponding to the point of significant increase in fault risk). This step utilizes extreme value theory to accurately model the tail characteristics (extreme anomalies) of the anomaly scores, avoiding the unsuitability of traditional fixed thresholds for complex operating conditions; the severity score quantifies the degree of anomaly, and the threshold is adaptively adjusted, providing a scientific basis for graded response.
[0068] S55. The process data is matched with the preset process rules through the process knowledge constraint reasoning layer to generate the process compliance judgment and compliance mark with the corresponding timestamp. The process data refers to the parameter information that reflects the process operation status, such as temperature range, current ripple, and pellet geometric size and shape index collected during the pellet making process of traditional Chinese medicine. As an example, suppose the preset process rules include: the normal preset rules for the hobbing speed are 300±5 r / min; the normal shot diameter is 4-6 mm; and the normal mold temperature is 55-60℃. If the current hobbing speed is 310 r / min > 305 r / min, it is judged as "non-compliant" and marked as "hobbing speed overspeed"; if the current mold temperature is 62℃ > 60℃, it is judged as "non-compliant" and marked as "temperature exceeding the limit"; if the current shot diameter is 5 mm, it is compliant and marked as "normal". This step introduces domain-specific process rules (such as speed and temperature thresholds), combining data-driven results with professional knowledge to reduce pure algorithmic misjudgments (e.g., abnormal vibration may be due to normal load fluctuations rather than a fault); compliance marking provides contextual verification for the anomaly type.
[0069] S56. The conflict coordination unit performs a consistency comparison on the classification candidate results, anomaly scores, anomaly severity scores and process compliance judgments. When a conflict of multiple judgment results is detected, the unit makes a decision according to the preset priority and time continuity rules and outputs the anomaly type judgment result with the corresponding timestamp. In this step, when a conflict is detected between multiple source judgment results, a decision is made according to preset priority and time continuity rules, and the anomaly type judgment result corresponding to the timestamp is output. For example, suppose there is the following conflict between multiple source results: Classification header: "Roller wear" (confidence 0.85); Process judgment: "Roller overspeed", "Temperature exceeding limit"; Anomaly severity: 80 points (above the alarm threshold of 60). The following conflict coordination rules can be set: Priority: Process knowledge judgment > Classification header result (because process rules more directly reflect the equipment status); Time continuity: "Roller overspeed" has been detected in the past 3 time windows, with a continuity of 100%, and is given priority. Decision result: The anomaly type is corrected to "Roller overspeed causing wear", retaining the severity of 80 points. This step can resolve contradictions in multi-source judgments (such as classification header misjudgment, sensor noise interference), ensure result consistency through priority and continuity rules, and improve the reliability of anomaly type judgment by combining process context with algorithm results.
[0070] S57. Output the anomaly type discrimination result and the anomaly severity score, and store the anomaly scoring time series, the anomaly severity score time series and the process compliance mark.
[0071] The output of this step includes the specific anomaly type, quantified severity, and process basis, providing clear handling guidance for operation and maintenance personnel; the auxiliary information supports tracing the cause of the anomaly, facilitating subsequent root cause location and linkage control.
[0072] This invention achieves a closed-loop processing across the entire chain, encompassing known category identification, unknown category detection, anomaly severity quantification, and process compliance determination, by introducing a classification head, an open set identification head, an extreme value theory calibration unit, a process knowledge constraint inference layer, and a conflict coordination unit into the discrimination module. The classification head utilizes precise mapping from deep features to the category space to ensure high-accuracy identification of known categories. The open set identification head effectively captures unknown categories and anomalous samples through category center vectors and distance metrics, significantly improving the detection capability for out-of-distribution data. The extreme value theory calibration unit combines generalized Pareto distribution to achieve accurate tail distribution modeling and quantile mapping, enabling anomaly scores to be interpretable and comparable across time windows. The process knowledge constraint inference layer integrates domain process rules with detection results, improving the professionalism and reliability of anomaly type discrimination. The conflict coordination unit resolves multi-source conflicts through priority and temporal continuity strategies, ensuring the consistency and stability of the output results.
[0073] In some implementations, after obtaining the anomaly type determination result and the anomaly severity score, the following steps are also included: S60. Based on the anomaly type discrimination result and anomaly severity score, construct the equipment-process-quality heterogeneity diagram and obtain the root cause localization result through graph neural network, calculate the risk function score, trigger linkage control command according to multi-level threshold and display it visually. Specifically, the system receives anomaly type discrimination results and anomaly severity scores, as well as process data, equipment structure information, process flow data, and quality indicator data, and simultaneously receives dynamic confidence coefficients. Based on the equipment structure information, process flow data, and quality indicator data, it constructs an equipment-process-quality heterogeneous graph, sets equipment component nodes, process nodes, and quality indicator nodes in the graph, and establishes directed edges reflecting time sequence and hierarchical edges reflecting process dependencies. The system maps anomaly type discrimination results, anomaly severity scores, and process data to corresponding nodes and edges according to timestamps, and uses dynamic confidence coefficients to assign weights to nodes related to sensor channels, forming a weighted equipment-process-quality heterogeneous graph. Bidirectional message passing is performed along the time path and process level path on the weighted equipment-process-quality heterogeneous graph. The anomaly correlation score of each node is calculated and a node score set is formed. Specifically, the calculation of the anomaly correlation score of each node refers to performing forward and reverse message passing on the time path and process level path as information propagation channels in the weighted equipment-process-quality heterogeneous graph. The state information of adjacent nodes and path weights are weighted and aggregated. The representation vector of the current node is updated in combination with the node's own characteristics. The similarity or correlation index is calculated based on the updated node representation vector and the preset anomaly pattern vector to obtain a value reflecting the degree of association between the node and the anomaly pattern. Finally, the complete node score set is formed. The current root cause localization result is determined based on the node score set, and the stability ratio of the root cause is statistically analyzed within a set continuous time window. The stability ratio is the percentage of times the root cause result remains consistent across multiple consecutive time windows. Based on the anomaly probability corresponding to the anomaly type discrimination result, the standardized value of the anomaly severity score, the loss weight coefficient obtained from the quality indicator deviation, the root cause stability ratio, and the weighted average of the top-level root cause node scores, the risk function value is calculated. A multi-level risk threshold set and corresponding hysteresis are set. When the risk function value reaches the first-level threshold and the time for continuously meeting the threshold condition is not less than the minimum trigger duration corresponding to the first-level threshold, the linkage control command corresponding to the risk level is executed, including speed reduction, encrypted sampling, shutdown, or maintenance. Before executing the linkage control command, a safety interlock verification and whitelist check are performed. The root cause location result, root cause stability ratio, risk function value and the linkage control command to be executed are recorded together. The root cause node, root cause path and location hot zone are displayed on the visualization interface. At the same time, the cooldown execution time of the linkage control command is set.
[0074] S70. Store and alarm the root cause localization results, linkage control commands, multi-channel time series data and joint feature vectors, compare the current and historical feature distribution changes to determine whether release drift has occurred, and collect and organize cross-membrane state scarce fault sample data in the offline stage.
[0075] Specifically, the root cause localization results, linkage control commands, multi-channel time-series data, and joint feature vectors are archived and stored according to a unified data structure. Real-time alarms are triggered when an anomaly is detected, and alarm information is synchronously pushed to the control platform and operation and maintenance terminal. A baseline is established based on the stored joint feature vectors and historical feature distributions. The degree of difference between the current operating cycle and the historical reference distribution is compared. When the degree of difference exceeds a preset threshold, it is determined that a distribution drift has occurred, and the time of the drift and the scope of its impact are recorded. In the offline stage, the multi-channel time-series data, joint feature vectors, and corresponding anomaly labels during the distribution drift period are sorted and filtered to remove invalid or noisy samples and retain effective data segments that can characterize cross-modal feature differences. The sorted dataset is labeled and classified according to cross-modal alignment rules to generate a cross-modal scarce fault sample dataset.
[0076] Based on the heterogeneous graph structure of equipment-process-quality and the dual-path message passing algorithm, this invention enables the system to achieve accurate root cause localization and trigger multi-level threshold control strategies based on risk assessment results, thereby reducing unnecessary downtime while ensuring production safety.
[0077] The method of the present invention will be further explained and illustrated below through specific embodiments: Example 1 To verify the feasibility of this invention in practice, it was applied to a large-scale traditional Chinese medicine pharmaceutical factory. The factory previously used a threshold alarm system based on a single vibration and temperature sensor, which could only provide preliminary warnings for some mechanical faults. It was almost ineffective in identifying abnormalities in pellet forming quality, unknown types of faults, and inconsistencies in multi-source signals. Especially with the production line operating continuously in three shifts, short-term anomalies occurring during the night shift were often amplified into equipment shutdowns or even large-scale scrapping due to delays in manual inspections, resulting in production losses.
[0078] After introducing the method of this invention, high-definition industrial cameras and multiple types of sensors are installed at key workstations of each traditional Chinese medicine pelleting machine, covering multi-channel signal acquisition including vibration, sound, current, and temperature. All process data from the pellet forming area (such as forming temperature, pellet diameter, roundness index, and current ripple amplitude) are included in the real-time monitoring range. All acquired signals are synchronized using a unified timestamp, ensuring comparison and analysis of multimodal data at the same time point at the millisecond level. A vision converter analyzes the image sequence in real time to identify pellet surface defects, size deviations, and color anomalies; a bidirectional long short-term memory network analyzes multi-channel time-series signals to capture vibration mode changes, abnormal noise characteristics, and current fluctuations.
[0079] In the fusion stage, this invention introduces a dynamic confidence coefficient weighting mechanism, enabling the system to automatically reduce the impact of a channel on the fusion result and maintain the stability of the overall judgment when short-term signal fluctuations or loss occur in the sensor. A cross-modal consistency mask is used to filter feature regions with high correlation between the image and the sensor signal, and dual-path attention fusion further extracts global and local anomaly features. In the discrimination stage, the classification head is responsible for identifying known anomaly types, such as uneven feeding, cutter failure, and mold jamming, while the open set recognition head identifies novel anomalies not found in the training set, and an anomaly severity score from 0 to 100 is obtained through extreme value theory calibration.
[0080] During a night shift production run, the system detected a minor appearance anomaly in the pellet forming area, accompanied by a slow rise in forming temperature and unstable current ripple. The anomaly severity score was 72, triggering a level-two response in the multi-level threshold system. This automatically issued commands to reduce speed and increase sampling density. Simultaneously, the root cause was located in the equipment-process-quality heterogeneity graph, indicating a fault in the mold temperature control circuit. Maintenance personnel completed the inspection and replacement of the temperature control module within 5 minutes of receiving the command, preventing a potential batch of defective products. According to production records, the processing time for this anomaly was reduced by an average of approximately 68% compared to the original system, and the scrap rate was reduced by approximately 54%.
[0081] Table 1. Comparison of production performance of pelleting machines before and after application of this invention.
[0082] As shown in Table 1, in the three months before the system was introduced, the pelleting machine experienced 29 anomalies during operation. The accuracy rate for detecting known anomalies was 91.4%, the recall rate for unknown anomalies was 0%, and the average response time was 9.8 minutes, resulting in a cumulative downtime of 36.5 hours. The pellet appearance defect rate and weight deviation exceedance rate were 0.83% and 0.74%, respectively. In the first month after the method of this invention was introduced, the total number of anomalies significantly decreased to 11, the accuracy rate for detecting known anomalies increased to 97.6%, the recall rate for unknown anomalies reached 90.1%, the average response time was shortened to 3.2 minutes, the downtime was reduced to 10.8 hours, the pellet appearance defect rate decreased to 0.42%, and the weight deviation exceedance rate decreased to 0.33%. The metrics continued to improve in the second and third months, with the recall rate for unknown anomalies reaching 92.7% and 94.4% respectively. The average response time was further reduced to 3.0 minutes, downtime decreased to 8.5 hours, and the appearance defect rate and weight deviation exceedance rate decreased to 0.36% and 0.28%, and 0.37% and 0.31%, respectively. Compared with the three-month average data before its introduction, the accuracy rate for detecting known anomalies increased by 6.7 percentage points, the recall rate for unknown anomalies increased by 92.4 percentage points, the average response time decreased by 68.4%, the total downtime decreased by 73.9%, and the pellet appearance defect rate and weight deviation exceedance rate decreased by 54.2% and 58.1%, respectively.
[0083] In summary, the method of this invention improves the comprehensiveness and accuracy of anomaly detection, shortens response time, significantly reduces downtime losses, and improves the appearance quality and weight consistency of pellet products, fully verifying its feasibility in actual production.
[0084] The online anomaly detection method for a traditional Chinese medicine pill-making machine in the embodiments of the present invention has been described above. The online anomaly detection device for a traditional Chinese medicine pill-making machine in the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 2 One embodiment of the online anomaly detection device for a traditional Chinese medicine pill-making machine in this invention includes: Alignment module 10 is used to collect multi-channel time-series data of the Chinese medicine pill-making machine during the production process, assign a unified timestamp mark and complete synchronous alignment. The multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals and temperature signals. The visual embedding vector generation module 20 is used to preprocess the image sequence in the multi-channel time series data, extract features from the preprocessed image sequence through the visual transformer model, and generate a visual embedding vector by calculating the extracted features through a multi-layer self-attention mechanism. The dynamic feature extraction module 30 is used to filter, normalize and construct sequences of vibration signals, sound signals, current signals and temperature signals in multi-channel time series data, extract time series dynamic feature representations through a bidirectional long short-term memory network, calculate sensor health scores and generate dynamic confidence coefficients; The fusion module 40 is used to input the visual embedding vector, temporal dynamic feature representation and dynamic confidence coefficient into the feature fusion layer, first perform gated weighted fusion, then perform cross-membrane attention fusion, and introduce cross-membrane consistency contrast constraint and energy constraint during the training phase to output joint feature vector; The anomaly output module 50 is used to input the joint feature vector into the discrimination module, which includes the classification head and the open set recognition head, to perform extreme value theory calibration on the anomaly score to obtain the anomaly severity score, and to complete the process compliance judgment in the parallel process knowledge constraint reasoning layer, and output the anomaly type discrimination result and the anomaly severity score.
[0085] Based on the same ideas as the methods in the above embodiments, the device provided by this application can implement the methods in the above embodiments. For ease of explanation, the structural schematic diagram of the device embodiment only shows the parts related to the embodiments of this application. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer modules than illustrated, or combine certain modules, or have different module arrangements.
[0086] Figure 2 The online anomaly detection device for the traditional Chinese medicine pill-making machine in this embodiment of the invention is described in detail from the perspective of modular functional entities. The online anomaly detection device for the traditional Chinese medicine pill-making machine in this embodiment of the invention is described in detail from the perspective of hardware processing.
[0087] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 111 (e.g., one or more processors) and a memory 121, and one or more storage media 130 (e.g., one or more mass storage devices) for storing application programs 133 or data 132. The memory 121 and storage media 131 can be temporary or persistent storage. The program stored in the storage media 130 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the electronic device 100. Furthermore, the processor 111 may be configured to communicate with the storage media 130 and execute the series of instruction operations in the storage media 130 on the electronic device 100.
[0088] Electronic device 100 may also include one or more power supplies 141, one or more wired or wireless network interfaces 151, one or more input / output interfaces 161, and / or one or more operating systems 131, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The device structure shown does not constitute a limitation on the electronic device 100, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0089] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the online anomaly detection method for a traditional Chinese medicine pill-making machine.
[0090] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0091] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0092] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for online anomaly detection in a traditional Chinese medicine pill-making machine, characterized in that, Including the following steps: Collect multi-channel time-series data of the Chinese medicine pill-making machine during the production process, assign a unified timestamp mark and complete synchronization alignment. The multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals and temperature signals. The image sequences in the multi-channel time series data are preprocessed, and the preprocessed image sequences are feature extracted through a visual transformer model. The extracted features are then calculated to generate visual embedding vectors through a multi-layer self-attention mechanism. The vibration, sound, current and temperature signals in the multi-channel time series data are filtered, normalized and sequenced. The time series dynamic features are extracted through a bidirectional long short-term memory network, and the sensor health score is calculated and dynamic confidence coefficient is generated. Visual embedding vectors, temporal dynamic feature representations, and dynamic confidence coefficients are input into the feature fusion layer. First, gated weighted fusion is performed, then cross-membrane attention fusion is executed. Cross-membrane consistency contrast constraints and energy constraints are introduced during the training phase to output a joint feature vector. The joint feature vector is input into the discrimination module, which includes a classification head and an open set recognition head. The anomaly score is calibrated using extreme value theory to obtain the anomaly severity score. In the parallel process knowledge constraint reasoning layer, the process compliance determination is completed, and the anomaly type discrimination result and anomaly severity score are output.
2. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 1, characterized in that, The process of preprocessing image sequences in multi-channel time-series data, extracting features from the preprocessed image sequences using a visual transformer model, and generating visual embedding vectors from the extracted features using a multi-layer self-attention mechanism includes the following steps: Denoising is performed on the image sequence of multi-channel time-series data to obtain a denoised image sequence, with each frame of the image corresponding to the acquisition timestamp; The denoised image sequence is normalized by linearly scaling the pixel values according to the preset minimum and maximum pixel value range to obtain the normalized image sequence. The normalized image sequence is divided into blocks according to a fixed pixel size to obtain a set of multiple numbered image blocks; The set of image patches is input into the visual transformer model, and each image patch is converted into an image patch embedding vector through the embedding mapping process; Add a location-corresponding encoding vector to each image patch embedding vector to obtain a set of image patch embedding vectors containing location information; The set of image patch embedding vectors containing location information is input into a multi-layer self-attention mechanism for calculation to obtain a visual embedding vector, which represents the image spatial features under the corresponding timestamp.
3. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 2, characterized in that, The set of image patch embedding vectors containing location information is input into a multi-layer self-attention mechanism for computation to obtain a visual embedding vector, including the following steps: The image patch embedding vector set containing location information is input into the first layer of the multi-layer self-attention mechanism, and the weighted feature representation is obtained by calculating the correlation between features. The feature representations output from the first layer are passed sequentially to the computational layers of the multi-layer self-attention mechanism. In each layer, the weight distribution between each feature is continuously updated, gradually enhancing the ability to model global spatial dependencies. After completing the calculation of all multi-layer self-attention mechanisms, the final visual embedding vector is output. The visual embedding vector integrates global context information and represents the image spatial features under the corresponding timestamp.
4. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 1, characterized in that, The vibration, sound, current, and temperature signals from multi-channel time-series data are filtered, normalized, and sequenced. Temporal dynamic features are extracted using a bidirectional long short-term memory network. Sensor health scores are calculated, and dynamic confidence coefficients are generated. The process includes the following steps: Vibration signals, sound signals, current signals, and temperature signals in multi-channel time-series data are filtered, normalized, and sequenced to form a channel sequence arranged by timestamps. The channel sequence is then input into a bidirectional long short-term memory network to obtain a time-series dynamic feature representation corresponding to the timestamps. Within a sliding time window, the amplitude interval dwell ratio, the rate of change of the spectral centroid, and the envelope mutation count are statistically analyzed to generate a stability sub-score. Using adjacent channels of the same type and redundant channels as references, the consistency of the ordering, peak synchronization and event overlap of the time-series dynamic features are compared to generate a consensus sub-score. Call the time-frequency fingerprint database of normal and fault conditions established during the debugging phase, perform fingerprint matching and classification on the current time window, and generate fingerprint matching sub-scores; Perform data integrity verification, detect missing samples, communication delays, saturation shearing, and range out-of-bounds errors, and output integrity flags; The stability sub-score, consensus sub-score, fingerprint matching sub-score, and integrity marker are sorted, voted, and fused to obtain the sensor health score. The sensor health score is hierarchically mapped and time-smoothed to generate a dynamic confidence coefficient. A one-to-one correspondence is established between the coefficient and the time-series dynamic feature representation, and the time series of the sensor health score and the dynamic confidence coefficient are recorded.
5. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 1, characterized in that, The visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficient are input into the feature fusion layer. First, gated weighted fusion is performed, followed by cross-membrane attention fusion. During the training phase, cross-membrane consistency contrast constraints and energy constraints are introduced to output a joint feature vector. The steps include: The visual embedding vector, temporal dynamic feature representation, and dynamic confidence coefficient are input into the feature fusion layer; The time-series dynamic feature representation is gated and weighted using dynamic confidence coefficients, so that the time-series dynamic feature representation of each channel is adjusted by the corresponding dynamic confidence coefficient, thus generating a weighted time-series dynamic feature representation. Within a sliding time window, a time alignment operation is performed on the visual embedding vector and the weighted temporal dynamic feature representation. The alignment strategy includes time delay compensation and missing segment interpolation to generate an aligned temporal dynamic feature representation. The cross-modal similarity matrix is calculated based on the visual embedding vector and the aligned temporal dynamic feature representation, and a cross-modal consistency mask is generated according to the similarity threshold. Cross-modal attention fusion is performed on the visual embedding vector and the aligned temporal dynamic feature representation in the global time window path and the local time window path, respectively, to obtain the fused output vector of the two paths; The fusion output vectors of the two paths are concatenated and merged, and the feature contribution is reduced in the low value region of the cross-modal consistency mask. At the same time, the fusion result is weighted again according to the dynamic confidence coefficient to obtain the candidate fusion vector. During the training phase, cross-modal consistency contrast constraints are applied to the candidate fusion vectors, and energy constraints are applied to the candidate fusion vectors. The constrained candidate fusion vectors are output as joint feature vectors, and the time series of the joint feature vectors and cross-modal consistency masks are stored together.
6. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 1, characterized in that, The joint feature vector is input into the discrimination module, which includes a classification head and an open set recognition head. The anomaly score is calibrated using extreme value theory to obtain an anomaly severity score. Then, in the parallel process knowledge constraint inference layer, process compliance determination is completed, and the anomaly type discrimination result and anomaly severity score are output. The steps include: The joint feature vector is input to the discrimination module, which consists of a classification head, an open set identification head, an extreme value theory calibration unit, a process knowledge constraint reasoning layer, and a conflict coordination unit. The classification head is used to classify the joint feature vector, and the classification candidate results and corresponding confidence scores of the known categories are output. The joint feature vector is subjected to open set discrimination by the open set identification head, and an anomaly score reflecting the degree of deviation of the sample from the normal category is generated, forming an anomaly score time series arranged in timestamp order; The extreme value theory calibration unit selects the tail high value samples of the abnormal scoring time series within the sliding time window, constructs the tail distribution model, generates the quantile mapping relationship based on the tail distribution model, converts the abnormal score corresponding to each time stamp into an abnormal severity score within a fixed interval, and adaptively generates the corresponding alarm level threshold and maintenance level threshold according to the tail distribution model. The process data is matched with the preset process rules through the process knowledge constraint reasoning layer to generate process compliance judgment and compliance mark with corresponding timestamp. The process data refers to the parameter information that reflects the process operation status, such as temperature range, current ripple, and pellet geometric size and shape index collected during the pellet making process of traditional Chinese medicine. The conflict coordination unit performs a consistency comparison of the classification candidate results, anomaly scores, anomaly severity scores, and process compliance judgments. When a conflict of multiple judgment results is detected, a decision is made according to a preset priority and time continuity rule, and the anomaly type judgment result corresponding to the timestamp is output. The anomaly type determination result and the anomaly severity score are output.
7. The online anomaly detection method for a traditional Chinese medicine pill-making machine according to claim 1, characterized in that, After obtaining the anomaly type identification results and anomaly severity scores, the following steps are also included: Based on the anomaly type discrimination results and anomaly severity scores, an equipment-process-quality heterogeneity diagram is constructed, and the root cause localization results are obtained through a graph neural network. The risk function score is calculated, and linkage control commands are triggered according to multi-level thresholds and displayed visually. The root cause localization results, linkage control commands, multi-channel time series data, and joint feature vectors are stored and alarmed. The current and historical feature distribution changes are compared to determine whether a release drift has occurred. In the offline stage, rare fault sample data across membrane states are collected and organized.
8. An online anomaly detection device for a traditional Chinese medicine pill-making machine, characterized in that, include: The alignment module is used to collect multi-channel time-series data of the Chinese medicine pill-making machine during the production process, assign a unified timestamp mark and complete synchronous alignment. The multi-channel time-series data includes image sequences, vibration signals, sound signals, current signals and temperature signals. The visual embedding vector generation module is used to preprocess image sequences in multi-channel time-series data. It extracts features from the preprocessed image sequences through a visual transformer model and calculates and generates visual embedding vectors from the extracted features through a multi-layer self-attention mechanism. The dynamic feature extraction module is used to filter, normalize, and construct sequences of vibration signals, sound signals, current signals, and temperature signals in multi-channel time-series data. It extracts time-series dynamic feature representations through a bidirectional long short-term memory network, calculates sensor health scores, and generates dynamic confidence coefficients. The fusion module is used to input the visual embedding vector, temporal dynamic feature representation and dynamic confidence coefficient into the feature fusion layer. First, gated weighted fusion is performed, then cross-membrane attention fusion is performed, and cross-membrane consistency contrast constraint and energy constraint are introduced during the training phase to output joint feature vector. The anomaly output module is used to input the joint feature vector into the discrimination module, which includes the classification head and the open set recognition head. It performs extreme value theory calibration on the anomaly score to obtain the anomaly severity score, and completes the process compliance determination in the parallel process knowledge constraint reasoning layer, outputting the anomaly type discrimination result and the anomaly severity score.
9. An electronic device, characterized in that, It includes a memory and at least one processor, wherein the memory stores computer-readable instructions; The at least one processor invokes the computer-readable instructions in the memory to perform the steps of the online anomaly detection method for a traditional Chinese medicine pill-making machine as described in any one of claims 1-7.
10. A computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement each step of the online anomaly detection method for a traditional Chinese medicine pill-making machine as described in any one of claims 1-7.
Citation Information
Patent Citations
Electrocardiogram data anomaly recognition method, device and equipment based on deep learning and storage medium
CN119442124A
Abnormality detection method, device and equipment for time series data and storage medium
CN120296622A
Abnormal behavior detection method and system for power sensitive data flow
CN120541717A
Multi-modal information tagging method, apparatus and device, and storage medium and product
WO2025148651A1
Cited By
Evaluation method and system of multi-mode deep pseudo data discrimination algorithm
CN121434688A
Quality defect detection method based on aluminum alloy profile for rail transit conductor rail
CN121595562A
Method for detecting quality defects of aluminum alloy profile for rail transit conductive rail
CN121595562B