Granule-based computing and attention mechanism-based cross-modal anomaly detection method and system
Patent Information
- Application Number
- CN202610865742.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-16
AI Technical Summary
[0005]针对现有技术中面对海量跨模态异构数据时存在的计算效率低下、特征融合困难等技术问题,本发明的目的在于提供一种基于粒球计算与注意力机制的跨模态异常检测方法及系统
[0018] By adopting the above scheme, the present invention has the following advantages and beneficial effects: (1) The present invention introduces particle sphere computing to process cross-modal heterogeneous data. By converting massive scattered data into a finite number of feature particles for representation, the computational complexity of distance measurement is greatly reduced, effectively overcoming the dimensionality curse of traditional methods when processing large-scale industrial data, and significantly improving the real-time performance of anomaly detection. (2) The present invention combines a multi-head attention mechanism to deeply fuse feature particles of different modalities, breaking down the modal barriers between heterogeneous data. It can adaptively capture the deep correlation of cross-modal features in spatiotemporal evolution, and improve the discrimination accuracy of minor anomalies and early faults in complex scenarios.
Smart Images

Figure CN122388658B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of industrial Internet of Things and data mining technology, specifically to a cross-modal anomaly detection method and system based on particle-sphere computing and attention mechanisms. Background Technology
[0002] With the rapid development of the Industrial Internet of Things (IIoT) and cyber-physical systems, complex equipment generates massive amounts of multi-source monitoring data during operation. This data typically exhibits significant cross-modal and heterogeneous characteristics, including both continuous multivariate time-series data (such as sensor temperature and vibration frequency) and discrete status data (such as equipment start-up and shutdown records and error alarm logs). How to quickly and accurately identify abnormal behavior or potential faults in equipment from this complex cross-modal data has become a key challenge in ensuring the safe operation of industrial systems.
[0003] Most existing anomaly detection methods are designed for single-modality or homogeneous data. When faced with massive cross-modal data, existing technologies typically suffer from two significant drawbacks: First, low computational efficiency. Traditional methods often employ point-by-point traversal calculations when calculating similarity or distance measures for heterogeneous data. When dealing with large-scale industrial time-series data, this is highly susceptible to the curse of dimensionality, resulting in extremely poor real-time performance and failing to meet the low-latency requirements of industrial early warning systems. Second, difficulty in cross-modal feature fusion. Because different modalities reside in different feature spaces, existing detection systems often only perform simple physical splicing, ignoring the deep correlations between different modal features during complex spatiotemporal evolution, leading to severely insufficient accuracy in identifying minor anomalies or early faults.
[0004] Therefore, there is an urgent need for an anomaly detection scheme that can efficiently process large-scale heterogeneous data and accurately integrate multimodal features, so as to break through the bottlenecks of existing technologies in terms of real-time performance and accuracy. Summary of the Invention
[0005] To address the technical problems of low computational efficiency and difficulty in feature fusion when dealing with massive cross-modal heterogeneous data in existing technologies, the present invention aims to provide a cross-modal anomaly detection method and system based on particle-sphere computation and attention mechanism.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: a cross-modal anomaly detection method based on particle-sphere computation and attention mechanism, comprising the following steps: S1. Acquire and preprocess cross-modal monitoring data: Collect continuous multivariate time series data and discrete state data from the target system, align them by time dimension and slice them by sliding window, and construct a cross-modal heterogeneous feature dataset. S2. Generate modal feature spheres: Map the cross-modal heterogeneous feature dataset to a multi-dimensional feature space, generate feature spheres specific to each modality based on sphere calculation and adaptive splitting methods, and extract the center coordinates and radius of each feature sphere; S3, Cross-modal attention computation and feature fusion: Construct a multi-head attention mechanism module to calculate the spatial correlation weights between feature particles of different modalities, and generate a cross-modal joint particle feature matrix through weighted fusion and channel splicing; S4. Anomaly Factor Assessment: Calculate the Euclidean distance from the features of the online test sample to the nearest neighbor cross-modal joint particle center, and calculate the anomaly factor score based on the relative relationship between this distance and the corresponding particle radius; S5. Anomaly Warning and Status Determination: The anomaly factor score is compared with the dynamic safety threshold set based on historical normal samples to generate a graded warning result.
[0007] Preferably, step S1 includes the following steps: (1) Collect cross-modal heterogeneous data: Read cross-modal raw data in real time from the sensor network and control system of the target industrial IoT device. The data includes continuous multi-dimensional time series data in the physical space and discrete heterogeneous state data in the information space. (2) Data cleaning and spatiotemporal alignment: Lagrange interpolation is used to fill in missing values in multivariate time series data; feature embedding technology is used to map discrete heterogeneous state data into numerical feature vectors; then, spatiotemporal alignment logic is introduced based on the highest sampling rate of continuous time series to broadcast discrete state data to the corresponding timestamp interval. (3) Sliding window truncation and normalization: In order to capture the temporal dynamic evolution characteristics of the data, a sliding window with a length of 1000 mm is used. Step size is A sliding time window is used to slice the time-aligned data; a maximum-minimum normalization method is used to eliminate the differences in physical dimensions between multi-source heterogeneous data. The normalization formula is:
[0008] in, Indicates the first The first feature dimension One original data value, and Each represents the number of elements within the current sliding window. Minimum and maximum values in each feature dimension The features are normalized; finally, the preprocessed data is packaged into a cross-modal heterogeneous feature dataset.
[0009] Preferably, step S2 includes the following steps: (1) Modal feature space mapping: The cross-modal heterogeneous feature dataset is independently input into the corresponding feature extraction network according to the modal category and mapped to the multi-dimensional feature space; in the initialization stage, the set of all sample points under the same modality is regarded as a global feature sphere; (2) Adaptive calculation of particle center and radius: For the first Characteristic spheres Its geometric center coordinates and its boundary radius The calculation formula is:
[0010]
[0011] in, For the first The total number of data points contained within each sphere To traverse the local index of the feature samples inside the sphere, This represents the coordinates of each sample point within the sphere from the center. The Euclidean distance; (3) Granulocyte splitting and joint evolution: Define the maximum radius tolerance threshold of the granule. With variance tolerance threshold When the current particle is detected radius Or the variance of the characteristic distribution within the granule If the distribution of data within the particle does not meet the high cohesion requirement, the K-Means++ clustering algorithm is used to split it into multiple sub-particles. This joint evolution and splitting process is executed iteratively until the radius and variance of all sub-particles in the current mode meet the preset threshold conditions.
[0012] Preferably, step S3 includes the following steps: (1) Spatiotemporal feature linear mapping: Extract the center coordinates of each modal feature sphere generated in step S2. This serves as the representative feature vector of the particle; subsequently, the feature particle vector of the first mode is multiplied by a learnable weight matrix and mapped to a query vector. The feature particle vectors of the second mode are then multiplied by their corresponding weight matrices and mapped to key vectors. Sum value vector In this way, a cross-modal attention query mechanism is constructed. The first modality is the modality corresponding to the group evolution characteristics of the continuous physical operating state of the equipment when the target system generates continuous multivariate time series data in the physical space. The second modality refers to the discrete state data generated by the system in the information space, and its feature spheres reflect the modality corresponding to the semantic clustering characteristics of discrete events after text embedding technology. (2) Multi-head attention space correlation calculation: The feature mapping space is divided into A separate attention head; in the first In each attention head, the original correlation degree between cross-modal feature particles is calculated. After scaling by the square root dimension and applying Softmax activation, the local cross-modal fusion feature matrix output by the m-th attention head is obtained. The calculation formula is as follows:
[0013] in, , , The first Each attention head corresponds to the query vector. Key vector Sum value vector The linear mapping parameter matrix, The feature dimension of the key vector is the size of the space. Normalized exponential activation function; (3) Cross-modal feature fusion: fusing all The output of each attention point is obtained through a concatenation function. Perform feature concatenation and multiply by the global output weight matrix. Generate a cross-modal joint particle-sphere feature matrix that combines temporal evolution and spatial topological characteristics. The formula is: .
[0014] Preferably, step S4 includes the following steps: (1) Nearest neighbor particle matching: real-time online sample data to be tested is obtained, and after the same spatiotemporal alignment and normalization preprocessing as in step S1, the current numerical feature vector to be tested is generated. The cross-modal joint particle feature matrix output in step S3 In the corresponding set of joint particles, calculate The Euclidean distances to the centers of each joint particle are used to filter out the target matching particle with the smallest distance and extract its center. With radius ; (2) Quantitative calculation of abnormal factors: Based on the relative degree of deviation of the test sample from the target matching particle, an abnormal factor score is constructed. The calculation formula is as follows:
[0015] (3) For the score of this abnormal factor ,like This indicates that the sample to be tested falls completely within the envelope boundary covered by normal characteristic spheres, possessing normal confidence level; if The ratio then increases linearly, reflecting the degree of local outliers in the joint feature space of heterogeneous data.
[0016] Preferably, step S5 includes the following steps: (1) Hyperparameter optimization and threshold setting: During the model training or calibration phase, a grid search strategy is used to optimize the hyperparameters of historical normal data, calculate the set of abnormal factor scores of historical normal samples, and extract their mean. with standard deviation Based on Chebyshev's inequality or confidence interval rule, a dynamic safety threshold is set. ; (2) Multi-level early warning generation: The abnormal factor scores of online test samples are generated. With dynamic security threshold Compare with its multiples; if This triggers a Level 1 observation and warning; if This triggers a Level 2 severe fault warning and simultaneously outputs the source modality dimension that caused the abnormally high score.
[0017] A cross-modal anomaly detection system based on particle-sphere computation and attention mechanism, the system being used to implement the cross-modal anomaly detection method based on particle-sphere computation and attention mechanism as described above, the system comprising: Data preprocessing module: used to acquire cross-modal monitoring data and perform time alignment and sliding window slicing to construct a cross-modal heterogeneous feature dataset; Particle sphere generation module: used to map the dataset to a multi-dimensional feature space, generate feature particles of each modality based on particle sphere calculation and adaptive splitting method, and extract center and radius parameters; Feature fusion module: Used to construct the multi-head attention mechanism module, calculate the spatial correlation weights between feature spheres of different modalities, and output the cross-modal joint sphere feature matrix; Anomaly assessment module: used to calculate the relative distance between the online test sample and the nearest neighbor cross-modal joint particle center, and output anomaly factor score; Early warning decision module: used to compare the abnormal factor scores with dynamic safety thresholds and generate graded abnormality early warning results.
[0018] By adopting the above scheme, the present invention has the following advantages and beneficial effects: (1) The present invention introduces particle sphere computing to process cross-modal heterogeneous data. By converting massive scattered data into a finite number of feature particles for representation, the computational complexity of distance measurement is greatly reduced, effectively overcoming the dimensionality curse of traditional methods when processing large-scale industrial data, and significantly improving the real-time performance of anomaly detection. (2) The present invention combines a multi-head attention mechanism to deeply fuse feature particles of different modalities, breaking down the modal barriers between heterogeneous data. It can adaptively capture the deep correlation of cross-modal features in spatiotemporal evolution, and improve the discrimination accuracy of minor anomalies and early faults in complex scenarios. Attached Figure Description
[0019] Figure 1 A flowchart illustrating the overall steps of a cross-modal anomaly detection method based on particle-sphere computation and attention mechanism provided in this embodiment of the invention; Figure 2 This is a schematic diagram of spatiotemporal alignment and sliding window slicing preprocessing for cross-modal heterogeneous data provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the bottom-up adaptive splitting and evolution mechanism of characteristic spheres provided in an embodiment of the present invention; Figure 4 This is a diagram of a cross-modal feature depth joint alignment network architecture based on a multi-head attention mechanism provided in an embodiment of the present invention. Figure 5 This is a logic diagram for anomaly factor calculation and multi-level early warning determination based on nearest neighbor particle matching provided in an embodiment of the present invention; Figure 6 This is a block diagram of a virtual module structure for a cross-modal anomaly detection system based on particle-sphere computation and attention mechanism, provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Reference manual attached Figure 1 The cross-modal anomaly detection method based on particle-sphere computation and attention mechanism disclosed in this embodiment is specifically implemented by the central processing unit or graphics processing unit of the computing device, and includes the following detailed steps: Step 1: Acquisition, cleaning, and spatiotemporal standardization preprocessing of cross-modal heterogeneous data. Specifically, the system's underlying acquisition module first reads the target system's cross-modal raw data in batches through common industrial interfaces (such as OPCUA, MQTT, etc.). This cross-modal raw data spans both physical and information spaces: in the physical space, it manifests as continuous multivariate time-series data collected by high-frequency sensors (e.g., bearing temperature, housing vibration acceleration, rotor three-phase current, etc.); in the information space, it manifests as discrete heterogeneous state data generated by the console or operating system (e.g., system fault alarm logs, switch action commands, maintenance event records, etc.).
[0022] To address missing values in multivariate time-series continuous data due to network jitter or occasional sensor hardware failures during the acquisition process, this embodiment preferably employs Lagrange interpolation for mathematical completion. For discrete, heterogeneous state data, a pre-trained text feature embedding network (such as Word2Vec or a lightweight BERT model) is used to statically map it into a low-dimensional, numerically dense feature vector.
[0023] Subsequently, to eliminate the impact of inconsistent sampling frequencies between different sensor modes, this embodiment introduces timestamp-based broadcast alignment logic. Using the sampling benchmark of the high-frequency continuous time series as the main axis, discrete state feature vectors are broadcast and filled into their corresponding time windows, achieving precise alignment of heterogeneous data in the time dimension. After alignment, a dynamic time window of length W and sliding step size S (where W and S are preset positive integer parameters) is constructed. The aligned cross-modal full data is then sliced to effectively capture the dynamic temporal characteristics of the data stream during its spatiotemporal evolution.
[0024] Furthermore, the mathematical expression for slicing the cross-modal full data using the sliding time window is as follows: Let the aligned cross-modal full data sequence be... ,in This is the total time step. The length elapsed is... Step size is After slicing the sliding window, the nth dynamic time window subsequence is extracted. Represented as:
[0025] in, Let be the index identifier of the sliding window, and satisfy . Through this slicing operation, the originally continuous and infinite data stream is transformed into multiple discrete local spatiotemporal feature matrices, which serve as the standard input for subsequent particle-sphere evolution.
[0026] Furthermore, due to temperature (degrees Celsius), vibration ( The dimensions of different physical indicators such as current (amperes) are vastly different. Directly calculating the distance would cause features in the high-value range to overwhelm features in the low-value range. Therefore, in this embodiment, the maximum and minimum values of each feature matrix after the sliding window is truncated are normalized. The normalization formula is shown in equation (1): (1) In equation (1), It has a unique physical meaning, specifically used to represent the current sliding window, the first... The numerical value of the i-th original data point in each feature dimension; This indicates that within the sliding window, the first... The minimum value of all original data points under each feature dimension; The corresponding value represents the position within the sliding window. The maximum value of all original data points under each feature dimension; These are standardized feature values mapped to the [0,1] interval after dimensionless processing. By traversing and executing equation (1), a structured cross-modal heterogeneous feature dataset is finally constructed.
[0027] Furthermore, to meet the training specifications and online evaluation requirements of machine learning models, this embodiment, after completing the above normalization preprocessing, divides the constructed cross-modal heterogeneous feature dataset into a training set, a validation set, and a test set according to the chronological order of the time series (or a random sampling ratio). The training set mainly consists of cross-modal feature data under historical normal operating conditions, specifically used to drive the initialization and adaptive splitting of feature spheres in subsequent steps, and to train the spatial mapping matrix parameters of the multi-head attention mechanism. The validation set is used to optimize hyperparameters such as the maximum radius threshold Rmax of the spheres and the number of attention heads H using a grid search strategy. The test set is used to simulate a real industrial online environment, inputting into the system to independently evaluate the accuracy of anomaly factor calculation and the generalization performance of multi-level early warning.
[0028] Step 2: Construct modal feature granular balls and perform coarse-grained adaptive information extraction. Traditional anomaly detection algorithms (such as KNN or isolated forests) typically require traversing the neighborhood distances between pairs of objects on the original scatter dataset, and their computational complexity increases exponentially with the amount of data samples. To overcome this computational bottleneck, this embodiment independently splits the normalized cross-modal heterogeneous feature dataset according to modal categories and projects them into their respective independent feature spaces. Granular-ball computing technology is used to achieve a coarse-grained, efficient, and adaptive representation of massive scatter points.
[0029] The bottom-up (or top-down) adaptive generation process of the feature sphere is as follows: In the algorithm initialization stage, the set of all standardized data sample points under a certain single mode in the current feature space is directly regarded as a unique initial global feature sphere. For any k-th feature sphere Gk in the space, the system needs to quantitatively extract two core physical parameters describing its geometric shell features, namely the center coordinates Ck of the sphere and the radius rk of the sphere enclosure. The mathematical calculation formulas are shown in Equation (2) and Equation (3) respectively: (2) (3) Regarding the symbols in equations (2) and (3), a unique reference for the exhaustive expression is explained here: Specifically used to refer to the first segment that has been partitioned in the current feature space. Characteristic particles, where the subscripts are... A unique identification index for the particle; This indicates that the first The absolute total number of standardized data sample points currently contained within each sphere; This is a local traversal vector symbol, specifically used for traversing and referring to objects belonging to a particle. The internal first Each feature data point, index This is the index for counting local samples inside the particle; Indicates the current particle The vector summation is performed on all data points within the range. To calculate the output of this first... The geometric centroid coordinate vector of each sphere; Represents local data points to the center of mass The standard 2 norm (i.e., Euclidean distance); This means finding the maximum distance from all data points within the sphere to the centroid; this maximum distance is defined as the [missing value]. The radius of the boundary of each characteristic sphere .
[0030] After calculating the current characteristic sphere of and Subsequently, the algorithm introduces a quality evaluation index to determine whether the sphere needs further splitting. The quality evaluation index is preferably determined by a preset maximum radius tolerance threshold. It is determined by both the variance of the data density distribution within the grain and the radius of the current grain. If the data point distribution within the sphere exhibits multimodal sparse characteristics, it indicates that the sphere cannot accurately fit the local normal data boundary. In this case, the algorithm adaptively initiates a splitting mechanism, introducing the K-Means++ clustering operator, with the current centroid... Based on this, it is forcibly split into two entirely new sub-spheres in space.
[0031] After calculating the current characteristic sphere of and Then, the algorithm introduces a quality evaluation index to determine whether the sphere needs further splitting. This quality evaluation index is based on a preset maximum tolerance threshold for radius. variance of the characteristic distribution within the grain This constitutes a dual-channel decision logic. Among these, the variance of the internal characteristic distribution of the particle... The mathematical definition formula is shown in equation (4). (4) In equation (4), Quantitatively characterized the first The degree of dispersion of data points within each sphere from the center. The preset variance tolerance threshold is... During the adaptive evolution process, the algorithm performs the following logical judgment: when it is detected that the current particle satisfies the condition... or meet the conditions When this occurs, it indicates that the data distribution within the sphere is too sparse to accurately fit the local high-density data boundary. At this point, the algorithm adaptively initiates a splitting mechanism, introducing the K-Means++ clustering operator, with the current centroid... Based on this, it is forcibly split into two entirely new sub-spheres in space.
[0032] Subsequently, for the newly generated sub-spheres, equations (2) and (3) are iteratively called again to update their respective center coordinates and radii, and the quality evaluation index is determined again. This splitting evolution process is executed in parallel and alternately in the feature space of each modality until all the divided sub-spheres can perfectly and compactly wrap the local data points, and their radii are all less than or equal to the preset tolerance threshold. Through this process, the original hundreds of thousands of discrete data points are successfully compressed and simplified into a very small number (such as dozens or hundreds) of highly cohesive feature spheres, realizing adaptive matrix dimensionality reduction and efficient lightweight representation of massive heterogeneous data while completely preserving the underlying multidimensional topological structure.
[0033] Step 3: Construct a multi-head attention mechanism module to perform deep joint alignment and multi-granularity fusion of cross-modal space features. In Step 2, although adaptive particle sphere computation simplifies the massive scatter data into a finite number of feature particles, achieving a lightweight representation of a single modal feature space, different modalities (such as continuous physical sensing signals and discrete information state vectors) exist in completely heterogeneous feature spaces, making direct physical concatenation or simple linear weighting impossible. Therefore, this embodiment introduces a multi-head attention mechanism, utilizing multi-spatial projection and dynamic weight allocation to break down the barriers between heterogeneous modalities. (See attached specification.) Figure 4 The specific implementation process of cross-modal feature alignment and deep fusion is as follows: First, the feature sphere parameters generated by each mode are transformed into a feature matrix of uniform dimension. In this embodiment, the sphere feature matrix of the first mode (such as a key continuous time-series mode like bearing core temperature) is preferably used as the query matrix. The sphere feature matrices of the second mode and other heterogeneous modes (such as discrete encoding vectors of equipment operating states) are respectively used as key matrices. AND-value matrix Subsequently, in order to capture the multivariate topological correlations of cross-modal features in spatiotemporal evolution in parallel across different subspaces, parameters were introduced. This represents the total number of attention heads (where...) (As a preset positive integer). The entire attention mapping space is divided equally into Each of the independent sub-channels, for the first... Each attention point (where the subscript) For the attention head, a unique global recognition index is provided, satisfying The mathematical formulas for calculating the local feature correlation degree and attention weight within the channel are shown in Equation (5): (5) For each matrix symbol and subscript parameter in equation (5), a unique reference explanation is performed using an exhaustive formula: Specifically used to represent the first The local cross-modal fusion feature matrix output by each attention head; , , These refer to the initial query matrix, key matrix, and value matrix constructed from the original heterogeneous modal feature particles, respectively. Specifically refers to the first A unique attentional focus used for querying the matrix. The learnable parameter weight matrix for linear spatial projection; Specifically refers to the first A unique attentional point used for key matrix The learnable parameter weight matrix for linear spatial projection; Specifically refers to the first A unique attentional focus used for value matrices The learnable parameter weight matrix for linear spatial projection; This indicates that a matrix transpose operation is performed on the projected key matrix; Represents the standard matrix multiplication operator; As a global scalar, its physical meaning is uniquely defined as the size of the subspace dimension of the projected feature vector. As a scaling factor, it is used to prevent the gradient from vanishing due to excessively large dot product results; This is a row-wise normalized activation function used to transform the original geometric correlations after spatial projection of different modal particles into a local attention weight matrix conforming to a probability distribution. All operations are performed in parallel. After extracting local features from each attention head, the algorithm introduces a feature concatenation operator to globally integrate these heterogeneous local correlation information scattered across different subspaces. The output matrices of each head are horizontally concatenated along the channel dimension, and a global output weight matrix is introduced. A final linear combination mapping is performed to generate a cross-modal joint particle-sphere feature matrix that combines temporal dynamic evolution characteristics with multi-dimensional spatial topological characteristics. The combination and fusion formula is shown in equation (6): (6) In equation (6), The physical meaning is uniquely limited to the physical meaning of the object. to this Each local fusion matrix performs a matrix concatenation operation along the channel dimension; This is the global feature transformation parameter matrix; This is the final output, a multi-granularity joint feature matrix that perfectly achieves spatial cross-modal alignment. Each vector in this matrix represents a high-order joint feature particle that integrates heterogeneous information from multiple sources.
[0034] Step 4: Match nearest neighbor joint spheres and quantitatively calculate the anomaly factor score for the test sample online. Traditional anomaly detection methods typically require distance comparison between the test sample and every scatter point in the training set during the inference phase, resulting in extremely high latency in real-time online detection. In this embodiment, the massive number of scatter points has been compressed into a very small number of cross-modal joint feature spheres during the training phase. Therefore, during online inference, only efficient nearest-neighbor particle matching needs to be performed. (See attached instruction manual.) Figure 5 The specific quantitative calculation process of the abnormal factors is as follows: When an IoT system receives a brand new real-time sample for testing online, it first performs timestamp alignment and maximum / minimum value normalization on the sample using the same logic as in step one, generating the current feature vector for testing. Subsequently, the computing device will Projecting onto the cross-modal joint particle-sphere feature space constructed by equation (6), calculate The system calculates the geometric Euclidean distance to the center coordinates of each joint feature particle. Through minimum value retrieval, the system adaptively selects the target-matching joint particle closest to the online test sample and extracts the center coordinates of that target-matching joint particle. and the corresponding boundary radius Based on this, in order to quantitatively characterize the outlier degree of the test sample deviating from the normal behavior boundary, an outlier factor score (OF) without neighborhood search is constructed, and its mathematical calculation formula is shown in Equation (7): (7) For each mathematical symbol in equation (7), a unique reference explanation of the exhaustive expression is performed here: It has a unique physical meaning and is specifically used to refer to the feature vector of the online inference test sample after it has been standardized. Specifically refers to the sample to be tested within the global joint particle sphere set. The centroid coordinate vector of the normal feature joint sphere with the closest geometric distance; Specifically refers to the scalar radius of the envelope boundary of the nearest neighbor normal feature joint sphere; This indicates the calculation of the standard L2 norm (i.e., Euclidean distance) between the sample point to be tested and the coordinates of the center of the sphere. Equation (7) outputs... This is a dimensionless, quantitative score of abnormal factors. Obviously, if... This indicates that the sample point to be tested falls completely within the safe envelope boundary covered by the normal characteristic spheres, and the system has a very high degree of normality confidence; if The scores are linearly amplified, intuitively and clearly reflecting the degree of local outlier in the heterogeneous data being tested, which deviates from the normal behavioral envelope.
[0035] Step 5: Introduce hyperparameter grid search optimization to construct a multi-level intelligent early warning and modal fault tracing response mechanism under dynamic thresholds. To enable the anomaly detection system to adaptively adapt to the fluctuation differences of different industrial equipment under different operating conditions, this embodiment introduces a grid search strategy during offline calibration or model training to optimize system hyperparameters (including the maximum tolerable threshold for particle splitting radius). and total number of attention heads Joint optimization is performed. The grid search aims to maximize the detection accuracy on the validation set and automatically determines the optimal combination of hyperparameters.
[0036] After the hyperparameters are determined, the entire set of historical normal samples is input into the system, and a set of historical normal and abnormal factor scores is calculated by iterating through equation (7). The mathematical mean of this set is then calculated. with standard deviation Based on the statistical confidence interval rule, the dynamic safety threshold for online detection is rigidly set to [value missing]. .
[0037] During the system's online operation, the early warning decision module will use the real-time anomaly factor scores calculated online. With the dynamic security threshold The system performs real-time interval comparison with its preset multiples and executes the following multi-level intelligent early warning response logic: When When the system determines that the current industrial equipment is operating normally, no alarms are triggered, and online monitoring continues; when When a slight deviation in behavior of the current industrial equipment is detected, a Level 1 (yellow) warning is triggered. At this time, the system automatically increases the sampling frequency of that channel and pushes a status observation prompt to the control console. If the system determines that a serious local outlier behavior or sudden failure has occurred in the current industrial equipment, it will immediately trigger a Level 2 (red) serious fault warning.
[0038] Furthermore, upon triggering the Level 2 Red Severe Fault Warning, the system initiates an abnormal modal fault tracing procedure to provide interpretable support required for industrial maintenance. Since the attention weight matrix in equation (5) of step three accurately records the contributions of different feature dimensions and different modal particles during fusion, the tracing procedure extracts the maximum attention weight coefficient within the current time window from equation (5) in reverse, thus pinpointing the cause of the fault. The optimal feature subset with abnormally high scores is matched with the corresponding original data modes. Finally, the system synchronously outputs the critical fault alarm signal and the accurate modal fault tracing results to the industrial control screen and the terminal interface of maintenance personnel, thereby guiding maintenance personnel to perform precise and efficient shutdown maintenance and troubleshooting for the hardware equipment corresponding to specific modes.
[0039] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A cross-modal anomaly detection method based on particle-sphere computation and attention mechanism, characterized in that, Includes the following steps: S1. Acquire and preprocess cross-modal monitoring data: Read cross-modal raw data in real time from the sensor network and control system of the target industrial IoT device. The cross-modal raw data includes continuous multivariate time series data in the physical space and discrete heterogeneous state data in the information space. After time dimension alignment and sliding window slicing, a cross-modal heterogeneous feature dataset is constructed. The multivariate time series data includes bearing temperature, housing vibration acceleration, and rotor three-phase current. The heterogeneous state data includes system fault alarm logs, switch action commands, and maintenance event records. S2, Generating Modal Feature Particles: Mapping the cross-modal heterogeneous feature dataset to a multi-dimensional feature space, generating feature particles specific to each modality based on particle calculation and adaptive splitting methods, and extracting the center coordinates and radius of each feature particle; wherein step S2 specifically includes: (1) Modal Feature Space Mapping: Inputting the cross-modal heterogeneous feature dataset independently into the corresponding feature extraction network according to modal category, and mapping it to a multi-dimensional feature space; in the initialization stage, the set of all sample points under the same modality is regarded as a global feature particle; (2) Adaptive Calculation of Particle Center and Radius: For the first Characteristic spheres Its geometric center coordinates and its boundary radius The calculation formula is: in, For the first The total number of data points contained within each sphere To traverse the local index of the feature samples inside the sphere, This represents the coordinates of each sample point within the sphere from the center. (3) Granulosa splitting and joint evolution: defining the maximum radius tolerance threshold of the granule. With variance tolerance threshold When the current particle is detected radius Or the variance of the characteristic distribution within the granule When the distribution of data within the particle does not meet the high cohesion requirement, the K-Means++ clustering algorithm is used to split it into multiple sub-particles. This joint evolution and splitting process is executed iteratively until the radius and variance of all sub-particles in the current mode meet the preset threshold conditions. S3, Cross-modal attention calculation and feature fusion: Construct a multi-head attention mechanism module to calculate the spatial correlation weights between feature spheres of different modalities, and generate a cross-modal joint sphere feature matrix through weighted fusion and channel concatenation; S3 includes the following steps: Spatiotemporal feature linear mapping: Extract the center coordinates of each feature sphere generated in step S2. This serves as the representative feature vector of the particle; subsequently, the feature particle vector of the first mode is multiplied by a learnable weight matrix and mapped to a query vector. The feature particle vectors of the second mode are then multiplied by their corresponding weight matrices and mapped to key vectors. Sum value vector In this way, a cross-modal attention query mechanism is constructed. The first modality is the modality corresponding to the group evolution characteristics of the continuous physical operation state of the equipment reflected by the feature spheres of the continuous multivariate time series data generated by the target system in the physical space. The second modality refers to the discrete state data generated by the system in the information space. Its feature spheres reflect the modality corresponding to the semantic clustering characteristics of discrete events after text embedding technology. (2) Multi-head attention space association calculation: The feature mapping space is divided into A separate attention head; in the first In each attention head, the original correlation degree between cross-modal feature particles is calculated. After scaling by the square root dimension and applying Softmax activation, the local cross-modal fusion feature matrix output by the m-th attention head is obtained. The calculation formula is as follows: in, , , The first Each attention head corresponds to the query vector. Key vector Sum value vector The linear mapping parameter matrix, The feature dimension of the key vector is the size of the space. Normalized exponential activation function; (3) Cross-modal feature fusion: all The output of each attention point is obtained through a concatenation function. Perform feature concatenation and multiply by the global output weight matrix. Generate a cross-modal joint particle-sphere feature matrix that combines temporal evolution and spatial topological characteristics. The formula is: ; S4. Anomaly Factor Assessment: Calculate the Euclidean distance from the features of the online test sample to the nearest neighbor cross-modal joint particle center, and calculate the anomaly factor score based on the relative relationship between this distance and the corresponding particle radius; S5. Anomaly Warning and Status Determination: The anomaly factor score is compared with the dynamic safety threshold set based on historical normal samples to generate a graded warning result for the target industrial IoT device.
2. The cross-modal anomaly detection method based on particle-sphere computation and attention mechanism according to claim 1, characterized in that, S1 includes the following steps: (1) Collect cross-modal heterogeneous data; (2) Data cleaning and spatiotemporal alignment: Lagrange interpolation is used to fill in missing values in multivariate time series data; For discrete heterogeneous state data, feature embedding technology is used to map it into numerical feature vectors; then, based on the highest sampling rate of the continuous time series, spatiotemporal alignment logic is introduced to broadcast the discrete state data to the corresponding timestamp interval. (3) Sliding window truncation and normalization: In order to capture the temporal dynamic evolution characteristics of the data, a sliding window with a length of 1000 mm is used. Step size is A sliding time window is used to slice the time-aligned data; a maximum-minimum normalization method is used to eliminate the differences in physical dimensions between multi-source heterogeneous data. The normalization formula is: in, Indicates the first The first feature dimension One original data value, and Each represents the number of elements within the current sliding window. Minimum and maximum values in each feature dimension The features are normalized; finally, the preprocessed data is packaged into a cross-modal heterogeneous feature dataset.
3. The cross-modal anomaly detection method based on particle-sphere computation and attention mechanism according to claim 1, characterized in that, S4 includes the following steps: (1) Nearest neighbor particle matching: real-time online sample data to be tested is obtained, and after the same spatiotemporal alignment and normalization preprocessing as in step S1, the current numerical feature vector to be tested is generated. The cross-modal joint particle feature matrix output in step S3 In the corresponding set of joint particles, calculate The Euclidean distances to the centers of each joint particle are used to filter out the target matching particle with the smallest distance and extract its center. With radius ; (2) Quantitative calculation of abnormal factors: Based on the relative degree of deviation of the test sample from the target matching particle, an abnormal factor score is constructed. The calculation formula is as follows: (3) For the score of this abnormal factor ,like This indicates that the sample to be tested falls completely within the envelope boundary covered by normal characteristic spheres, possessing normal confidence level; if The ratio then increases linearly, reflecting the degree of local outliers in the joint feature space of heterogeneous data.
4. The cross-modal anomaly detection method based on particle-sphere computation and attention mechanism according to claim 3, characterized in that, S5 includes the following steps: (1) Hyperparameter optimization and threshold setting: During the model training or calibration phase, a grid search strategy is used to optimize the hyperparameters of historical normal data, calculate the set of abnormal factor scores of historical normal samples, and extract their mean. with standard deviation Based on Chebyshev's inequality or confidence interval rule, a dynamic safety threshold is set. ; (2) Multi-level early warning generation: The abnormal factor scores of online test samples are generated. With dynamic security threshold Compare with its multiples; if This triggers a Level 1 observation and warning; if This triggers a Level 2 severe fault warning and simultaneously outputs the source modality dimension that caused the abnormally high score.
5. A cross-modal anomaly detection system based on particle-sphere computation and attention mechanism, characterized in that, The system is used to implement the cross-modal anomaly detection method based on particle-sphere computation and attention mechanism as described in any one of claims 1-4, and the system includes: Data preprocessing module: used to acquire cross-modal monitoring data and perform time alignment and sliding window slicing to construct a cross-modal heterogeneous feature dataset; Particle sphere generation module: used to map the dataset to a multi-dimensional feature space, generate feature particles of each modality based on particle sphere calculation and adaptive splitting method, and extract center and radius parameters; Feature fusion module: Used to construct the multi-head attention mechanism module, calculate the spatial correlation weights between feature spheres of different modalities, and output the cross-modal joint sphere feature matrix; Anomaly assessment module: used to calculate the relative distance between the online test sample and the nearest neighbor cross-modal joint particle center, and output anomaly factor score; Early warning decision module: used to compare the abnormal factor scores with dynamic safety thresholds and generate graded abnormality early warning results.
Citation Information
Patent Citations
Image noise mark feature selection method and system, storage medium and computer
CN120894642A
Graph neural network defense method based on attribute enhanced PPR and pellet clustering
CN121683871A