Data anomaly prediction method based on multi-modal fusion
Patent Information
- Application Number
- CN202611091969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-10-09
AI Technical Summary
[0006]本发明的一个目的在于提出基于多模态融合的数据异常预测方法,针对现有技术中多源数据格式差异大、模态时间错位、轻微异常证据难以累积以及单一数据源分析难以提前识别组合风险的问题,提出了按实体对象和时间窗口归一化组织多源数据、提取多模态特征、通过时滞置信矩阵进行跨模态门控对齐、构建动态图序列和异常证据记忆图并利用图时序转换器输出异常趋势、异常分数、风险类型、贡献模态和置信区间的技术方案,本发明具备提前识别由多类轻微异常共同引发的质量、设备或业务风险并提高预测结果可追溯性的技术效果
[0056] 1. By normalizing and organizing structured business data, log text, sensor time-series data, and manually labeled information according to entity objects and time windows, and retaining missing markers, collection time, channel identifiers, and labeling sources, data from different sources can form associative multimodal sample units, providing a definite data foundation for subsequent unified feature extraction and anomalous evidence propagation.
Smart Images

Figure CN122885751A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring and anomaly prediction of industrial data, and in particular to a data anomaly prediction method based on multimodal fusion. Background Technology
[0002] In enterprise operations and industrial production scenarios, business systems, log systems, sensor acquisition systems, and manual quality inspection and annotation systems continuously generate structured business data, log text, sensor time-series data, and manually annotated information. Existing anomaly detection solutions typically model for a single data source, such as setting thresholds for business metrics, performing rule matching on log events, performing time-series prediction on sensor curves, or using manual annotations as the basis for offline verification.
[0003] The aforementioned methods struggle to handle differences in entity identification, acquisition time, sampling frequency, semantic granularity, and annotation delays across multiple data sources, making it difficult to promptly align and accumulate evidence of minor anomalies across different modalities. When quality fluctuations, equipment malfunctions, or business risks are triggered by multiple minor anomalies, the anomaly score for a single modality may not reach the alarm threshold. Existing methods are prone to issues such as delayed anomaly trend identification, insufficient explanation of risk sources, and unstable threshold calibration criteria.
[0004] Furthermore, there is often a time misalignment between log events, sensor mutations, and manual annotations, and manual annotations may also contain noise or delayed feedback. Without a unified modeling mechanism for time-delay relationships, entity relationships, process relationships, sensor channel relationships, and historical anomaly prototypes, multimodal evidence is difficult to propagate along business processes and temporal adjacency relationships, and anomaly trend prediction results are also difficult to provide traceable contribution modes and confidence intervals.
[0005] Therefore, a data anomaly prediction method is needed to address the shortcomings of existing technologies. Summary of the Invention
[0006] One objective of this invention is to propose a data anomaly prediction method based on multimodal fusion. Addressing the problems in existing technologies such as large differences in multi-source data formats, temporal misalignment of modalities, difficulty in accumulating evidence of minor anomalies, and the inability of single-source data analysis to identify combined risks in advance, this invention proposes a technical solution that normalizes multi-source data by entity object and time window, extracts multimodal features, performs cross-modal gating alignment using a time-delay confidence matrix, constructs dynamic graph sequences and anomaly evidence memory graphs, and utilizes a graph time-series converter to output anomaly trends, anomaly scores, risk types, contributing modalities, and confidence intervals. This invention has the technical effect of identifying quality, equipment, or business risks caused by multiple minor anomalies in advance and improving the traceability of prediction results.
[0007] This invention provides a data anomaly prediction method based on multimodal fusion, comprising: S1, normalizing and organizing structured business data, log text, sensor time-series data, and manually labeled information according to entity objects and time windows to generate multimodal sample units; S2, extracting business state change features, log semantic event features, sensor multi-scale time-frequency features, and labeled prototype features from the multimodal sample units respectively to generate a multimodal feature sequence; S3, performing cross-modal comparison mapping on the multimodal feature sequence, and generating a time-delay confidence matrix indexed by modality pairs and offset windows according to candidate offset windows, mapping each modality feature to the same S4. Based on the aligned evidence sequence, construct a dynamic graph sequence containing entity nodes, process nodes, log event nodes, sensor channel nodes, and labeled prototype nodes. The historical anomaly prototype nodes, labeled prototype nodes, and the current window node form an anomaly evidence memory graph, and output a graph sequence with memory edge weights. S5. Input the graph sequence with memory edge weights into the trained graph time-series converter for cross-node aggregation and cross-time propagation to obtain a cumulative anomaly vector. Output the anomaly trend, anomaly score, risk type, contribution mode, and confidence interval for future time windows through a trend prediction head.
[0008] Optionally, S1 includes:
[0009] Configure unified entity identifiers, data source identifiers, and collection timestamps for structured business data, log text, sensor time-series data, and manually labeled information;
[0010] Divide the time window according to the preset window length and preset sliding step size;
[0011] Data belonging to the same entity and whose collection timestamps fall within the same time window are written into the same multimodal sample unit;
[0012] Write the missing marker and missing duration for the missing modality;
[0013] Numerical fields are standardized according to the mean and standard deviation of the training sample distribution; category fields are generated with category codes; event time and event template identifiers are retained in log text; channel identifiers, sampling frequency and sampling sequence are retained in sensor time series data; and annotation category, annotation time and annotation source are retained in manually annotated information.
[0014] Optionally, S2 includes:
[0015] Business state change characteristics are generated based on the field differences, state transition counts, and process node dwell times in the continuous time window of structured business data.
[0016] Perform event template matching and context encoding on log text to generate log semantic event features;
[0017] Statistical quantities, frequency band energy, and abrupt change point locations are calculated for sensor time-series data according to a preset scale window to generate multi-scale time-frequency characteristics of the sensor.
[0018] Aggregate the features of samples with the same annotation category from manually annotated information to generate annotation prototype features;
[0019] The business status change features, log semantic event features, sensor multi-scale time-frequency features, and labeled prototype features are sorted by entity object and time window to obtain the multimodal feature sequence.
[0020] Optionally, S3 includes:
[0021] Each modal feature in the multimodal feature sequence is input into the corresponding modal encoder to obtain a modal embedding vector of the same dimension;
[0022] The modal encoder is updated by comparing the loss, using different modal embedding vectors of the same entity object and within the same time window as positive sample pairs, and different modal embedding vectors of the entity object or time window as negative sample pairs.
[0023] The alignment confidence of each modality pair under each candidate offset window is calculated based on the absolute value of the time difference between the log event time, the sensor mutation point time, and the manual annotation time, respectively, and the candidate offset window.
[0024] The time-delay confidence matrix is generated by performing a normalized exponential transformation on the alignment confidence scores of the same mode pair.
[0025] The aligned evidence sequence is obtained by gating and weighting the modal embedding vectors under different candidate offset windows according to the time-delay confidence matrix.
[0026] Optionally, S4 includes:
[0027] Entity relationships, business process adjacency relationships, sensor channel association relationships, and time adjacency relationships are respectively converted into entity edges, process edges, channel edges, and time edges in a dynamic graph sequence;
[0028] Add the historical anomaly prototype, manually annotated prototype, current entity node, log event node, and sensor channel node to the anomaly evidence memory graph;
[0029] The memory edge weights are determined based on the alignment confidence in the time-delay confidence matrix, the prototype similarity between the current window node and the historical abnormal prototype node, the business process adjacency marker, and the time sequence marker.
[0030] Anomalies whose single-modal anomaly scores are less than a preset alarm threshold and not less than the p-th percentile of the historical background distribution of that modality are written into the anomaly evidence memory graph, where p is a preset percentile parameter.
[0031] Write the memory edge weights into the corresponding edge attributes to obtain the graph sequence with memory edge weights.
[0032] Optionally, S5 includes:
[0033] In the graph sequence with memory edge weights, the anomaly evidence vectors of adjacent nodes are weighted and aggregated according to the attributes of each edge to generate a node aggregation vector.
[0034] The node aggregation vectors of the same entity object are temporally encoded according to the time window order to generate the entity temporal representation;
[0035] The entity temporal representation is fused with the representation of the corresponding historical anomaly prototype node in the anomaly evidence memory graph to obtain the cumulative anomaly vector;
[0036] The cumulative anomaly vector is input into the trend prediction head, which outputs the anomaly trend category, anomaly score, risk type label, contribution weight of each modality, and upper and lower bounds of the confidence interval for the future time window.
[0037] Furthermore, the trained graph time-series converter and the trend prediction head are obtained by using historical multimodal samples with future time window anomaly labels, risk type labels, and anomaly occurrence times as training samples.
[0038] The training samples are processed through steps S1 to S4 to obtain a training graph sequence;
[0039] Input the training graph sequence into the graph time series converter and the trend prediction head to obtain the predicted anomaly score, the predicted risk type, and the predicted anomaly occurrence time.
[0040] The model parameters are updated by weighting and summing the classification loss between the predicted anomaly score and the anomaly label, the classification loss between the predicted risk type and the risk type label, the time deviation loss between the predicted anomaly occurrence time and the labeled anomaly occurrence time, and the comparison loss, until the change in the weighted sum on the validation sample set within a preset number of rounds is less than the convergence threshold.
[0041] Furthermore, the confidence interval is determined by obtaining the true anomaly score for the future time window and the anomaly score output by the trend prediction head from the calibration sample set;
[0042] Calculate the absolute value of the difference between the two as the non-compliant score;
[0043] The quantile values are determined from the distribution of the non-compliant scores according to the target coverage.
[0044] Subtract the quantile value from the outlier score output by the trend prediction head to obtain the lower bound of the confidence interval, and add the quantile value to the outlier score output by the trend prediction head to obtain the upper bound of the confidence interval.
[0045] When the time range of the calibration sample set is updated, the quantile values are re-determined based on the updated non-consistent fraction distribution;
[0046] Furthermore, the contribution weights of each modality are determined as follows: the edge attention weights of the graph time-series converter during cross-node aggregation and the temporal attention weights during cross-time propagation are recorded;
[0047] The attention weights of nodes, edges, and time windows corresponding to the same modality source are summed and normalized to obtain the initial value of the modality contribution;
[0048] The initial value of the modality contribution is corrected based on the missing duration written in step S1 and the alignment confidence generated in step S3 to obtain the contribution weight of each modality.
[0049] The contributing mode is determined by ranking the modes that rank first according to their contribution weight values.
[0050] Furthermore, the abnormal scores are used for threshold calibration, specifically including: statistically analyzing the distribution of abnormal scores and confirmed abnormal labels for historical time windows within each update cycle;
[0051] Candidate alarm thresholds are generated based on the qth quantile value of the abnormal score distribution, where q is a preset percentile parameter.
[0052] Calculate the false alarm rate and missed alarm rate corresponding to the candidate alarm thresholds based on the confirmed abnormal labels;
[0053] When the false alarm rate is not greater than the preset false alarm rate threshold and the false alarm rate is not greater than the preset false alarm rate threshold, the candidate alarm threshold is written into the preset alarm threshold of the next update cycle.
[0054] When the false alarm rate or the false alarm rate exceeds the corresponding threshold, the current preset alarm threshold is maintained and the samples of the update cycle are written into the calibration sample set.
[0055] The beneficial effects of this invention are:
[0056] 1. By normalizing and organizing structured business data, log text, sensor time-series data, and manually labeled information according to entity objects and time windows, and retaining missing markers, collection time, channel identifiers, and labeling sources, data from different sources can form associative multimodal sample units, providing a definite data foundation for subsequent unified feature extraction and anomalous evidence propagation.
[0057] 2. By generating a time-delay confidence matrix based on log event time, sensor mutation time, and manual annotation time during cross-modal alignment, and by gating and weighting the modal embedding vectors under the candidate offset window, modal features with sampling frequency differences, event delays, or annotation lags can be mapped to a unified anomaly evidence space, thereby reducing the impact of time misalignment on anomaly trend prediction.
[0058] 3. By constructing a dynamic graph sequence containing entity nodes, process nodes, log event nodes, sensor channel nodes, and labeled prototype nodes, and introducing an anomaly evidence memory graph composed of historical anomaly prototypes, manually labeled prototypes, and current window nodes, anomaly evidence that does not reach the alarm threshold but meets the historical background quantile condition can be accumulated and propagated, thereby supporting the joint output of future time window anomaly trends, risk types, contribution modes, and confidence intervals. Attached Figure Description
[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0060] Figure 1 This is a flowchart of a data anomaly prediction method based on multimodal fusion.
[0061] Figure 2 This is a flowchart of step S3 of the present invention, which involves traversing candidate offsets, calculating the confidence matrix, and outputting aligned evidence using gated weighting. Detailed Implementation
[0062] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0063] refer to Figures 1-2A data anomaly prediction method based on multimodal fusion includes: S1, normalizing and organizing structured business data, log text, sensor time-series data, and manually labeled information according to entity objects and time windows to generate multimodal sample units; S2, extracting business state change features, log semantic event features, sensor multi-scale time-frequency features, and labeled prototype features from the multimodal sample units to generate multimodal feature sequences; S3, performing cross-modal comparison mapping on the multimodal feature sequences, and generating time-delay confidence matrices indexed by modality pairs and offset windows according to candidate offset windows, mapping each modality feature to the same dimension. S4. Based on the aligned evidence sequence, construct a dynamic graph sequence containing entity nodes, process nodes, log event nodes, sensor channel nodes, and labeled prototype nodes. The historical abnormal prototype nodes, labeled prototype nodes, and current window nodes form an abnormal evidence memory graph, and output a graph sequence with memory edge weights. S5. Input the graph sequence with memory edge weights into the trained graph time-series converter for cross-node aggregation and cross-time propagation to obtain a cumulative abnormal vector. Output the abnormal trend, abnormal score, risk type, contribution mode, and confidence interval of the future time window through the trend prediction head.
[0064] In this specific embodiment, S1 includes:
[0065] Structured business data, log text, sensor time-series data, and manually labeled information are normalized and organized according to entity objects and time windows to generate multimodal sample units. Entity identifiers, data source identifiers, and collection timestamps are uniformly configured for all four data types. The entity identifier uses a unified entity identifier and is denoted as [insert identifier here]. The entity mapping table maps any original identifier from the business system primary key, equipment asset number, and work order number to the same entity. The data source identifier is enumerated and written into the src field, where the structured business data corresponds to... Log text correspondence Sensor time series data correspondence Corresponding to manually labeled information The collected timestamps are uniformly set to millisecond timestamps and written into the ts field, and time zone conversion is performed using UTC time base to ensure cross-system time consistency.
[0066] The time window uses a preset window length. With preset sliding step size To fix the reference time As the starting point for window division, the time interval (ts) of any record will fall into the interval. Records are written to the same time window, among which According to Start with step size Incrementally obtained, and the multimodal sample unit is defined as the key. Data Structures This data structure contains a window start and end time field, win_start. With win_end Each structure is configured with four substructures: structured business data payload (biz), log event payload (log), sensor channel payload (sens), and manual annotation payload (anno).
[0067] For missing modalities, write the missing flag and the duration of the missing data, where the missing flag is written to the field `miss_flag`. If and only if the corresponding mode is in When there is no record of the flag "miss_flag" And write the missing duration field miss_dur When a record exists for the corresponding modality, set the miss_flag. And write miss_dur This ensures that subsequent steps can consistently identify missing states and missing spans;
[0068] Standardize the numeric fields in the structured business data and write them back to the Biz substructure. Calculate the mean for each numeric field in the training sample set beforehand. with standard deviation This is then solidified into a standardized parameter table, and the original value of this field is used during online normalization. Calculate the standardized results And replace the original value, the standardization satisfies the formula ,in This represents the value of the standardized numeric field. This indicates the original value of the field in the current record. This represents the mean of the field on the training sample set. This indicates that the standard deviation of this field on the training sample set is greater than 0;
[0069] Generate category codes for the category fields in the structured business data and write them back to the Biz substructure. The category codes are constructed offline from a category vocabulary based on the training sample set. The encoding is given as an integer, and unregistered categories are uniformly mapped to the encoding 0 to maintain a stable encoding space.
[0070] The log text retains the event time and event template identifier and writes them. The substructure contains the event time (ts) of the log line and the event template identifier (tpl_id) field, which is defined by a fixed log template library. Given that each template in the library has a unique integer number, the matching of log lines to templates is implemented using a template parser based on a Drain structure, and its parsing parameters are fixed as tree depth. Similarity threshold Maximum number of branches The parser outputs a unique tpl_id for each log line and writes it along with the event time to the list field log.events to support aggregation of multiple events in the same window;
[0071] The sensor time-series data retains the channel identifier, sampling frequency, and sampling sequence and writes them into the sens substructure, where the channel identifier is written into the field. The sampling frequency is obtained by mapping from the sensor point table and written into the field fs in Hertz, and the sampling sequence is written into the field seq and determined by all ts falling into the range. The sampled values are concatenated in ascending order of time. If there are multiple channels, multiple (ch_id, fs, seq) triples are stored in the sens.channels list.
[0072] Manually annotated information retains the annotation category, annotation time, and annotation source, and writes them into the `anno` substructure. The annotation category is written into the `label_id` field and is determined by a fixed annotation category table. The mapping results in the labeling time being taken from the ts of the labeling record, and the labeling source being written into the label_src field and enumerated encoding is used to distinguish between quality inspectors, automatic backfilling and external audit sources. If there are multiple labels in the same window, they are written into the anno.labels list in ascending order of time.
[0073] After the window ends, The data are serialized into multimodal sample units required for training or inference and output to the downstream step S2 to ensure that subsequent feature extraction has a deterministic, reproducible and traceable data organization basis in both the entity object and time window dimensions.
[0074] In this specific embodiment, S2 includes:
[0075] Based on multimodal sample units Extract business status change features, log semantic event features, sensor multi-scale time-frequency features, and labeled prototype features to form a multimodal feature sequence, where each entity is identified as... And the starting point of the time window is of First, process the structured business data payload (biz) to generate a feature vector of business state changes. The Biz pre-configures field settings table and distinguishes between sets of numeric fields. Category field collection The process node field `node_id` and the status field `state_id` are used to calculate the field differences for consecutive time windows after sorting the `biz.records` within the current window in ascending order by `ts`, and then written to the records. For each numerical field, the current window statistic includes the window's last value, the window mean, and the value relative to the previous window. The differences of the corresponding statistics are concatenated in field order. For each category field, the number of times `state_id` changes within the current window is calculated as the state transition count and concatenated. The dwell time of each process node is counted using `node_id` as the key and written into a fixed-length vector in dictionary order of process nodes. The dwell time is obtained by accumulating the TS difference between two adjacent `biz.records`, and only the part that falls within the window is counted for the part that crosses the window. Duration;
[0076] Processing log event payloads To generate log semantic event feature vectors The event template identifiers tpl_id in log.events are arranged in ascending order of event time, and the length does not exceed [a certain value]. The template sequence is truncated to the latest 64 events if it exceeds the limit, and the missing part is filled with template identifier tpl_id. The template sequence is padded and a valid length mask is generated simultaneously. The template sequence is then mapped to a vector sequence through an embedding layer with the embedding dimension set to 1. The embedding layer parameter is size A trainable lookup table matrix is used, and vectors are retrieved using tpl_id as the index. The embedded vector sequence is then added to a sinusoidal positional code based on event-relative time and input into a two-layer Transformer context encoder to obtain a context-sensitive representation. Each layer of the Transformer contains a multi-head self-attention network and a feedforward network, with the number of multi-heads set to [value missing]. , feedforward hidden dimension set to Layer normalization employs a Pre-LN structure, and Dropout with a dropout rate of 0.1 is applied after the attention and feedforward sub-layers. The encoded output is obtained by mask mean pooling of the effective length mask. And maintain the dimension as ;
[0077] Process the sensor channel load sens to generate multi-scale time-frequency feature vectors of the sensor. Sort sens.channels in ascending order by channel identifier ch_id and fix the maximum number of channels to be processed. The sampled sequence for each channel is divided into three scale windows. Any excess channels are truncated to the first 16 channels, and any insufficient channels are padded with zero-channel features. The upper segment calculates statistics and frequency band energy, and extracts the locations of abrupt change points. The statistics include mean, standard deviation, minimum, maximum, and root mean square, concatenated in scale order. Frequency band energy is calculated by performing a real-valued Fast Fourier Transform on each signal segment and extracting the low-frequency band. Mid-frequency band High frequency band The power spectral density is accumulated to obtain a three-dimensional energy vector, which is then stitched together in scale order. The location of the mutation point is obtained by CUSUM mutation detection, and the time position of the mutation point relative to the start of the window is divided by the window length. Normalization to Interval Union and First Option Each mutation point location is written into the feature vector; if there are fewer than three, they are padded with -1. Finally, the features from each channel are concatenated in channel order to obtain the final result. ;
[0078] To process manually labeled payloads (anno) to generate labeled prototype features, first... The features of the current window samples are obtained by concatenating them in a fixed order. This serves as input for prototype aggregation, and the prototype table is maintained using the label category `label_id` of each label in `anno.labels` as the key. For each labeled category The prototype vector is updated according to the exponential moving average rule and the update coefficient is fixed. Its update formula is:
[0079] ;
[0080] in Indicates the label category is The The updated annotation prototype vector. Indicates the label category is The The updated annotation prototype vector. This represents the exponential moving average coefficient for prototype updates. The entity identifier is And the starting point of the time window is The current window sample feature vector, This indicates the category code corresponding to the labeled category. This indicates the cumulative number of updates for the same label category. If there are no annotations within the window, the prototype update will not be triggered, and the annotation prototype feature vector corresponding to the entity in that window will be updated. Set to an all-zero vector; if a label exists, then... The category corresponding to the last label in this window The current prototype Write as This ensures that each window corresponds to only one specific labeled prototype feature;
[0081] Finally Together with the missing flag written in step S1, according to entity identifier Starting point of the time window Sort the data to form a multimodal feature sequence of the entity and output it to step S3.
[0082] In this specific embodiment, S3 includes:
[0083] Perform cross-modal contrastive mapping and time-delay gated alignment on the multimodal feature sequences to obtain aligned evidence sequences, where the modality set is defined as:
[0084] ;
[0085] Configure a mode encoder for each mode The modal feature is mapped to a modal embedding vector of the same dimension, which is also used as an anomaly evidence vector and denoted as . ,in Indicates modal identifier, Indicates a unified entity identifier. Indicates the start point of the time window. Indicates the time window sequence number. Represents a unified embedding dimension; a structured business modal encoder. A two-layer fully connected network is used with a fixed hidden dimension of 256, the activation function is ReLU, and LayerNorm is applied to the output. The input is the business state change feature. and output ;
[0086] Log Modal Encoder Using a linear projection layer from Projected to Apply LayerNorm to the output to output ;
[0087] Sensor Modal Encoder A two-layer fully connected network with a fixed hidden dimension of 512 and GELU activation function is used, with LayerNorm applied to the output. The input is multi-scale time-frequency features from the sensor. and output ;
[0088] Annotation prototype modal encoder A two-layer fully connected network with a fixed hidden dimension of 256, ReLU activation function, and LayerNorm applied to the output, is used. The input consists of labeled prototype features. and output ;
[0089] The four modal encoders mentioned above are updated through cross-modal contrastive learning during the offline training phase, with training batches consisting of multiple... Composed of and each Within the context, any two modality embedding vectors of different modalities are selected as positive sample pairs, while modality embedding vectors with different entity identifiers or different time window indices are selected as negative sample pairs. Cosine similarity is used for comparison, and the comparison temperature parameter is fixed. At the same time, when S1 writes the corresponding mode missing flag (miss_flag) The modality is not included in the construction of positive sample pairs, and its modality embedding vector is fixed as an all-zero vector to ensure consistency in missing data handling between training and inference;
[0090] After completing the comparison mapping, time-delay alignment is performed, and the candidate offset window set is defined as follows: and and for each entity With window Determine the representative event times within each modality. The representative event time of the business modality is taken from the center of the window. and The representative event time of the log modality is taken as log. The median of the event times in the events list, and when there are no events in the window, is taken as... The representative event time of the sensor modality is taken as the absolute time corresponding to the first mutation point obtained by CUSUM mutation detection in the sensor, and when there is no mutation point in the window, it is taken as... The representative event time of the prototype modality is taken as the time of the last label in anno.labels, and when there are no labels in the window, the time is taken as the time of the last label. ;
[0091] Based on the above representative event times, each pair of modes With each candidate offset Calculate alignment confidence and generate a time-delay confidence matrix indexed by mode pairs and offset windows. Its calculation uses a normalized exponential transformation and satisfies the formula:
[0092] ;
[0093] in The entity identifier is And the time window number is Time modality With mode In offset Alignment confidence under the following conditions This indicates the modal identifier used as an alignment reference. Indicates the modal identifier being aligned. and This represents the time offset of the candidate offset window, in seconds. Representing modes Representative event time, Representing modes Representative event time, This represents the absolute value operation. Represents an exponential function. Represents temperature parameters on a time scale. Indicates all candidate offsets The normalization term for summation;
[0094] Subsequently, the cross-window modal embedding vectors of the aligned modes are gated and weighted according to the time-delay confidence matrix to generate an alignment evidence vector, where the offset is... Mapped to window index offset and and from the same entity window Take the aligned modality embedding vector If participating in weighted average, Beyond the Entity The available window range will then Set it to an all-zero vector and treat the term as missing to ensure that the aligned output dimension is fixed;
[0095] For each entity With window embedding vectors with business modalities As an alignment reference, the above-mentioned gating weighting is performed on the log modality, sensor modality, and labeled prototype modality respectively to obtain the following results. , , and will By time window number The ascending order is used to align the evidence sequence for output to step S4.
[0096] In this specific embodiment, S4 includes:
[0097] Receive aligned evidence sequences, and in each time window A dynamic graph snapshot containing entity nodes, process nodes, log event nodes, sensor channel nodes, and labeled prototype nodes is constructed to form a dynamic graph sequence. Simultaneously, an anomaly evidence memory graph consisting of historical anomaly prototype nodes, labeled prototype nodes, and the current window node is constructed, and the memory edge weights are written into the graph edge attributes. The dynamic graph snapshot is defined by window index. A heterogeneous graph data structure is used to record node types, edge types, node characteristics, and edge attributes, with entity nodes identified by a unified entity identifier. Create and denote the index as Its node features are written into the business modality anomaly evidence vector. This enables entity nodes to carry evidence related to changes in the business status of the window;
[0098] Process nodes are identified by business process nodes. Create and denote the index as The business process node identifier The structured business data field node_id from step S1 is given and taken as follows: The node_id with the longest dwell time is used as the flow position of the window entity, and the flow node features are obtained using a trainable flow embedding vector table. By index Dimensions obtained by table lookup and embedded ;
[0099] Log event nodes are recreated and recorded as unique events within each window based on the event template identifier tpl_id. ,in This indicates the first [number] in this window. Each unique tpl_id represents a log event node feature derived from a window-level aligned log anomaly evidence vector. With template embedding vector The template embedding vector is obtained by element-wise addition and its size is... The trainable lookup table matrix is given, and the template is written into the node attributes. The number of occurrences within the specified range is used for subsequent aggregation;
[0100] Sensor channel nodes are identified by channel identifier Create and record as ,in The channel identifier is derived from sens.channels in step S1, and the channel node features are window-aligned with the sensor anomaly evidence vector. With channel embedding vector The channel embedding vector is obtained by element-wise addition and is determined by its size. The trainable lookup table matrix is given and compared with the channel upper limit in step S2. Consistent;
[0101] Annotation prototype nodes are coded according to annotation category Create and record as The node features are written into the annotation prototype vector maintained in step S2. And when the prototype has not yet been initialized Set it to an all-zero vector to ensure a fixed dimension;
[0102] Based on the above nodes, create a current window node for each entity and window, and denote it as follows: Its node characteristics are composed of and The cumulative state of multimodal anomaly evidence in this window is obtained by adding elements together;
[0103] Entity relationships, business process adjacency relationships, sensor channel association relationships, and temporal adjacency relationships are converted into entity edges, process edges, channel edges, and time edges, respectively. Entity edges connect different entity nodes under the same production line identifier and are determined by the master data relationship table. Process edges connect process nodes that are adjacent in terms of process routing and are determined by the process adjacency matrix. Channel edges connect channel nodes with a Pearson correlation coefficient of not less than 0.7 within the offline statistical period and are determined by the channel association matrix. Temporal edges connect entity nodes of the same entity in adjacent windows and form... Directed edges can be used to explicitly express temporal relationships;
[0104] The abnormal evidence memory graph is defined as a heterogeneous graph structure maintained across windows and denoted as . Historical anomaly prototype nodes are tagged by risk type. Create and record as Its node features are historical anomaly prototype vectors. The historical anomaly prototype vector occurs when a confirmed anomaly occurs and the risk type is... Use the current window node The node features are updated with an exponential moving average coefficient of 0.02 and remain unchanged when there are no confirmed anomalies, and the size of the risk type label set is fixed at 8 to ensure that the number of prototype nodes is determined.
[0105] Based on the time delay confidence matrix Generate the alignment confidence scalar for the current window ,in Take the modal pairs (biz, log), (biz, sens), and (biz, proto) in the candidate offset window set. The arithmetic mean of the maximum alignment confidence over the window is used to reflect the multimodal temporal alignment reliability.
[0106] In and join in Then, memory edges are created for the current window node, each historical anomaly prototype node, and each labeled prototype node, and the weights of these memory edges are written into the edge attribute field. The memory edge weights are determined by alignment confidence, prototype similarity, business process adjacency markers, and time sequence markers. Prototype similarity is calculated using the cosine similarity between the current window node features and the prototype node features, and denoted as... The business process adjacency marker is used to indicate whether the current entity's process node and the prototype's associated process node are adjacent in the process adjacency matrix, and is denoted as . The time sequence marker is used to indicate whether the time of the exception corresponding to the prototype node is earlier than the start time of the current window. And record it as The memory edge weights are calculated using the following formula:
[0107] ;
[0108] in Represents a node With nodes Remember the edge weights and write them into the edge attributes. Indicates the current window node Represents historical anomaly prototype nodes Or mark the prototype node The Sigmoid function maps real numbers within parentheses to the range 0 to 1. This represents the weight of the alignment confidence term and takes a value of [value]. This represents the weight of the prototype similarity term and takes the value of . This indicates the weight of the process adjacency marker item and has a value of 1.0. This represents the weight of the time sequence marker and has a value of 0.5. The entity identifier is And the time window number is Alignment confidence scalar Represents a node With nodes cosine similarity, Represents a node With nodes Binary markers indicating whether processes are adjacent. Represents a node Is it earlier than the current window's binary tag? Indicates a unified entity identifier. Indicates the time window sequence number;
[0109] To achieve the goal of "writing anomalous evidence that does not reach the alarm threshold but meets the historical background quantile condition into the anomalous evidence memory map in a single modality", this step involves each modality... Configure a single-modal anomaly scoring network And output the single-mode anomaly score The single-modal anomaly scoring network adopts a two-layer perceptron structure and has an input dimension of . The hidden layer width is 64, the activation function is GELU, and the output layer uses Sigmoid to ensure score normalization. The business modality input is... Log modal input is Sensor modal input is The prototype modal input is labeled as The preset alarm threshold is set to 0.9, and a percentile parameter is set. For each modality, calculate the first modality anomaly score on the single-modal anomaly score distribution of historical normal samples. The quantile value is used as the historical background distribution threshold for this mode and is denoted as... If and only if there exists a mode that satisfies and At that time, the current window node Mark nodes as retrievable evidence nodes and in the abnormal evidence memory graph Twenty consecutive time windows are reserved for cross-window accumulation and propagation, while the memory edges between this node and each prototype node are configured as described above. Write edge attributes And simultaneously write the current dynamic graph snapshot. This forms a graph structure with memorized edge weights, which is then output in window order. The graph sequence with memory edge weights is passed to S5.
[0110] In this specific embodiment, S5 includes:
[0111] Graph sequences with memory edge weights Input a pre-trained graph time-series transformer with fixed parameters to achieve cross-node aggregation and cross-time propagation, and output the anomaly trend, anomaly score, risk type, contribution mode, and confidence interval for future time windows, where each graph snapshot... Includes node feature dimensions Heterogeneous nodes and those containing edge attributes The heterogeneous edges, the The memory edge weights written in step S4 have values ranging from 0 to 1;
[0112] Cross-node aggregation in each Internally, perform attention-based neighborhood aggregation on each node and assign edge attributes. Used as an attention gating coefficient, during aggregation, edge type embeddings are configured for different edge types and jointly generated with the features of adjacent nodes to produce attention scores, thereby enabling entity edges, process edges, channel edges, time edges, and memory edges to participate in information propagation under the same attention framework; the graph temporal transformer is constructed by concatenating two layers of graph attention coding layers and two layers of temporal Transformer coding layers, with the number of attention heads in the graph attention coding layer set to... Each head dimension is The feedforward hidden dimension is 512, the dropout rate is 0.1, and the output of each layer is processed by LayerNorm to stabilize the training distribution. The temporal Transformer coding layer uses causal masks and is ordered according to time window numbers. Only allow information to propagate from past windows to the current window, and the number of attention heads is also set to [value missing]. The feedforward hidden dimension is 512, the dropout rate is 0.1, and the output of each layer is processed by LayerNorm;
[0113] After completing each After cross-node aggregation, from Read entity nodes The aggregation representation serves as the node aggregation vector for that entity within that window, and the same entity is identified according to the time window order. The node aggregation vectors form the entity temporal input sequence, which is then fed into the temporal Transformer to obtain the entity temporal representation;
[0114] The entity temporal representation is compared with the abnormal evidence memory map. The historical anomaly prototype nodes in the data are fused to obtain the cumulative anomaly vector, where the historical anomaly prototype nodes are indexed by risk type label and their prototype vectors are denoted as . During fusion, for all risk type labels corresponding to The similarity between the entity temporal representation and the entity temporal representation is calculated, normalized to weights, and then weighted and summed to obtain the memory context vector. The memory context vector and the entity temporal representation are then concatenated by channel and projected onto a fully connected layer. The accumulated anomaly vector is obtained; the accumulated anomaly vector is input into a trend prediction head for multi-task output. The trend prediction head is constructed using a shared two-layer perceptron trunk and three output branches. The hidden dimension of the shared trunk is set to 128, the activation function is GELU, and a dropout rate of 0.1 is applied. The first output branch is for future... The system identifies three abnormal trend categories within a given time window and uses Softmax output. The second output shows the future trend. The outlier score is calculated for each time window and a sigmoid output is used to ensure the outlier score falls between 0 and 1. The third output shows the future... Each time window has a risk type label, the number of risk types is set to 8, and Softmax is used for output;
[0115] The contribution weights of each modality are generated by the internal attention weights of the graph temporal transformer and obtained through missing and alignment corrections. During cross-node aggregation, attention weights related to edges are recorded, and during cross-temporal propagation, attention weights related to time steps are recorded. These attention weights are then mapped back to the modality set according to node type and edge type. After summing the initial values of each modal contribution and normalizing them, the contribution of the missing mode is attenuated based on the missing duration miss_dur written in S1, and then based on the time delay confidence matrix obtained in S3. The modal contribution with low alignment confidence is attenuated, and then normalized again to obtain the contribution weight of each modality. The modality with the largest contribution weight is determined as the contributing mode.
[0116] The confidence interval is determined using the distribution of non-conformity scores in the calibration sample set and is updated as the time range of the calibration sample set is updated. Specifically, the non-conformity score is calculated by taking the outlier score output from the trend prediction head in the calibration sample set and calculating the absolute value of the difference between it and the true outlier score in the corresponding future time window. This difference is then used to determine the target coverage. Determine the quantile value from the non-consistent fractional distribution and denot it as . And for each predicted outlier score Output the lower bound of the confidence interval and the upper bound of the confidence interval And satisfy ,in The entity identifier is And the time window number is The lower bound of the confidence interval for the outlier score output at that time. This represents the upper bound of the confidence interval for the corresponding outlier score. This represents the outlier score output by the trend prediction head. Indicates target coverage rate The quantile value determined by the non-consistent fractional distribution of the calibration sample set. This represents the target coverage parameter and is fixed at 0.9.
[0117] In this specific embodiment, the graph time series converter and the trend prediction head are obtained through offline supervised training, and the parameters are fixed after training and used for the inference output of S5;
[0118] The training data consists of historical multimodal samples and is labeled with a uniform entity identifier. Starting point of the time window The organization, each training sample simultaneously carries the future... The anomaly label, risk type label, and anomaly occurrence time are defined for each time window. The anomaly label indicates whether an anomaly will occur in a future time window and is denoted as [missing information]. The risk type label is used to indicate the risk type corresponding to the future time window and is denoted as . The time of the anomaly is the actual timestamp of the first occurrence of the anomaly in the future, and is denoted as . ;
[0119] In each training iteration, steps S1 to S4 are executed sequentially for all samples within the batch to obtain a training graph sequence with memory edge weights. ,Will The input graph timing converter obtains the cumulative anomaly vector, and the input trend prediction head obtains the predicted anomaly score. Predicting risk types With the predicted time of anomaly occurrence ,in A continuous value between 0 and 1 For the probability distribution of 8 risk types, The time offset prediction is in seconds and relative to the current window start point. Perform regression;
[0120] The training objective function employs a weighted sum of classification loss, risk type classification loss, time bias loss, and contrastive loss, and is applied to the graph time series converter, trend prediction head, and modal encoder in S3. Joint update, classification loss uses binary cross-entropy and... Supervision Risk type classification loss uses cross-entropy and is based on Supervision Time deviation loss is adopted The absolute error is used to constrain the anomaly occurrence time for localization. The contrast loss follows the cross-modal contrastive learning settings of step S3 and uses temperature parameters. To bring cross-modal embeddings of the same entity and the same window closer together, and cross-modal embeddings of different entities or different windows further apart;
[0121] The total loss function is denoted as And satisfy ,in This represents the total loss used for backpropagation. This represents the classification loss weight for the anomaly label and has a value of 1.0. Indicates the predicted outlier score With abnormal tags Binary cross-entropy loss between them This represents the risk type classification loss weight and its value is... Indicates the type of predicted risk Risk type label Cross-entropy loss between This represents the time deviation loss weight and has a value of 0.5. Indicates the predicted time of anomaly occurrence. Compared with the actual time of occurrence of the anomaly The absolute error loss between them This indicates the comparison loss weights and their values are... Indicates cross-modal contrast loss;
[0122] The optimizer uses AdamW and has a fixed learning rate. Weight decay is A sequence of 64 consecutive time windows with a batch size of 32 and each batch containing 32 entities was used, and the gradient norm was pruned to 1.0 to suppress gradient explosion.
[0123] The training process involves dividing the samples into a training set and a validation set in chronological order, and monitoring is performed on the validation set. The maximum number of training rounds is set to 50, and the convergence criterion is the validation set. The change over five consecutive rounds was less than the convergence threshold. Training stops and the validation set is saved when the convergence criterion is met. The model parameters corresponding to the minimum number of rounds are used as the trained graph time series converter and trend prediction head.
[0124] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0125] This invention connects multi-source data normalization organization, cross-modal time delay alignment, dynamic graph sequence modeling, and graph time-series converter prediction into a continuous technical processing chain. This enables structured business status changes, log semantic events, sensor time-frequency changes, and manually labeled prototypes to participate in the calculation in the same abnormal evidence space, thereby forming early predictions for quality fluctuations, equipment status anomalies, or business risks caused by multiple minor anomalies.
[0126] This invention sets up a time-delay confidence matrix and a gated weighting mechanism in the cross-modal alignment module, and introduces an anomaly evidence memory graph before the graph time-series converter. The memory edge weights are updated by alignment confidence, prototype similarity, business process adjacency relationship and time sequence relationship, so that the algorithm structure can handle time misalignment and annotation noise scenarios, and provide traceable intermediate evidence for threshold calibration, risk interpretation and confidence interval generation.
Claims
1. A data anomaly prediction method based on multimodal fusion, characterized in that, include: S1. Normalize and organize structured business data, log text, sensor time series data and manually labeled information according to entity objects and time windows to generate multimodal sample units; S2. Extract business status change features, log semantic event features, sensor multi-scale time-frequency features, and labeled prototype features from the multimodal sample units to generate a multimodal feature sequence; S3. Perform cross-modal comparison mapping on the multimodal feature sequence, and generate a time-delay confidence matrix indexed by modality pairs and offset windows according to the candidate offset windows, mapping each modal feature to an anomaly evidence vector of the same dimension to obtain an aligned evidence sequence; S4. Construct a dynamic graph sequence containing entity nodes, process nodes, log event nodes, sensor channel nodes, and labeled prototype nodes based on the aligned evidence sequence. Then, form an anomaly evidence memory graph by combining historical anomaly prototype nodes, labeled prototype nodes, and the current window node, and output a graph sequence with memory edge weights. S5. Input the graph sequence with memory edge weights into the trained graph time-series converter for cross-node aggregation and cross-time propagation to obtain the cumulative anomaly vector. Then, output the anomaly trend, anomaly score, risk type, contribution mode, and confidence interval of the future time window through the trend prediction head.
2. The method according to claim 1, characterized in that, S1 includes: Configure unified entity identifiers, data source identifiers, and collection timestamps for structured business data, log text, sensor time-series data, and manually labeled information; Divide the time window according to the preset window length and preset sliding step size; write data belonging to the same entity object and whose collection timestamps fall into the same time window into the same multimodal sample unit; write missing markers and missing durations for missing modalities; Numerical fields are standardized according to the mean and standard deviation of the training sample distribution; category fields are generated with category codes; event time and event template identifiers are retained in log text; channel identifiers, sampling frequency and sampling sequence are retained in sensor time series data; and annotation category, annotation time and annotation source are retained in manually annotated information.
3. The method according to claim 1, characterized in that, S2 includes: Business state change characteristics are generated based on the field differences, state transition counts, and process node dwell times in the continuous time window of structured business data. Perform event template matching and context encoding on log text to generate log semantic event features; Statistical quantities, frequency band energy, and abrupt change point locations are calculated for sensor time-series data according to a preset scale window to generate multi-scale time-frequency characteristics of the sensor. The features of samples with the same annotation category in the manually annotated information are aggregated to generate annotation prototype features; the business status change features, the log semantic event features, the sensor multi-scale time-frequency features, and the annotation prototype features are sorted according to entity objects and time windows to obtain the multimodal feature sequence.
4. The method according to claim 1, characterized in that, S3 includes: inputting each modal feature in the multimodal feature sequence into the corresponding modal encoder to obtain modal embedding vectors of the same dimension; using different modal embedding vectors within the same entity object and the same time window as positive sample pairs, and using modal embedding vectors with different entity objects or time windows as negative sample pairs, updating the modal encoder by comparison loss; calculating the alignment confidence of each modal pair under each candidate offset window based on the absolute value of the time difference after adding the log event time, sensor mutation point time, and manual annotation time to the candidate offset window respectively; performing a normalized exponential transformation on each alignment confidence of the same modal pair to generate the time-delay confidence matrix; and performing gated weighting on the modal embedding vectors under different candidate offset windows according to the time-delay confidence matrix to obtain the alignment evidence sequence.
5. The method according to claim 1, characterized in that, S4 includes: Entity relationships, business process adjacency relationships, sensor channel association relationships, and temporal adjacency relationships are respectively converted into entity edges, process edges, channel edges, and temporal edges in a dynamic graph sequence; historical anomaly prototypes, manually annotated prototypes, current entity nodes, log event nodes, and sensor channel nodes are added to the anomaly evidence memory graph; memory edge weights are determined based on the alignment confidence in the time-delay confidence matrix, the prototype similarity between the current window node and the historical anomaly prototype node, business process adjacency markers, and temporal sequence markers; anomaly evidence where the single-modality anomaly score is less than a preset alarm threshold and not less than the p-th quantile of the historical background distribution of that modality is written into the anomaly evidence memory graph, where p is a preset percentile parameter; the memory edge weights are written into the corresponding edge attributes to obtain the graph sequence with memory edge weights.
6. The method according to claim 1, characterized in that, S5 includes: in the graph sequence with memory edge weights, weighting and aggregating the anomalous evidence vectors of adjacent nodes according to the attributes of each edge to generate a node aggregation vector; performing time-series encoding on the node aggregation vectors of the same entity object according to the time window order to generate an entity time-series representation; fusing the entity time-series representation with the representation of the corresponding historical anomalous prototype node in the anomalous evidence memory graph to obtain the cumulative anomalous vector; inputting the cumulative anomalous vector into the trend prediction head, and outputting the anomalous trend category, anomalous score, risk type label, contribution weight of each modality, and upper and lower bounds of the confidence interval for the future time window.
7. The method according to claim 6, characterized in that, The trained graph time-series converter and the trend prediction head are obtained as follows: historical multimodal samples with future time window anomaly labels, risk type labels, and anomaly occurrence times are used as training samples; the training samples are processed through steps S1 to S4 to obtain a training graph sequence; the training graph sequence is input into the graph time-series converter and the trend prediction head to obtain the predicted anomaly score, predicted risk type, and predicted anomaly occurrence time. The model parameters are updated by weighting and summing the classification loss between the predicted anomaly score and the anomaly label, the classification loss between the predicted risk type and the risk type label, the time deviation loss between the predicted anomaly occurrence time and the labeled anomaly occurrence time, and the contrast loss, until the change in the weighted sum on the validation sample set within a preset number of rounds is less than the convergence threshold.
8. The method according to claim 6, characterized in that, The confidence interval is determined as follows: The true outlier scores for future time windows and the outlier scores output by the trend prediction head are obtained from the calibration sample set; the absolute value of the difference between the two is calculated as the non-consistent score; quantiles are determined from the distribution of the non-consistent scores according to the target coverage; the lower bound of the confidence interval is obtained by subtracting the quantiles from the outlier scores output by the trend prediction head, and the upper bound of the confidence interval is obtained by adding the quantiles to the outlier scores output by the trend prediction head; when the time range of the calibration sample set is updated, the quantiles are re-determined based on the updated distribution of the non-consistent scores.
9. The method according to claim 6, characterized in that, The modal contribution weights are determined as follows: the edge attention weights of the graph time-series converter during cross-node aggregation and the temporal attention weights during cross-time propagation are recorded; the attention weights of nodes, edges, and time windows corresponding to the same modal source are summed and normalized to obtain the initial modal contribution value; the initial modal contribution value is corrected according to the missing duration written in step S1 and the alignment confidence generated in step S3 to obtain the modal contribution weights; the contributing modality is determined according to the modality whose contribution weight value is ranked first.
10. The method according to claim 8, characterized in that, The anomaly scores are used for threshold calibration, specifically including: statistically analyzing the anomaly score distribution and confirmed anomaly labels for historical time windows within each update cycle; generating candidate alarm thresholds based on the q-th quantile of the anomaly score distribution, where q is a preset percentile parameter; calculating the false negative rate and false positive rate corresponding to the candidate alarm thresholds based on the confirmed anomaly labels; when the false negative rate is not greater than a preset false negative rate threshold and the false positive rate is not greater than a preset false positive rate threshold, writing the candidate alarm thresholds into the preset alarm thresholds of the next update cycle; when the false negative rate or the false positive rate exceeds the corresponding threshold, maintaining the current preset alarm thresholds and writing the samples of that update cycle into the calibration sample set.