A real-time data processing system and method
By performing time synchronization correction and pattern parsing on multi-source data, combined with cross-source interleaving processing and partitioned aggregation inference algorithms, the problem of event misalignment caused by time offset in multi-source data processing is solved, thereby improving the real-time performance and accuracy of data processing and ensuring the stability and processing efficiency of the system.
Patent Information
- Application Number
- CN202511461801.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies lack mechanisms to quickly eliminate minute time shifts and maintain event consistency in multi-source data processing, making it difficult to balance accuracy and stability in real-time data processing. This is especially true in millisecond-level response scenarios, which can easily lead to a decrease in the real-time performance of data processing results or a delay in triggering response actions.
By introducing a time synchronization correction mechanism to standardize and correct data from multiple sources, and combining pattern parsing to extract key time points and data sensitivity factors, an initial data sequence with emergency weights is formed. Then, through cross-source interleaving processing and partition aggregation inference algorithms, abnormal candidate segments are identified, a local correlation network with weight gradients is constructed, and finally, a corrected data processing sequence is generated.
It effectively eliminates the event misalignment problem caused by time drift at different acquisition ends, improves the consistency and accuracy of data fusion, enhances the real-time performance and stability of the system, and significantly improves processing efficiency and the ability to adapt to complex application scenarios.
Smart Images

Figure CN120929729B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to a real-time data processing system and method. BACKGROUND
[0002] With the rapid development of distributed computing and real-time streaming processing, real-time data processing methods have been widely applied in financial transaction monitoring, industrial production scheduling, network security protection and information services. The existing technology usually realizes the collection and merging of multi-source data through message queues and stream processing frameworks, but there are still some subtle but key problems in specific applications. For example, when continuous data streams from different collection ends arrive at the processing system within the same time window, due to insufficient clock synchronization accuracy or uneven transmission delay, the same event is easily divided into different processing fragments, resulting in deviation of statistical results and mismatch of decision logic. Although such deviation only represents the error of a small number of samples when the overall data volume is large, in scenarios requiring millisecond-level response, it will directly cause the real-time performance of data processing results to decline or the triggering delay of response actions. The existing technology lacks a processing mechanism that can quickly eliminate minor time offsets and maintain event consistency during data fusion, making it difficult to balance the accuracy and stability of real-time data processing. Therefore, it is necessary to design a real-time data processing system and method that improves the real-time performance of data processing. SUMMARY
[0003] In view of the deficiencies of the prior art, the present application provides a real-time data processing system and method, which has the advantage of improving the real-time performance of data processing and solving the problems in the background art.
[0004] To achieve the above-mentioned purpose of improving the real-time performance of data processing, the present application provides the following technical scheme: a real-time data processing method, comprising the following steps:
[0005] Instantaneously capturing continuous data streams from multiple source collection ends, and extracting key time points, data sensitive factors and processing boundary conditions triggered by events through a pattern analysis mechanism, thereby generating an initial data sequence with emergency weight;
[0006] Cross-source interleaving processing of the initial data sequence, combining the offset trajectory of the historical data period, the buffer interval of the adjacent data channel and the conflict frequency of multi-source input, to form a composite feature set with time sequence level characteristics;
[0007] On the basis of the composite feature set, using a partition aggregation deduction algorithm to extract the dynamic occupation path of the data event in the processing period, and marking out potential abnormal candidate fragments by judging the discrete deviation of the dynamic occupation path and the initial processing configuration;
[0008] The abnormal candidate segments are re-labeled, and key nodes with instant response characteristics, non-key nodes allowing delay and controlled scheduling nodes are fused within a specified time window to construct a local correlation network with weight gradient;
[0009] According to the local correlation network, data processing units with resource mismatch in the target period are detected and identified, and a corrected data processing sequence is derived in combination with the conflict shock trend caused by system fluctuations.
[0010] Preferably, the process of generating an initial data sequence with an emergency weight is as follows:
[0011] In the input stage of the multi-source acquisition end, the data acquisition time is standardized and corrected through a unified time synchronization mechanism;
[0012] Based on the pattern analysis mechanism, the event trigger information of the input data is preprocessed, and multi-dimensional parameters including key time points, data sensitive factors and processing boundary conditions are extracted;
[0013] An integrated scoring function is used to calculate the multi-dimensional parameters, assign emergency weight labels to each event, and sort the events according to the scoring results to form a continuous initial data sequence.
[0014] Preferably, the process of cross-source interleaving processing of the initial data sequence is as follows:
[0015] According to the trigger time and source of each data event in the initial data sequence, data segments of different sources are combined in an interleaved manner;
[0016] The time sequence relationship of each data segment is corrected in combination with the time offset trajectory, channel delay and data synchronization error in the historical data period;
[0017] The conflict frequency between multiple source inputs is analyzed at the same time, and the segments with frequent conflicts are marked as special nodes, so as to preserve the dependence and difference characteristics between data in the cross-source combination process, and finally generate a time sequence set with interleaving structure.
[0018] Preferably, the process of forming a composite feature set with time sequence level characteristics is as follows:
[0019] The data events after cross-source interleaving processing are mapped to a multi-dimensional feature space, and the emergency weight, trigger time and channel conflict frequency are fused;
[0020] In the feature space, a hierarchical aggregation algorithm is used to hierarchically divide adjacent events to generate multi-layer data clusters with upper and lower relationships;
[0021] By aligning different data clusters through normalization and temporal sorting mechanisms, a composite feature set that can reflect the hierarchical dependence and time series characteristics between data events is constructed.
[0022] Preferably, the process of extracting the dynamic path occupied by data events during the processing cycle using the partition aggregation inference algorithm is as follows:
[0023] Based on the composite feature set, the event nodes are divided into different partition regions, and the nodes in each region are aggregated and calculated.
[0024] The inference algorithm is used to simulate the occupancy behavior of event nodes on the time axis, and to generate the dynamic occupancy path of each node in different processing cycles;
[0025] By comparing the position and trend of different nodes in the aggregation results, potential path deviations and abnormal nodes can be identified.
[0026] Preferably, the process of marking potential anomalous candidate segments is as follows:
[0027] Compare the dynamic path occupancy with the preset standard processing configuration;
[0028] Using structural similarity calculation methods, discrete deviations and abnormal nodes appearing in the path are detected;
[0029] Combined with a threshold determination mechanism, data segments with deviations are anomaly-marked. The adaptive threshold is dynamically adjusted based on historical data distribution characteristics, real-time input fluctuation range, and frequency of anomaly occurrence, eliminating the need for manual reconfiguration during scene switching. Data segments with deviations are recorded as candidate segments.
[0030] Preferably, the process of re-classifying and labeling abnormal candidate segments is as follows:
[0031] Based on the key event nodes involved in the abnormal candidate segments, determine the core time intervals in which the key event nodes occurred;
[0032] Within the core time interval, critical nodes with instant response characteristics, non-critical nodes that allow for delays, and controlled scheduling nodes are integrated.
[0033] By assigning weight levels to different nodes based on their impact on the processing results, the abnormal candidate segments are re-graded and labeled.
[0034] Preferably, the process of constructing a local correlation network with weighted gradients is as follows:
[0035] The anomaly fragments and related nodes that are further labeled are mapped to the associated network to form a network graph containing high-priority and low-priority nodes;
[0036] Based on the dependencies and conflict characteristics between nodes, an edge set is established to form a multi-dimensional directed network;
[0037] By assigning priority differences among nodes through a weight gradient mechanism, a local association network that can reflect event dependencies and conflict distribution is constructed.
[0038] Preferably, the process of deriving the corrected data processing sequence is as follows:
[0039] Identify the key nodes with the highest resource utilization in a locally interconnected network;
[0040] By combining historical scheduling records and real-time input data, the resource competition relationship between key nodes and adjacent nodes is analyzed;
[0041] Identify and mark data processing units that have obvious resource mismatches or conflicting usage;
[0042] Based on the detected resource mismatch units, and combined with the priority weights and dependency paths of nodes in the local network, a dependency loop detection algorithm is used to identify potential cyclic structures in the local network. When a dependency loop is detected, a loop decoupling strategy is used, including priority breaking, virtual node insertion, or strong connectivity component decomposition, to eliminate the risk of deadlock and to make feasible adjustments to the task execution order.
[0043] By utilizing a dynamic scheduling intervention mechanism, the resource allocation ratio is optimized and the dependencies between nodes are coordinated to generate a corrected sequence of data processing results.
[0044] A real-time data processing system, comprising:
[0045] Sequence generation module: Captures and parses continuous data streams from multiple acquisition terminals, extracts key time points and data factors, generates an initial data sequence with emergency weights, and enables a cache queue to temporarily store interrupted paths when data access is abnormal, and supplements the processing after the acquisition terminal recovers.
[0046] Feature fusion module: Performs cross-source interleaving processing on the initial data sequence, combines historical offsets and conflict frequencies to form a composite feature set with temporal hierarchical characteristics;
[0047] Path inference module: Based on the composite feature set, it runs the partition aggregation inference algorithm, extracts the dynamic occupied path of data events and judges the deviation, marks abnormal candidate segments, and restores the last effective inference state through the snapshot backup mechanism when the operation is abnormal or crashes.
[0048] Hierarchical association module: It performs hierarchical labeling on abnormal candidate segments, merges different types of nodes, and constructs a local association network with weight gradients;
[0049] Correction output module: Based on the local correlation network detection resource mismatch unit, and combined with the conflict trend, the corrected data processing sequence is deduced. At the same time, a result verification and rollback mechanism is added in the output stage. When the corrected output has errors or delays, it will roll back to the most recent stable result.
[0050] Compared with the prior art, the present invention provides a real-time data processing system and method, which has the following beneficial effects:
[0051] This invention introduces a multi-source data time synchronization correction and pattern parsing mechanism, which effectively avoids event misalignment caused by time drift at different acquisition ends during the data acquisition stage, thereby improving the consistency and accuracy of data fusion. Through cross-source interleaving processing combined with historical offset trajectory analysis, it not only achieves time-series correction of multi-channel inputs but also preserves the dependencies and differences between events in the interleaving structure, making subsequent analysis more hierarchical. By dynamically extracting the occupied paths of data events in the processing cycle through a partitioned aggregation inference algorithm and combining it with difference discrimination to mark abnormal segments, it can achieve early warning of minor deviations and potential conflicts, improving the system's sensitivity to anomalies. The construction of hierarchical labeling and local association networks enables the system to allocate different weight levels based on the immediacy, latency, and controllability of events, thereby reflecting priority differences in resource scheduling and effectively alleviating resource mismatch problems. By dynamically scheduling intervention and inferring the corrected processing sequence, it ensures the real-time performance and stability of the data processing process while avoiding processing delays or redundant overhead caused by resource conflicts, significantly improving the overall system's processing efficiency and adaptability in complex application scenarios. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the method of the present invention;
[0053] Figure 2 This is a schematic diagram of the system of the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] Example 1: Please refer to Figure 1 As shown in the figure, a real-time data processing method according to an embodiment of the present invention includes the following steps:
[0056] S1: Real-time capture of continuous data streams from multiple acquisition points, and extraction of key event triggering points, data sensitivity factors and processing boundary conditions through pattern parsing mechanism, thereby generating an initial data sequence with emergency weights.
[0057] The process of generating the initial data sequence with urgent weights in S1 is as follows:
[0058] During the input phase at the multi-source acquisition end, a unified time synchronization mechanism is used to standardize and correct the data acquisition time. During the data input process at the multi-source acquisition end, due to differences in hardware clock accuracy, network transmission delay, and environmental interference, the timestamps of the same event in different sources often have millisecond to second-level offsets. In order to ensure that subsequent cross-source data can be compared and processed on the same time axis, this method introduces a unified time synchronization mechanism in the input phase, such as a clock calibration scheme based on network time protocol or precision time protocol. By periodically broadcasting time synchronization signals, detecting clock drift of the acquisition devices at the end, and dynamically adjusting the acquisition timestamps, all data are standardized to a unified reference clock before entering the processing flow.
[0059] Furthermore, when a network interruption occurs at the multi-source acquisition terminal, this method will not directly discard the data during the network outage, but will instead compensate and correct it using the following strategies:
[0060] Local clock maintenance mechanism: During network outages, the acquisition device continues to rely on the local clock to record timestamps and estimates its cumulative offset through a drift monitoring model;
[0061] Breakpoint compensation mechanism: After the network is restored, the system first obtains the reference time of the device before and after the network outage, and calculates the timestamp correction value during the network outage based on the historical offset trajectory and time drift curve.
[0062] Interpolation and realignment strategy: Based on the calculated correction value, the data during the network outage period is interpolated and replayed to maintain the temporal continuity of the data with other normal sources, thereby avoiding cross-source comparison errors caused by clock errors;
[0063] Through the above mechanism, this method can not only achieve unified alignment of cross-source time under normal circumstances, but also ensure the time sequence consistency of offline data and normal data in extreme scenarios.
[0064] The event triggering information of the input data is preprocessed based on the pattern parsing mechanism to extract multi-dimensional parameters, including key time points, data sensitivity factors, and processing boundary conditions. In the time-synchronized data stream, the event triggering data of various types is first decomposed into a structured form through the pattern parsing mechanism. For example, for environmental parameters collected by sensors (temperature, humidity, current changes, etc.), business events captured by log collection terminals (task submission, abnormal interruption, user request, etc.), or behavioral fragments (image frames, voice commands) in multimedia streams, parsing algorithms based on regular expression matching, feature extraction, or machine learning classifiers are used to identify key time points of the event (such as start time, peak point, end point), data sensitivity factors related to the priority of event processing (such as whether the data involves security or affects the core process), and important processing boundary conditions affecting task scheduling (such as timeout limits, resource upper and lower limits, dependency constraints). Through the extraction of multi-dimensional parameters, the originally unstructured or semi-structured input data is transformed into a standardized feature set that can be used for priority determination.
[0065] A comprehensive scoring function is used to calculate multidimensional parameters, assigning urgency weight labels to each event, and sorting the events according to the scoring results to form a continuous initial data sequence. After parameterizing key time points, sensitive factors, and boundary conditions, the urgency score of each event is calculated using the following formula:
[0066] ;
[0067] In the formula, Let i be the urgency score for event i. For time sensitivity, For data sensitivity, As for the urgency of the boundary conditions, , , These are the weighting coefficients for time sensitivity, data sensitivity, and boundary condition urgency, respectively; the weighting coefficients satisfy... It can be set in two ways:
[0068] Fixed configuration: Preset according to the application scenario, such as industrial monitoring scenario. , , ;
[0069] Adaptive configuration: Using the entropy weight method, weights are dynamically allocated based on the information entropy of each indicator in the sample set;
[0070] Each event is weighted based on a comprehensive scoring function. The comprehensive scoring function can be combined with a multi-factor linear weighting model or a dynamic weighting strategy based on entropy weighting. The time sensitivity (such as timeout risk), data sensitivity (such as whether it involves critical resources or security risks), and boundary condition urgency (such as the severity of resource restrictions) of the event are assigned weight coefficients and normalized to output a comprehensive urgency score. Each event is sorted from high to low according to this score. High-scoring events enter the data processing channel first, while low-scoring events enter the buffer or delayed processing sequence.
[0071] S2: Perform cross-source interleaving processing on the initial data sequence, and combine the offset trajectory of historical data cycles, the buffer between adjacent data channels, and the collision frequency of multi-source inputs to form a composite feature set with time-series hierarchical characteristics.
[0072] The process of cross-source interleaving of the initial data sequence in S2 is as follows:
[0073] Based on the trigger time and collection source of each data event in the initial data sequence, data fragments from different sources are interleaved and combined. After obtaining the initial data sequence with emergency weight labels, the trigger timestamp and collection source identifier attached to each event are parsed and classified. Different collection sources may include video surveillance, sensor devices, business logs, or network traffic monitoring terminals. The system interleaves and combines these data fragments from different sources on a unified time axis, that is, it merges events according to the order of trigger time while keeping their original source attributes intact. In order to avoid data loss or duplication, an event index table and a source mapping table are introduced during the interleaving process to aggregate events from different sources in the same time period into a logical fragment group, thereby ensuring that the contribution of each collection source can still be distinguished after data fusion.
[0074] By combining the time offset trajectory, channel delay, and data synchronization error that appear in the historical data period, the timing relationship of each data segment is corrected. Since multi-source acquisition systems inevitably experience network congestion, processing delay, and device clock drift during operation, the arrival time of the same event may be offset in different sources. After interleaving and combining, it is necessary to use the time offset trajectory recorded in the historical data period (such as the long-term +5ms drift of a certain type of sensor), channel delay characteristics (such as an average delay of 20ms for a certain network link), and synchronization error distribution model (such as a normal distribution offset range of ±2ms) to dynamically correct the timing relationship of the interleaved data segments. In specific implementation, sliding window comparison, Kalman filtering, or offset correction algorithms are used to align the triggering order of events from different sources, ensuring that the interleaved segment sequence can more accurately reflect the actual event occurrence order, effectively reducing the event misalignment caused by delay or error, and avoiding misjudgment of the order of events in subsequent processing stages.
[0075] Simultaneously, the system analyzes the frequency of conflicts between multiple input sources, marking frequently conflicting segments as special nodes. This preserves the dependencies and differences between data during cross-source combination, ultimately generating a time-series set with an interleaved structure. In the corrected interleaved sequence, the system further performs conflict detection on data segments from different sources, determining whether multiple events within the same time window are competing for the same processing resources, occupying the same time slot, or mutually exclusively affecting the same control unit. For example, in traffic signal control, there might be simultaneous vehicle passage events from geomagnetic sensors and vehicle-free events from cameras; these conflicting events need to be identified. The system statistically analyzes the frequency of these conflicts, marking high-frequency conflicting segments as special nodes, while explicitly preserving their dependencies and differences within the interleaved structure.
[0076] The process of forming a composite feature set with temporal hierarchical characteristics in S2 is as follows:
[0077] The system maps cross-source interleaving data events to a multi-dimensional feature space, integrating urgency weights, trigger times, and channel conflict frequencies. After the cross-source interleaving step, the system obtains an event sequence that has undergone time correction and conflict labeling. Each event is no longer just a single point in time but contains multiple measurable attributes, such as urgency weight labels (reflecting the importance of the event), trigger times (the specific location on a unified time axis), channel conflict frequencies (the number of times events from different sources conflict within that time period), and optional additional parameters such as acquisition source type and delay correction. To facilitate subsequent complex hierarchical analysis, the system maps these attributes to a multi-dimensional feature space. The specific implementation includes setting independent coordinate axes for each parameter and mathematically expressing the event in the form of feature vectors. For example, an event can be represented as a vector (W, T, C, S), where W represents the weight, T represents the trigger time, C represents the conflict frequency, and S represents the acquisition source category. Through this vectorized modeling, the original time series is expanded into high-dimensional data points, laying the foundation for hierarchical clustering and hierarchical dependency analysis.
[0078] In the feature space, a hierarchical aggregation algorithm is used to hierarchically partition adjacent events, generating multi-level data clusters with hierarchical relationships. After mapping events to the feature space, a hierarchical aggregation algorithm or a partitioning aggregation method based on temporal proximity is used to progressively aggregate adjacent events. In the specific implementation, a multi-dimensional distance metric between events is calculated, where urgency weight, trigger time, and conflict frequency occupy different computational weights to ensure that key parameters are highlighted during the partitioning process. The algorithm follows a proximity-to-distance approach, first aggregating events with the smallest distance into a low-level cluster, and then gradually merging them into higher-level clusters, thus forming a tree-like or hierarchical structure. The final result is a series of data clusters with hierarchical relationships: the bottom-level clusters represent event groups that are highly similar in time and features, while the upper-level clusters represent macro-level event patterns across time periods or sources. This hierarchical partitioning not only reveals the local correlations between events but also constructs multi-level dependencies, which is helpful for subsequent path deduction and anomaly detection.
[0079] By aligning different data clusters through normalization and temporal sorting mechanisms, a composite feature set reflecting the hierarchical dependencies and time-series characteristics between data events is constructed. Since data clusters from different sources may exhibit inconsistent scales and uneven time spans during their formation, the aggregation results need to be standardized and sorted. Through normalization, parameters on different dimensions (such as urgency weight and conflict frequency) are scaled within a range to ensure their values fall within a uniform range (e.g., 0–1), thus preventing a single dimension's excessively large value from dominating the entire aggregation result. A temporal sorting mechanism is introduced based on the aggregated clusters, rearranging and aligning each data cluster on a unified time axis according to the event trigger time or the average time of the cluster center. This ensures logical consistency of the hierarchical clusters in time. The final result is a composite feature set that not only preserves the hierarchical dependencies between multi-source events but also clearly presents the chronological order of the time series. This allows it to serve as a foundational input in subsequent dynamic path deduction, improving the accuracy of abnormal segment identification.
[0080] S3: Based on the composite feature set, the dynamic occupancy path of data events in the processing cycle is extracted using the partition aggregation inference algorithm, and potential abnormal candidate segments are marked by judging the discrete deviation between the dynamic occupancy path and the initial processing configuration.
[0081] The process in S3 of extracting the dynamic path occupancy of data events during the processing cycle using the partition aggregation inference algorithm is as follows:
[0082] Based on the composite feature set, event nodes are divided into different partition regions, and aggregation calculations are performed on the nodes in each region. After hierarchical processing, the composite feature set already contains the multidimensional attributes and temporal relationships of event nodes. Based on feature parameters such as time span, emergency weight distribution, and conflict frequency, the time axis is divided into several partition regions. For example, the segmentation is performed by using signal period, data acquisition period, or calculation batch as boundaries. Event nodes falling into the same region are classified into that partition, and aggregation calculations are performed based on the multidimensional features of the nodes (such as the average emergency weight, the sum of conflict frequencies, and the distribution range of time points) to obtain the aggregated feature vector or center point parameter representing the region. This processing method can effectively reduce the data scale, avoid the complexity of analyzing each node one by one, and maintain the representativeness of events in local temporal sequence and feature dimensions.
[0083] The algorithm simulates the occupancy behavior of event nodes on the time axis, generating dynamic occupancy paths for each node in different processing cycles. The aggregated feature vector of each partition is mapped onto a unified time axis, and the occupancy range of the region is defined according to the start and end boundaries of the partition. The simulation is carried out step by step on the time axis to simulate the occupancy of each node in different processing cycles, such as the start point, duration, and resource utilization of a node in a cycle. The simulation process can adopt a sliding window mechanism and a time step iteration mechanism to ensure that the simulation results can cover all possible states in the cycle. The final result is a series of dynamic occupancy paths, which are essentially the time segments and occupancy state sequences experienced by each event node in the cycle evolution process, used to characterize the dynamic behavior pattern of the node over time.
[0084] By comparing the positions and trends of different nodes in the aggregation results, potential path deviations and abnormal nodes are identified. After the dynamic occupancy path is generated, horizontal and vertical comparisons are performed on the trajectories of different nodes. Horizontal comparison mainly analyzes the spatial distribution differences of different nodes on the path within the same period. For example, if the occupancy position of some nodes is significantly advanced or delayed, it indicates that there is an abnormal time deviation. Vertical comparison tracks the changing trend of the same node in multiple periods. For example, if a node shows an abnormal increase in occupancy time in multiple periods, it indicates abnormal resource consumption or system bottleneck problems. Statistical analysis is performed on the path, such as calculating the similarity, offset, and frequency of anomalies between paths, to quantify the degree of anomaly. Through the above analysis, abnormal nodes that deviate from the normal trajectory can be identified and marked as candidates for subsequent re-leveling processing.
[0085] The process of marking potential abnormal candidate segments in S3 is as follows:
[0086] The dynamic occupancy paths are compared with the preset standard processing configuration. A standard processing configuration model is established, which records the ideal occupancy position, duration, and resource allocation status of each event node within the processing cycle under normal conditions. Each dynamic occupancy path generated by the partition aggregation inference algorithm is mapped to the time axis and resource dimension corresponding to the standard configuration model. A node-by-node comparison is performed to calculate the time offset (difference between the actual occupancy start and end time and the standard time), position deviation (difference in the relative order of the node in the resource allocation sequence), and occupancy status difference (difference between the actual resource consumption and the standard allocation value) of each node. By quantifying these differences, it is possible to initially identify which nodes or time segments exhibit abnormal behavior. Partitioning criteria: A fixed time window length is used to divide data events into continuous partitions on the time axis. Each partition contains all event nodes within the time window. Aggregation calculation operator: Within each partition, the average resource occupancy rate, maximum latency, and conflict frequency of the nodes are statistically analyzed. Using the time window length as the basic step size, the occupancy situation of each partition is inferred forward sequentially to form a dynamic occupancy path trajectory across cycles.
[0087] Using structural similarity calculation methods, discrete deviations and abnormal nodes appearing in the path are detected; when detecting differences between dynamically occupied paths and standard processing configurations, an improved cosine similarity is used, with the formula:
[0088] ;
[0089] In the formula, P is the node occupancy vector of the dynamic occupancy path within the time window, Q is the reference vector of the standard processing configuration within the time window, n is the number of feature dimensions, and k is the feature weight adjustment factor.
[0090] After completing the difference comparison, a structured analysis is performed on the dynamically occupied paths. Each path is represented as a sequence of nodes or a graph structure. Graph structure similarity algorithms (such as weighted graph matching based on node attributes, shortest path similarity, or adjacency matrix comparison) are used to calculate the similarity between the path and the standard configuration. The location, order, dependencies, and occupation time of each node are evaluated to identify nodes that have structural mutations or misalignments, such as nodes being advanced or delayed, repeated, or missing. Through this structured similarity analysis, not only can time deviations be captured, but also discrete abnormal nodes caused by resource conflicts or processing anomalies can be identified.
[0091] By combining a threshold determination mechanism, data segments with deviations are anomaly-labeled. The adaptive threshold is dynamically adjusted based on historical data distribution characteristics, real-time input fluctuation range, and frequency of anomalies, eliminating the need for manual reconfiguration during scene switching. Data segments with deviations are recorded as candidate segments. After obtaining the deviation value and structural similarity index of each node or segment, a threshold determination mechanism is pre-set, including time offset threshold, resource usage deviation threshold, and structural similarity lower limit. Each data segment is checked item by item. When any deviation index exceeds the preset threshold, the data segment is marked as an anomalous data segment. Information such as the start and end time of the anomalous segment, the nodes involved, the deviation type, and the degree of deviation are recorded to form a list of candidate anomalous segments for subsequent re-level labeling and local association network construction. In this way, potential anomalous segments can be systematically and quantitatively identified.
[0092] S4: Re-label the abnormal candidate segments and merge key nodes with immediate response characteristics, non-key nodes that allow delay and controlled scheduling nodes within a specified time window to construct a local association network with weighted gradients.
[0093] The process of re-classifying and labeling abnormal candidate segments in S4 is as follows:
[0094] Based on the key event nodes involved in the candidate anomaly segments, the core time intervals in which the key event nodes occur are determined. All key event nodes are extracted from the list of candidate anomaly segments. These nodes are usually characterized by significantly deviating from the standard configuration in terms of occupied time, abnormal resource usage, or significantly low structural similarity. The time attributes of these key nodes are analyzed, including start time, end time, and relative position within the processing cycle. By statistically analyzing the offset amplitude and frequency of nodes in different cycles, the core time intervals in which anomalies occur can be determined, that is, the time period in which key nodes appear in concentration and have the greatest impact on the processing results. The core interval will serve as the basis for subsequent hierarchical labeling and local association network construction to ensure that the labeling covers the most critical anomaly events.
[0095] Within the core time interval, key nodes with immediate response characteristics, non-key nodes that allow for delays, and controlled scheduling nodes are integrated. Within the defined core time interval, node types are classified. Nodes that have an immediate impact on the processing results are identified as key nodes, which typically involve high-priority tasks or time-sensitive events. Nodes whose time consumption can be flexibly adjusted and have a smaller impact on the overall processing results are marked as non-key nodes, which can be assigned lower weights in the hierarchical labeling. Nodes controlled by the system scheduling strategy (such as resource allocation nodes that can be adjusted by the scheduler) are identified as controlled scheduling nodes. By aggregating the information of these nodes, a multi-attribute node set is formed within the core time interval, achieving unified integration of different types of nodes.
[0096] The system assigns weights to different nodes based on their impact on the processing results, thus completing the re-classification and labeling of candidate anomalies. The fused node set undergoes quantitative evaluation, with each node's weight calculated based on multi-dimensional indicators, including urgency weight, time offset, resource contention intensity, dependency on other nodes, and historical anomaly frequency. Nodes are assigned levels based on these indicators, such as high, medium, and low, or continuously divided into levels by percentage, to reflect their actual impact on the processing results. After weight allocation, each node within a candidate anomaly segment will have a clear level label, achieving re-classification and labeling.
[0097] The process of constructing a local correlation network with weighted gradients in S4 is as follows:
[0098] The re-graded and labeled anomaly fragments and their related nodes are mapped to the associated network to form a network graph containing high- and low-priority nodes. All node information, including key nodes, non-key nodes, and controlled scheduling nodes, is extracted from the re-graded and labeled anomaly candidate fragments. At the same time, the weight level and time attribute of each node are retained. These nodes are mapped to a node set, and an initial network framework is constructed with nodes as vertices. The position of nodes in the network can be arranged according to their time sequence, weight level, and dependency relationship to ensure that high-priority nodes are prominently reflected in the network topology, while low-priority nodes are in auxiliary or peripheral positions. Through this mapping method, a multi-level and multi-priority network graph is initially formed, which not only retains the classification information of nodes, but also lays the foundation for the establishment of the edge set.
[0099] Based on the dependencies and conflict characteristics between nodes, an edge set is established to form a multi-dimensional directed network. On the basis of the preliminary network graph, the dependencies between nodes (such as sequence, resource sharing, and processing logic transmission) and conflict characteristics (such as time overlap, resource contention, and frequency of abnormal triggering) are analyzed. For each pair of nodes with logical or conflicting relationships, a directed edge is established in the network. The direction of the edge represents the direction of dependency transmission or priority transmission. The attributes of the edge can include multi-dimensional information such as conflict intensity, dependency tightness, and historical abnormal frequency. In this way, the original set of nodes is transformed into a multi-attribute, directed network structure that can clearly reflect the sequence, dependency logic, and potential conflict situations between events.
[0100] By assigning priority differences among nodes through a weight gradient mechanism, a local association network reflecting event dependencies and conflict distribution is constructed. After the directed network is constructed, a weight gradient mechanism is introduced based on the weight level of the nodes and the dependency and conflict attributes of the edges. The influence of high-priority nodes is passed down along the network edges, while the weight values of low-priority nodes in the network remain relatively low or are adjusted according to the dependency relationship. The initial weight value of the nodes, the cumulative influence along the path, the weighted attenuation of the conflict intensity, and the summarization of the multi-path dependencies that appear in the network are calculated. After the weight gradient allocation, each node not only has its own weight information, but also reflects its influence on other nodes in the local network. The final local association network can fully reflect the dependency relationship, conflict distribution, and priority level of events in the abnormal candidate fragments.
[0101] S5: Based on the local correlation network, detect and identify data processing units with resource mismatch within the target period, and deduce the corrected data processing sequence by combining the conflict oscillation trend caused by system fluctuations.
[0102] The process of deriving the corrected data processing sequence in S5 is as follows:
[0103] In the local relational network, the key nodes with the highest resource utilization are identified. Based on the constructed local relational network, the resource utilization of all nodes is quantitatively evaluated. The evaluation indicators include CPU computing utilization, memory cache consumption, bandwidth usage, and I / O channel utilization. The average and peak utilization values of the nodes in different time windows are calculated by time series analysis. Combined with the priority weight of the nodes, the resource consumption curve is weighted and corrected to ensure that the resource utilization characteristics of high-priority nodes are identified. Through sorting and threshold screening, the key nodes with the highest resource utilization and significant impact on the overall processing performance are determined and used as the starting point for correcting the data processing sequence.
[0104] By combining historical scheduling records and real-time input data, the resource competition relationship between key nodes and adjacent nodes can be analyzed. The real-time input data is dynamically parsed to extract the latest task arrival rate, event trigger frequency and resource demand. By integrating historical records with real-time input, a resource competition model between key nodes and adjacent nodes can be constructed. By using conflict frequency statistics, resource utilization curve fitting and competition intensity index calculation methods, it can be analyzed which nodes are competing for resources or handling conflicts within the same time window.
[0105] Data processing units with obvious resource mismatch or conflict are identified and marked. If the resource demand of a node exceeds the system's available resource threshold, or multiple nodes compete for the same resource in a high-intensity manner during the same time period, it is determined that the node or combination of nodes has a resource mismatch and conflict problem. These nodes are assigned abnormal labels and marked in the local network with special markings, such as red marking high-risk conflict nodes and yellow marking medium-level resource mismatch nodes. After this processing, the bottlenecks and conflict hotspots of resource allocation can be clearly displayed in the entire network.
[0106] Based on the detected resource mismatch units, and combining the priority weights and dependency paths of nodes in the local associative network, a dependency loop detection algorithm is used to identify potential cyclic structures in the local network. When a dependency loop is detected, a loop decoupling strategy is used, including priority breaking, virtual node insertion, or strong connectivity component decomposition, to eliminate the risk of deadlock and to make feasible adjustments to the task execution order. After marking conflict and mismatch units, a task reordering algorithm is invoked to adjust the execution paths in the local associative network. The dependency relationships between high-priority nodes and other nodes are analyzed to ensure that logical integrity is not compromised when adjusting the task order. Based on resource consumption and conflict intensity, the execution time of mismatched nodes is shifted forward or backward to reduce the probability of concurrent conflicts. Combining the node weight gradient, priority is given to ensuring the processing time of critical nodes and immediate response nodes, while deferred nodes are allocated to time intervals with lower conflict. Through this process, an optimized execution order is generated, which can effectively alleviate resource consumption conflicts.
[0107] By utilizing a dynamic scheduling intervention mechanism, the resource allocation ratio is optimized and the dependencies between nodes are coordinated to generate a corrected data processing result sequence. The load curves of each resource are monitored in real time, and the resource allocation ratio between critical and non-critical nodes is dynamically adjusted. When a sudden increase in resource demand is detected at a node, temporary reallocation of bandwidth, CPU cores, or memory is automatically triggered. If a node dependency path is blocked, the stability of the overall dependency link is maintained by adjusting the task concurrency or delaying low-priority tasks. Through this mechanism, the final output is a corrected data processing sequence. While ensuring logical integrity and correct priority, the sequence significantly reduces the occurrence rate of resource mismatch and conflict, thereby improving the overall real-time processing capability and stability of the system.
[0108] Example 2: Figure 2 As shown, a real-time data processing system includes:
[0109] Sequence generation module: Captures and parses continuous data streams from multiple acquisition terminals, extracts key time points and data factors, generates an initial data sequence with emergency weights, and enables a cache queue to temporarily store interrupted paths when data access is abnormal, and supplements the processing after the acquisition terminal recovers.
[0110] Feature fusion module: Performs cross-source interleaving processing on the initial data sequence, combines historical offsets and conflict frequencies to form a composite feature set with temporal hierarchical characteristics;
[0111] Path inference module: Based on the composite feature set, it runs the partition aggregation inference algorithm, extracts the dynamic occupied path of data events and judges the deviation, marks abnormal candidate segments, and restores the last effective inference state through the snapshot backup mechanism when the operation is abnormal or crashes.
[0112] Hierarchical association module: It performs hierarchical labeling on abnormal candidate segments, merges different types of nodes, constructs a local association network with weight gradients, and enables interpolation and redundancy compensation strategies when nodes are missing or data is incomplete.
[0113] Correction output module: Based on the local correlation network detection resource mismatch unit, and combined with the conflict trend, the corrected data processing sequence is deduced. At the same time, a result verification and rollback mechanism is added in the output stage. When the corrected output has errors or delays, it will roll back to the most recent stable result.
[0114] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0115] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A real-time data processing method, characterized in that, Includes the following steps: The system captures continuous data streams from multiple acquisition points in real time and extracts key event triggering points, data sensitivity factors, and processing boundary conditions through a pattern parsing mechanism, thereby generating an initial data sequence with emergency weights. The initial data sequence is subjected to cross-source interleaving processing, and combined with the offset trajectory of historical data cycles, the buffer between adjacent data channels, and the collision frequency of multi-source inputs, a composite feature set with time-series hierarchical characteristics is formed. The process of performing cross-source interleaving on the initial data sequence is as follows: Based on the trigger time and collection source of each data event in the initial data sequence, data fragments from different sources are interleaved and combined. By combining the time offset trajectory, channel delay and data synchronization error that appear in the historical data cycle, the timing relationship of each data segment is corrected; Simultaneously, the frequency of conflicts between multiple input sources is analyzed, and fragments with frequent conflicts are marked as special nodes. This preserves the dependency and difference characteristics between data during cross-source combination, and finally generates a time series set with an intertwined structure. Based on the composite feature set, the dynamic occupancy path of data events in the processing cycle is extracted using the partition aggregation inference algorithm, and potential abnormal candidate segments are marked by judging the discrete deviation between the dynamic occupancy path and the initial processing configuration. The process of extracting the dynamic path occupied by data events during the processing cycle using the partition aggregation inference algorithm is as follows: Based on the composite feature set, the event nodes are divided into different partition regions, and the nodes in each region are aggregated and calculated. The inference algorithm is used to simulate the occupancy behavior of event nodes on the time axis, and to generate the dynamic occupancy path of each node in different processing cycles; By comparing the position and trend of different nodes in the aggregation results, potential path deviations and abnormal nodes can be identified. The abnormal candidate segments are re-labeled and then fused within a specified time window. Key nodes with immediate response characteristics, non-key nodes that allow delay, and controlled scheduling nodes are used to construct a local association network with weighted gradients. The process of constructing a locally related network with weighted gradients is as follows: The anomaly fragments and related nodes that are further labeled are mapped to the associated network to form a network graph containing high-priority and low-priority nodes; Based on the dependencies and conflict characteristics between nodes, an edge set is established to form a multi-dimensional directed network; By assigning priority differences among nodes through a weight gradient mechanism, a local association network that can reflect event dependencies and conflict distribution is constructed. Based on the local correlation network, data processing units with resource mismatches within the target period are detected and identified. Combined with the conflict oscillation trend caused by system fluctuations, the corrected data processing sequence is deduced.
2. The real-time data processing method according to claim 1, characterized in that, The process of generating an initial data sequence with urgency weights is as follows: During the input phase at the multi-source acquisition terminals, the data acquisition time is standardized and corrected through a unified time synchronization mechanism; Based on the pattern parsing mechanism, the event triggering information of the input data is preprocessed to extract multi-dimensional parameters, including key time points, data sensitivity factors and processing boundary conditions. The multidimensional parameters are calculated using a comprehensive scoring function, each event is assigned an emergency weight label, and the events are sorted according to the scoring results to form a continuous initial data sequence.
3. The real-time data processing method according to claim 2, characterized in that, The process of forming a composite feature set with temporal hierarchical characteristics is as follows: The data events after cross-source interleaving are mapped to a multi-dimensional feature space, and the emergency weight, trigger time and channel conflict frequency are integrated; In the feature space, a hierarchical aggregation algorithm is used to divide adjacent events into hierarchical levels, generating multi-level data clusters with hierarchical relationships; By aligning different data clusters through normalization and temporal sorting mechanisms, a composite feature set that can reflect the hierarchical dependence and time series characteristics between data events is constructed.
4. The real-time data processing method according to claim 3, characterized in that, The process of identifying potential anomalous candidate segments is as follows: Compare the dynamic path occupancy with the preset standard processing configuration; Using structural similarity calculation methods, discrete deviations and abnormal nodes appearing in the path are detected; By combining an adaptive threshold determination mechanism, data segments with deviations are marked as anomalies. The adaptive threshold is dynamically adjusted based on the distribution characteristics of historical data, the fluctuation range of real-time input, and the frequency of anomalies, without the need for manual reconfiguration when switching scenes, and the data segments with deviations are recorded as candidate segments.
5. The real-time data processing method according to claim 4, characterized in that, The process of re-classifying and labeling abnormal candidate segments is as follows: Based on the key event nodes involved in the abnormal candidate segments, determine the core time intervals in which the key event nodes occurred; Within the core time interval, critical nodes with instant response characteristics, non-critical nodes that allow for delays, and controlled scheduling nodes are integrated. By assigning weight levels to different nodes based on their impact on the processing results, the abnormal candidate segments are re-graded and labeled.
6. The real-time data processing method according to claim 5, characterized in that, The process of deriving the corrected data processing sequence is as follows: Identify the key nodes with the highest resource utilization in a locally interconnected network; By combining historical scheduling records and real-time input data, the resource competition relationship between key nodes and adjacent nodes is analyzed; Identify and mark data processing units that have obvious resource mismatches or conflicting usage; Based on the detected resource mismatch units, and combined with the priority weights and dependency paths of nodes in the local network, a dependency loop detection algorithm is used to identify potential cyclic structures in the local network. When a dependency loop is detected, a loop decoupling strategy is used, including priority breaking, virtual node insertion, or strong connectivity component decomposition, to eliminate the risk of deadlock and to make feasible adjustments to the task execution order. By utilizing a dynamic scheduling intervention mechanism, the resource allocation ratio is optimized and the dependencies between nodes are coordinated to generate a corrected sequence of data processing results.
7. A real-time data processing system, applied to the method as described in any one of claims 1-6, characterized in that, include: Sequence generation module: Captures and parses continuous data streams from multiple acquisition terminals, extracts key time points and data factors, generates an initial data sequence with emergency weights, and enables a cache queue to temporarily store interrupted paths when data access is abnormal, and supplements the processing after the acquisition terminal recovers. Feature fusion module: Performs cross-source interleaving processing on the initial data sequence, combines historical offsets and conflict frequencies to form a composite feature set with temporal hierarchical characteristics; Path inference module: Based on the composite feature set, it runs the partition aggregation inference algorithm, extracts the dynamic occupied path of data events and judges the deviation, marks abnormal candidate segments, and restores the last effective inference state through the snapshot backup mechanism when the operation is abnormal or crashes. Hierarchical association module: It performs hierarchical labeling on abnormal candidate segments, merges different types of nodes, constructs a local association network with weight gradients, and enables interpolation and redundancy compensation strategies when nodes are missing or data is incomplete. Correction output module: Based on the local correlation network detection resource mismatch unit, and combined with the conflict trend, the corrected data processing sequence is deduced. At the same time, a result verification and rollback mechanism is added in the output stage. When the corrected output has errors or delays, it will roll back to the most recent stable result.
Citation Information
Patent Citations
Multi-scale joint optimization multivariable time sequence anomaly detection method and system
CN118484756A
Emergency treatment data quality control analysis method and device based on big data
CN120183596A