Multi-source heterogeneous data-based industry return on investment real-time acquisition method

Through UTC benchmark and adjustable gain normalization of multi-source timestamps, combined with buffer queue and watermark threshold reordering, the time reference maintenance problem of multi-source heterogeneous data under high concurrency conditions is solved, high-precision timing alignment and distributed fault tolerance are achieved, and real-time calculation accuracy of financial transactions and return on investment is improved.

CN120295993AActive Publication Date: 2025-07-11JIANGXI LAYOUT DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510379568.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In multi-source heterogeneous data processing, it is difficult for the prior art to maintain unified time reference and high-quality output of data under high concurrency conditions, especially in node failure or network jitter, resulting in data loss or time stamp divergence, affecting the real-time calculation accuracy of high-frequency trading and return on investment in the financial field.

Method used

UTC reference, adjustable gain and power index are used to normalize multi-source timestamps, combine buffer queues and watermark thresholds to reorder out-of-order data, interpolation and conflict processing are performed, and the timestamps are fine-tuned using external high-precision reference signals, and fault-tolerant consensus and difference repair are triggered in case of failure to ensure data consistency and availability.

Benefits of technology

Maintain high-precision timing alignment and high availability in massive concurrent scenarios, reduce the risk of data loss or accuracy attenuation caused by timing inconsistency and node failure, and improve the accuracy and robustness of real-time analysis of multi-source heterogeneous data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295993A_ABST
    Figure CN120295993A_ABST
Patent Text Reader

Abstract

The invention discloses an industry return on investment real-time acquisition method based on multi-source heterogeneous data, and relates to the technical field of data analysis and processing, and the method comprises the following steps: normalizing a multi-source timestamp and retaining metadata information based on a UTC reference, an adjustable gain and a power exponent; in the streaming processing, reordering of out-of-order data and marking of expired data are realized by setting a buffer queue and a watermark threshold value; when null value fields or cross-source conflicts are detected, interpolation and conflict processing are carried out, and complete data subjected to consistency correction are output; then, an external high-precision reference or a cross-correlation function is used for further fine tuning the timestamp at a millisecond level, and if the adjustment amplitude exceeds a safety boundary, the timestamp is marked as suspicious; and finally, the corrected data is distributed to multiple nodes according to a fragment mapping strategy, and fault-tolerant consensus and difference repair are triggered under fault duration judgment, so that high-precision time sequence alignment and high availability in a massive concurrent scene are kept, data loss or precision attenuation caused by time sequence inconsistency and node faults is avoided, and the method can be widely applied to high-frequency analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis and processing, and specifically to a method for real-time obtaining of industry return on investment based on multi-source heterogeneous data. Background Art

[0002] In the current era of highly integrated data and real-time information interconnection, multi-source heterogeneous data generated by a large number of different industries and fields exhibits the characteristics of large scale, diverse structural forms, and high update frequencies. Especially in financial and industrial scenarios, data sources may include various channels such as trading systems, Internet of Things devices, policy announcements, news information flows, etc. Due to obvious differences in data accuracy, time zones, and release times among different channels, if a unified time mark and multi-source alignment mechanism are not established at the initial stage, subsequent real-time decision-making analysis often faces serious out-of-order and missing problems. At the same time, to meet the high-frequency processing requirements of seconds or even sub-milliseconds, the system not only needs to be able to perform streaming management and fault tolerance on large-scale concurrent data, but also needs to take into account alignment with external high-precision signals (such as synchronous clocks or high-frequency trading benchmarks). If there is a lack of effective caching and watermarking strategies during this process, or if distributed fault tolerance cannot be achieved in case of failures, it is extremely easy to cause distortion of key indicators (such as industry return on investment or other core business parameters), thus affecting the timeliness and accuracy of decision-making. Therefore, constructing a real-time data processing solution that can achieve an accurate unified time benchmark under multi-source heterogeneous inputs and has high availability capabilities in aspects such as streaming processing, data interpolation, high-precision alignment, and distributed fault tolerance has become a key technical direction concerned in fields such as financial transactions, industrial automation, and the Internet of Things.

[0003] In the Chinese patent invention with the authorization announcement number CN117171534B, a method, system, device, and medium for obtaining multi-source heterogeneous data of a numerically controlled machine tool are proposed, belonging to the technical field of data acquisition. The method includes: identifying the data types included in the multi-source heterogeneous data to be collected, and generating corresponding acquisition tasks according to the data types; executing the acquisition tasks, collecting data from heterogeneous data sources, parsing the communication protocols used by the heterogeneous data sources, and extracting their valid data; performing dimensionality reduction processing on the valid data of the heterogeneous data sources to convert it into low-dimensional data of the heterogeneous data sources with interpretability; performing data synchronization processing on the low-dimensional data of the heterogeneous data sources, and after the processing is completed, sending it to a data receiving end and performing data consistency verification. By optimizing the structural design and parsing for different protocols, the efficiency and accuracy of data transmission can be improved, and the real-time performance and synchronization of multi-source heterogeneous data can be ensured.

[0004] However, combining the above actual application scenarios and the above existing technologies:

[0005] When performing real-time processing and analysis on multi-source heterogeneous data, one of the core technical difficulties that urgently need to be solved is how to perform deeper refinement and reliable distribution of the preliminarily aligned and interpolated data under different network environments and high concurrency conditions, ensuring that a unified time reference and high-quality data output can still be maintained when node failures or load surges occur. For example, when the system has completed the second-level or millisecond-level time alignment of multi-source heterogeneous data, if it is deployed in a distributed cluster but lacks a multi-node fault tolerance mechanism and a clock synchronization consensus algorithm between nodes, it will fall into a dilemma of increased latency, data loss, or divergence of cross-node timestamps when individual nodes fail or network jitters occur, thus making it impossible to maintain the stable continuation of the high-precision alignment results achieved by consuming a large amount of computing resources in the early stage in a large-scale environment. This problem is particularly prominent in the financial field: in high-frequency trading, whether it is the real-time calculation of investment returns or risk control warnings, it is necessary to maintain the consistency of cross-node data at sub-millisecond time intervals. Once a node goes down and there is no compensation mechanism, the overall calculation accuracy and availability will be greatly impacted, making it difficult to meet the stringent requirements of actual production scenarios.

[0006] Therefore, the present invention provides a method for real-time obtaining of industry investment return rate based on multi-source heterogeneous data. Summary of the Invention

[0007] (1) Technical Problems to be Solved

[0008] Aiming at the deficiencies of the prior art, the present invention provides a method for real-time obtaining of industry investment return rate based on multi-source heterogeneous data, based on the UTC reference and adjustable gain, power exponential normalization of multi-source timestamps and retaining metadata information; in stream processing, by setting buffer queues and watermark thresholds, reordering of out-of-order data and marking of expired data are realized; when detecting null fields or cross-source conflicts, interpolation and conflict handling are performed to output complete data corrected for consistency; subsequently, the timestamps are further fine-tuned at the millisecond level using an external high-precision reference or cross-correlation function, and if the adjustment amplitude exceeds the safety boundary, it is marked as suspicious; finally, the corrected data is distributed to multiple nodes according to the shard mapping strategy, and fault tolerance consensus and difference repair are triggered under the determination of the fault duration, so as to maintain high-precision time series alignment and high availability in a massive concurrency scenario, avoid data loss or accuracy decay caused by time series inconsistency and node failures, and solve the technical problems proposed in the background art.

[0009] (2) Technical Solutions

[0010] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for real-time obtaining of industry investment return rate based on multi-source heterogeneous data, including,

[0011] After receiving multi-source heterogeneous data, based on the unique identification ID assigned to each data source i, and its local or server timestamp information, perform preliminary normalization conversion on time zone and precision with the help of UTC benchmark, adjustable gain coefficient and power exponent, and mark the converted time as the standard timestamp Meanwhile, supplement necessary metadata information;

[0012] When the data marked with the standard timestamp formally enters the streaming processing engine, create a buffer queue for it, judge the tolerance delay range according to the watermark threshold, perform chronological reordering and expiration marking on the out-of-order records within a given window, and mark the completed synchronization batch with a batch mark;

[0013] When null fields or cross-source data conflicts are found in the sorted batch, perform missing value calculation on the missing part according to the interpolation method, reduce the weight of abnormal records based on the conflict rules, output the complete chronological data after consistency correction, and retain the interpolation mark for reference during subsequent timestamp fine-tuning;

[0014] After receiving the interpolated data without serious contradictions, use an external high-precision reference signal or the cross-correlation function of the internal trigger sequence, combine adjustable parameters to fine-tune the second-level or millisecond-level offset and record the adjustment amount. If the correction amplitude exceeds the safety boundary, mark it as fine-tuned but suspicious, and output the final chronological data with finer granularity, which can meet the accurate time series requirements of high-frequency or low-tolerance scenarios;

[0015] After completing the refined timestamp alignment, distribute the corrected data to multiple nodes according to the sharding mapping strategy, and maintain the sharding status for each node. Once a node failure or abnormal load is detected and exceeds the fault judgment threshold, trigger the fault tolerance consensus mechanism, perform sharding migration and difference repair, and continuously maintain the high-availability real-time data output after alignment in a multi-node environment.

[0016] Preferably, for each data source S i allocate a unique identification ID i , and establish metadata M i in the database or cache, including:

[0017] The geographical location or time zone label of the data source, the description of the time precision of the data reported by the data source. If the source does not have an accurate timestamp locally, the local time needs to be noted here; store the above information in the index table in the form of <ID i ,M i >;

[0018] Preferably, according to the registered <ID i ,M i > information, for the data source that normally reports local time Perform the most preliminary conversion according to the time zone difference or precision hint, map it to a unified reference axis for benchmarking, and obtain a preliminary standard timestamp And replace the local timestamp or placeholder timestamp in the original record with it;

[0019] Preferably, the data record that has completed benchmarking will have an updated standard timestamp Put <ID i , Df i ) into the index record table in sequence, where Df i is a reference pointing to a specific data record; if several data from the same source or multiple sources are detected under the same standard timestamp, add a batch mark to the index item;

[0020] Preferably, for each record <ID i , Df i > that has completed the first step of processing, establish a unified streaming access point, temporarily store it in a data structure called a buffer queue, and divide it into a single queue mode or a multi-queue mode according to business requirements;

[0021] Set a capacity limit Θ cap for the buffer queue, and this capacity limit Θ cap can be determined according to the available memory or the estimated traffic;

[0022] Introduce an adaptive control function Γ buff (·), and automatically make decisions on expansion or contraction according to the deviation between the current queue length l(t) and the expected capacity during operation. Its form can be defined as:

[0023] Γ buff (l(t)) = μ(l(t) - l opt ) r + δ

[0024] In the formula: l(t) represents the number of records in the queue at time t, l opt represents the capacity that the queue is expected to maintain under ideal load, μ and δ are tuning parameters, and r is the power exponent;

[0025] Preferably, in the real-time stream processing framework, define a watermark function Ω wm (t) to represent the estimate of the maximum event time of the data that has arrived, and expand it into a form including a non-linear delay margin Ω wm (t):

[0026]

[0027] In the formula: Represents the maximum value of the standard timestamps of all records in the current buffer queue. BQ is the buffer queue, and Φ(ε, t) is a custom delay margin function, whose inputs are the current time t and the allowed delay error ε, in order to reserve a certain waiting space for high-delay data;

[0028] Φ(ε, t) = ε(1 + κln(1 + t - t0))

[0029] Where κ is an adjustable gain coefficient, and t0 represents the start or a certain reference time; the watermark threshold δ can have a higher tolerance for delay fluctuations in the later stage of operation, to avoid some data being prematurely determined as late due to accidental delays;

[0030] When obtaining a data record (ID, Df i ), its standard timestamp is compared in real time with the current watermark Ω wm (t);

[0031] If then this record is regarded as late data, a delay mark is added to it, and it is marked in the metadata M i ; if then it is regarded as arriving normally and no late processing is done;

[0032] The output includes the updated queue status, including the delay marks of the records, the watermark value Ω wm (t), etc.;

[0033] Preferably, according to the <ID, Df i , delay mark> information in the queue, all normal records are sorted in ascending order according to the standard timestamp and packaged into logical batch units for unified output:

[0034]

[0035] After sorting is completed and a certain batch output condition is reached, missing data interpolation and consistency correction are performed on this batch;

[0036] For records with delay mark = true, the following two strategies can be adopted:

[0037] Delayed output strategy: It is temporarily stored in the late area and continues to wait for a certain period of time, and then it is judged whether it arrives within the next watermark cycle and re-sorted. If it still cannot catch up with the next batch within the delay period, an expired mark is added to the metadata;

[0038] Directly marked as discarded: If such late data is determined to exceed the tolerable time window Ω wm (t) - Δ tolIf it takes too long, add a discarded flag to it and do not send it to subsequent steps to save computing resources, but a most concise piece of information can be retained, such as <ID i , >, for subsequent log analysis or data traceability;

[0039] After reordering is completed and batch units are generated each time, trigger an internal event to inform subsequent links that operations such as missing value imputation can be performed; record the late-arriving or discarded data in the stream monitoring log as well;

[0040] Preferably, retrieve each record from the batch data batch unit or the postponed waiting area; for data in a discarded or expired state, only retain the minimum reference information: for normal or postponed output data, check each record one by one to see if the required key fields exist;

[0041] If it is found that a certain record <ID i , Df i ) lacks a numerical value field or its value is marked as a null value, it is considered imputable for missing values. If a data source has no data arrival at all within a certain time period, it will also be classified into the category of imputable for missing values; according to the data source metadata M i , attach this category of information to the record, and the missing values can be divided into: critical missing values and general missing values;

[0042] Preferably, for data with general missing values, use low-cost methods for rapid filling. This includes forward value extension and interval interpolation; for fields with critical missing values, machine learning or rule algorithms can be introduced to predict the missing values with high precision:

[0043] If external features can be obtained, they can be incorporated into the kernel function or combined with a regression model to further improve the prediction accuracy; after imputation, add a filled flag to the corresponding record and write the specific imputation method into the metadata; if the imputation fails, add an unfilled flag and retain the null value in the result set;

[0044] If a certain record has both a delay flag and a null value field, it is preferably regarded as missing data and can enter the imputation algorithm; if the record is lost, no imputation process will be triggered anymore. In this way, a priority table can be established for all flags, indicating their compatible or mutually exclusive relationships to avoid conflicts;

[0045] Preferably, for values from different data sources within the same time window or adjacent time windows, if there are logical conflicts, they should be identified and adjusted;

[0046] Based on the constraints or association rules predefined in the metadata M i , check each imputed or original value record. If a conflict is found, mark it in the following ways: slightly inconsistent and severely inconsistent;

[0047] If a record is marked as seriously inconsistent, a penalty coefficient η ∈ (0, 1) can be applied to its value in subsequent operations to form an adjusted result

[0048]

[0049] where γ is a safety benchmark value;

[0050] For slightly inconsistent records, only monitoring or recording can be done, without immediately reducing the weight. The final validity or corrected value of each record will be updated, and a conflict mark or adjustment information will be attached to the metadata;

[0051] Preferably, in combination with the metadata M i The recorded time precision information is used to detect whether the actual acquisition precision of the data source S i is lower than the requirement. If a data source is marked as millisecond precision, but it requires microsecond-level alignment in actual applications and the precision does not meet the requirement, then a precision warning or a low-precision mark will be added at this stage;

[0052] If an external synchronization signal is available, the time of this external signal is recorded as T ext (t); It can be compared with its own time at this time If the deviation between the two exceeds the preset deviation threshold Δ sys , a fine calibration algorithm will be triggered at the next moment to provide a basis for subsequent fine-tuning; within the preset deviation threshold Δ sys , no major calibration will be done to save overhead or avoid interference; when the external signal is unavailable or unstable, this link will be skipped and the subsequent correction will be completed only based on the internal correlation analysis;

[0053] Preferably, cross-evaluation is performed on different data sources within the same time window or triggered by the same event. The correlation degree is calculated by using the labeled data, and a metric function γ(τ; S i , S j ) with an integral kernel is constructed. Let W be the time window for alignment evaluation, then it is defined as:

[0054]

[0055] In the formula: and are unrefined timestamps, K(·) is a kernel function used to characterize the contribution of timestamp differences to the alignment degree; ω(t) is a weight function;

[0056] If a high-precision reference is available, first perform an overall correction on the current clock according to the external signal, and then perform the above cross-correlation calculation on the local drift between different data sources; if the external signal is missing or unstable, directly estimate the offset between sources based on the metric function γ;

[0057] Store the calculation results in the deviation mapping table (offset mapping table), and record the optimal lag of each pair of data sources (S i ,S j ); Denote the deviation related to the external reference as According to all the optimal lags in the offset mapping table Find the synchronization order that can globally align the vast majority of sources most closely;

[0058] Determine a synchronization path through the minimum spanning tree or minimum cycle cover method, and fine-tune the timestamps of each source in sequence to avoid creating new conflicts caused by changing all sources simultaneously at once;

[0059] Output the deviation mapping table (offset mapping table) and the synchronization path or priority list;

[0060] Preferably, according to the synchronization path, correct the timestamps of each data source S i according to its optimal deviation relative to the high-precision reference and adjacent sources Let the standard timestamp represent the preliminary standard timestamp of the current source S i at time t, and construct the following refined model in the form of integral correction

[0061]

[0062] where: Err i (s) represents the time deviation function measured for source S i at time s; Φ i (Err i (s), s) is the correction kernel function, defined in the form of multi-variable input and non-linear output:

[0063] Φ i (x, s) = δ i (s) exp(-γ|x| v ), γ > 0, v ≥ 1

[0064] where: δ i (s) can be regarded as a time-varying weight, and Λ i is the overall scaling factor;

[0065] Set a safety margin for each record. If the correction amount exceeds this margin, manual intervention is required or excessive adjustments are recorded separately. After completing the fine-tuning of all data sources, indicate in the record metadata that the timestamp has been corrected and record the new timestamp As the final result, provide the adjustment log when outputting the result data;

[0066] Preferably, according to the time alignment result, use the sharding mapping table II (shard) Perform sharding according to the business dimension or time period and allocate to different computing nodes {N1, N2, …, N m};

[0067] Maintain the time offset function Δ n (t) locally, where n ∈ {1, …, m} represents the node number and t represents the actual processing time; To ensure that the time offset function Δ n (t) is adaptive to network fluctuations or node clock errors, adopt the following differentiable kernel integral form:

[0068]

[0069] where: κ n is the scaling coefficient of node N n ; Ω is the sliding window size; φ n (u) represents the transient clock deviation of node N n from the high-precision reference at time u; is the kernel function;

[0070] When any node detects that the load is approaching the upper limit during the communication peak period, it sends a sharding migration request to the central scheduler or other nodes, triggering the fault-tolerant consensus process;

[0071] Preferably, when any node Node x experiences an irrecoverable failure or a high load exceeding the tolerance threshold, detect the abnormality through the sharding status record associated with node Node x and send a failure alarm to other available nodes and the scheduling center;

[0072] After receiving the failure alarm, other nodes trigger the fault-tolerant consensus process using the broadcast time offset function Δ n (t) and the sharding mapping information to determine the optimal sharding takeover plan;

[0073] If node Node x recovers within a short time, set its sharding status to temporarily offline, and other nodes save a small amount of key incremental data for it first;

[0074] If it exceeds the fault judgment threshold ε failIf the duration is not restored, its shards are officially migrated to the available node Node y , by the node Node y Based on the time offset function Δ y (t) and the external calibration reference to perform time alignment and status continuation on the migrated data;

[0075] After the data is re-grouped, differential patching is performed on the trace data lost or missed during the fault period.

[0076] (III) Beneficial effects

[0077] The present invention provides a method for real-time obtaining of the industry return on investment based on multi-source heterogeneous data, having the following beneficial effects:

[0078] By adaptively fusing multi-source heterogeneous data, stream buffering and watermarking, missing value imputation and conflict verification, high-precision timestamp refinement, and distributed fault tolerance, the real-time consistency, integrity, and availability of multi-source heterogeneous data in the second-level or even millisecond-level environment are significantly improved.

[0079] By adopting the unified benchmark UTC for the local timestamps of the data source S i and cooperating with the adjustable gain coefficient κ and the power exponent p for time zone difference and precision scaling, the clock deviations of each source can be quickly eliminated at the initial access, so that each record has a standard time stamp With the help of the buffer queue and the watermark threshold δ strategy in stream processing, the correct event timing can still be maintained in the scenarios of out-of-order data and frequent delays, and the data exceeding the delay tolerance is marked as late or discarded with a delay mark, ensuring more reliable subsequent calculations. In the imputation and cross-source verification stage, general missing values and critical missing values are distinguished, and the obvious abnormal or untrustworthy information is moderately de-weighted in combination with the conflict record penalty factor η, which not only ensures the high integrity of the data at the numerical level, but also lays a clean foundation for the next timestamp refinement.

[0080] External high-precision signals and related deviation amounts are introduced to perform microsecond-level correction on the timestamps of each source. When the correction amplitude exceeds the safety boundary, it is marked to achieve a balance between precise alignment and robustness. In the distributed fault tolerance stage, through the node shard status record and the fault duration threshold, automatic migration and asynchronous takeover of data shards are realized, supplemented by the kernel integral offset function to maintain cross-node clock synchronization, and the continuity and high availability of the global alignment result can still be ensured when the nodes are overloaded or faulty.

[0081] Unify the time zone and precision at the beginning of multi-source access, mitigate the risk of out-of-order delay through streaming watermarks and buffering, strengthen the internal data consistency through targeted interpolation and conflict handling, consolidate high-precision alignment with refined timestamp fine-tuning, and cooperate with multi-node fault tolerance and sharded distribution to stabilize the data stream persistence in high-volume concurrent scenarios. Thereby, significantly reduce the risk of data loss or precision attenuation caused by time series inconsistency and node failures in high-frequency environments, effectively improve the accuracy and robustness of real-time analysis of multi-source heterogeneous data, and provide a more reliable time series basis and distributed parallel guarantee for industry decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 It is a schematic flow chart of the method for real-time obtaining of the industry investment return rate of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0084] Please refer to Figure 1 , the present invention provides a method for real-time obtaining of the industry investment return rate based on multi-source heterogeneous data, including: Step 1, after receiving multi-source heterogeneous data, based on the unique identifier ID assigned to each data source i , and its local or server timestamp information, preliminarily normalize and convert the time zone and precision with the UTC reference and adjustable gain coefficient and power exponent, and mark the converted time as the standard timestamp At the same time, supplement the necessary metadata information;

[0085] The above Step 1 includes the following contents:

[0086] Step 101, Data source identification and metadata injection

[0087] For each data source S i Assign a unique identifier ID i , where i represents the data source number, and the value range is {1, 2,..., N}, and N is the number of data sources; establish a metadata structure M in the database or cache i , and its content includes:

[0088] The geographical location or time zone label of the data source (recorded if known, marked as unknown if unknown), the time precision description of the data reported by the data source (such as seconds, milliseconds or microseconds, etc.), if the source has no accurate timestamp locally, it is necessary to indicate the local time here; record the above information with <IDi , M i in the form of < is stored in the index table, and all subsequent operations on the same data source can be retrieved quickly through the ID i to retrieve the corresponding metadata M i ;

[0089] By assigning a unique identifier ID i to each data source S i , and establishing a metadata structure M i , data can be accurately distinguished and key information such as time zone and precision can be recorded at the initial stage; special marker processing for local time enables data sources with missing or ambiguous timestamps to be retained for subsequent correction instead of being discarded rashly; the method of prior metadata injection can help the system flexibly identify data types and time characteristics in subsequent steps, improving processing efficiency and error localization ability.

[0090] Step 102, Local Time Parsing and Preliminary Benchmarking

[0091] When receiving a new data record, according to the registered <ID i , M i information, read the precision, time zone or marker status of the data source S i . If it is marked as local time, replace the local time with the arrival time of the receiving server;

[0092] For data reporting local time normally , perform the most preliminary conversion according to the time zone difference or precision hint, and map it to a unified reference axis; for example, first convert the local time to an intermediate time T A ; On this basis, use the following formula for benchmarking to obtain a preliminary standard timestamp

[0093]

[0094] where: T B is the UTC reference defined in advance by the system;

[0095] α i is the scaling factor for the data source S i , used to comprehensively consider the precision (seconds, milliseconds, etc.) and time zone deviation of the source. Its value range can be determined according to business requirements, for example, (0, 2]; p is an adjustable power exponent, and its value can be between 1 and 2. In ordinary scenarios, p = 1 can be taken; if non-linear mapping of clock offset such as acceleration / deceleration is required, then p > 1 can be set; ω i is the compensation term for server time drift, which can be set to 0 if there is no such requirement;

[0096] Finally, a standard timestamp is obtained And replace the local timestamp or placeholder timestamp in the original record to achieve unified benchmarking;

[0097] Compared with the traditional method of adding or subtracting time zones + precision conversion, the above formula introduces the power exponent p and the compensation amount ω i , and sets an independent scaling factor α for each data source i Therefore, it can flexibly handle multiple precisions or multiple potential offset modes, improving the adaptability of time alignment. If some data sources need special treatment (for example, there is obvious drift in geographically remote monitoring devices), it can be adjusted by separately adjusting its α i or ω i , without interfering with other data sources;

[0098] Step 103, Timing index and output connection

[0099] All data records that have completed benchmarking will have updated standard timestamps Put <ID i , Df i ) into the index record table in sequence, where Df i is a reference pointing to the specific data record (including the data content itself and its source information);

[0100] This index record table will sort and store the standard timestamps of different data sources. If it is detected that there are several data from the same source or multiple sources under the same standard timestamp, a batch mark can be added to the index item;

[0101] Use the index record table to index and store <ID i , Df i ), ensuring that the order of all data is consistent during query and subsequent processing. Setting batch marks can identify multiple data arriving at the same moment, facilitating batch processing and bandwidth optimization. Pass all metadata information and correction results to the next step together, so that the system does not need to repeat parsing the details of the first step later, saving computing resources.

[0102] Step 2. When the data marked with the standard timestamp formally enters the streaming processing engine, create a buffer queue for it and judge the tolerance delay range according to the watermark threshold, reorder the out-of-order records in time sequence and mark the expiration within the given window, and mark the sorted synchronization batch with a batch mark;

[0103] The above Step 2 includes the following contents:

[0104] Step 201, Buffer window initialization and dynamic capacity management

[0105] For each record that has completed the first - step processing <ID i , Df i > establish a unified streaming access point and temporarily store it in a data structure named buffer queue. According to business requirements, it can be divided into a single - queue mode to uniformly process all data sources; or a multi - queue mode, where separate buffers are established for different IDs i At the implementation level, a circular queue or a deque can be used to improve concurrency efficiency;

[0106] To prevent queue congestion caused by peak data - source periods, a capacity upper limit Θ needs to be set for the buffer queue cap , and this capacity upper limit Θ cap can be determined according to available memory or estimated traffic. For example, Θ cap can take values in the range of [10 4 , 10 6 records; where: if the peak concurrent volume of the data source is X, Θ can be initially set as cap ≈2×X;

[0107] Introduce an adaptive control function Γ buff (·) to automatically make expansion or contraction decisions according to the deviation between the current queue length l(t) and the desired capacity during operation. Its form can be defined as:

[0108] Γ buff (l(t))=μ(l(t)-l opt ) r +δ

[0109] In the formula: l(t) represents the number of records in the queue at time t, l opt represents the capacity that the queue is expected to maintain under ideal load. μ and δ are tuning parameters used to control the speed of expansion (positive value) or contraction (negative value): r is the power exponent. If r > 1, stronger penalties can be imposed on situations exceeding the target capacity to accelerate expansion or restriction;

[0110] Through the adaptive control function Γ buff (·), buffer resources can be automatically increased during peak traffic (if hardware conditions permit), and redundant memory can be reclaimed during traffic decline to maintain good real - time performance and resource utilization;

[0111] By setting the buffer queue and its capacity upper limit Θ cap, to prevent data loss in high-concurrency scenarios, ensure real-time performance and integrity, use a unified or source-separated buffer queue to ensure that data will not be lost due to out-of-order or instantaneous congestion before being processed, thereby maintaining data integrity. Dynamic capacity management enables the system to automatically adjust the queue size in a flexible resource environment, avoiding frequent overflows or performance waste caused by fixed-size buffers. Through the adaptive control function Γ buff (l(t)) function, adaptively regulate the queue load, rather than a simple linear threshold strategy, so as to have more flexible control capabilities in high-concurrency or extreme scenarios.

[0112] Step 202, Watermark Generation and Late Arrival Detection

[0113] In a real-time stream processing framework (such as Apache Flink, Spark Streaming), define the watermark function Ω wm (t) to represent the estimate of the maximum event time of the arrived data, and expand it to the form including the non-linear delay margin Ω wm (t):

[0114]

[0115] In the formula: represents the maximum value of the standard timestamps of all records in the current buffer queue, used to represent the time upper limit of the latest event. BQ is the buffer queue, and Φ(ε, t) is a custom delay margin function, whose inputs are the current time t and the allowed delay error ε, in order to reserve a certain waiting space for high-latency data. The specific form of the custom delay margin function Φ(ε, t) can adopt piecewise exponential or polynomial to appropriately relax the watermark trigger threshold during extreme traffic fluctuations. For example:

[0116] Φ(ε, t) = ε(1 + κln(1 + t - t0))

[0117] Among them, κ is an adjustable gain coefficient, and t0 represents the start or a certain reference time; the watermark threshold δ can have a higher tolerance for delay fluctuations in the later stage of operation, avoiding some data being prematurely determined as late arrivals due to occasional delays;

[0118] When obtaining the data record (ID, Df i ), compare its standard timestamp with the current watermark Ω wm (t) in real time;

[0119] If then this record is regarded as late-arriving data, add a delay mark to it, and mark it in the metadata M i ; if then it is regarded as a normal arrival and no late arrival processing is performed;

[0120] Records marked with a delay are still retained in the buffer queue for a short period of time, waiting for the downstream to decide whether to discard them or perform special processing; the output includes the updated queue status, including the delay marks of the records and the watermark value Ω wm (t), etc.;

[0121] Watermark function Ω wm (t) combined with the custom time delay margin Φ(ε, t) can dynamically adapt to network congestion or data source delay fluctuations, without overly discarding late records, marking late data with delay marks for subsequent steps to determine whether to continue to retain or mark;

[0122] The watermark mechanism enables real-time stream processing to have a certain degree of tolerance for out-of-order and delayed data, avoiding overall calculation errors or premature output due to the time delay of a single source. Combining late arrival determination with the adaptive time delay margin function Φ enables the system to adapt to various network conditions and data source differences, providing sufficient time for subsequent interpolation and correction; upgrading the common fixed watermark subtraction to Φ in the form of a logarithm or polynomial maintains high efficiency in stable scenarios and provides higher tolerance in high-delay abnormal situations.

[0123] Step 203, Time Series Reordering and Calculation Trigger

[0124] Each data record already has a normal or late mark. Now, according to the <ID, Df i , delay mark> information in the queue, all normal records are sorted in ascending order according to the standard timestamp and packed into logical batch batch units for unified output:

[0125]

[0126] After sorting is completed and a certain batch output condition is met (for example, the number of records reaches a predefined threshold or the watermark trigger point is sensed), missing data interpolation and consistency correction are performed on this batch;

[0127] For records with delay mark = true, the following two strategies can be adopted:

[0128] Postponed output strategy: Temporarily store it in the late arrival area and continue to wait for a certain period of time, then determine whether it arrives within the next watermark cycle and perform reordering. If it still cannot catch up with the next batch within the delay period, an expired mark is added to the metadata;

[0129] Directly mark and discard: If such late data is determined to exceed the tolerable time window Ω wm (t) - Δ tolIf it takes too long, add a discarded flag to it and do not send it to subsequent steps to save computing resources, but a copy of the most concise information can be retained, such as <ID i , >, for subsequent log analysis or data traceability;

[0130] After each reordering and generation of batch units is completed, trigger an internal event to inform subsequent processes that operations such as missing value imputation can be performed; record late or discarded data in the flow monitoring log as well;

[0131] By packing normal records in an orderly manner, it can be ensured that downstream algorithms do not have to separately concern themselves with out-of-order issues during processing, thus simplifying the implementation and improving the accuracy of missing data identification. Implementing hierarchical processing (delaying or directly discarding) of late data can avoid resource waste caused by unlimited waiting while ensuring integrity, and also enable the system to have a controllable adaptive ability to extreme network delays. It can retain some suspicious data for backtracking and avoid resource waste caused by unlimited waiting.

[0132] Add two modes: delayed output strategy and direct marked discard, so that late data can be treated differently according to the degree of delay, thereby flexibly setting the waiting time or tolerance according to different business requirements.

[0133] Step 3: When null value fields or cross-source data conflicts are found in the sorted batch, perform missing value calculations on the missing parts according to the imputation method, and reduce the weights of abnormal records based on the conflict rules, output the complete time-series data after consistency correction, and retain the imputation mark for reference during subsequent timestamp fine-tuning;

[0134] The said Step 3 includes the following contents:

[0135] Step 301: Missing record identification and classification

[0136] Retrieve each record from the batch data batch unit or the delayed waiting area; for data in the discarded or expired state, only retain the minimum reference information; for normal or delayed output data, check one by one whether the required keyword fields (such as observed values, text payloads, etc.) exist;

[0137] If it is found that a certain record <ID i , Df i ) lacks numerical fields or its value is marked as null, it is considered imputable missing. If a certain data source has no data arrival within a certain time period, it will also be classified into the category of imputable missing;

[0138] According to the data source metadata M iFor information such as importance, data type, time-frequency characteristics, etc., attach this category of information to the record. The missing values can be classified into: Critical missing values: such as core financial indicators, or keyword fields that have a significant impact on subsequent calculations; General missing values: relatively less important fields, such as optional additional attributes.

[0139] Through systematic scanning and classification, it is possible to quickly identify which missing values need to be prioritized for imputation and which can tolerate null values, improving the imputation efficiency and reducing unnecessary computational overhead for less important fields. For different categories, different strategies can be adopted during subsequent imputation, which is also convenient for priority management in case of imputation failure; Combine the missing value categories with the data source metadata M i to achieve differential processing based on data importance or impact degree in the same missing value situation, which is much more flexible and precise than the unified perspective of imputing whenever there is a missing value.

[0140] Step 302, Multi-strategy Imputation and External Feature Prediction

[0141] For data with general missing values, use low-cost methods for rapid filling, including:

[0142] Previous value extension: directly take the value in the previous non-null record of the same data source as the current filling value;

[0143] Interval interpolation: If there are valid values at adjacent times t1 and t2, then obtain the value at the missing time t according to a linear ratio:

[0144] X i (t) = X i (t1) + Λ[X i (t2) - X i (t1)]

[0145] where Λ is For fields with critical missing values, machine learning or rule algorithms can be introduced to predict the missing values with high precision: The following kernel regression imputation formula is as follows:

[0146]

[0147] In the formula: t m represents the time to be imputed; represents the data source S i the set of known sample indices within adjacent time periods; κ(·) is the kernel function, used to measure that the weights of historical observations closer in distance are larger;

[0148] ρ is the distance scaling coefficient, used to control the decay rate of the kernel function;

[0149] If external features (such as market fluctuations and upstream and downstream indicators of the industrial chain) are available, they can be incorporated into the kernel function or combined with a regression model to further improve the prediction accuracy; after imputation, a filled marker is added to the corresponding record, and the specific imputation method is written into the metadata; if the imputation fails (such as insufficient historical data), an unfilled marker is added and a null value is retained in the result set;

[0150] Applying low-cost simple imputation and machine learning-driven high-order imputation to different importance fields separately can not only balance computational efficiency but also improve the accuracy of key variable imputation. Introducing the kernel function κ and the scaling coefficient ρ brings stronger adaptability to the imputation process and can better reflect the time series pattern and similarity weight compared to traditional mean or linear methods;

[0151] Adopting the form of kernel regression in predictive imputation and allowing external features to be incorporated into the construction of the kernel function breaks through the limitations of conventional difference or moving average and is suitable for handling complex data missing scenarios with multiple sources and dimensions.

[0152] If a record has both a delay marker and a null value field, it is preferably regarded as missing data and can enter the imputation algorithm; if the record is lost, no imputation process will be triggered. In this way, a priority table can be established for all markers to indicate their compatibility or mutual exclusion relationships and avoid conflicts;

[0153] Step 303, Cross-source Consistency Check and Data Validity Adjustment

[0154] For values from different data sources within the same time window or adjacent time windows, if there are logical conflicts, they should be identified and adjusted. For example: the product relationship between trading volume and trading amount should be within a certain fluctuation range; if there is an extreme deviation between the upstream raw material price and the downstream production volume, it is marked as slightly inconsistent;

[0155] Based on the constraints or association rules predefined in the metadata M i check the imputed or original records. If a conflict is found, it is marked as follows:

[0156] Slightly inconsistent: Minor conflicts that can be ignored; Seriously inconsistent: Serious contradictions that need to reduce their weight or be specially processed in subsequent analysis;

[0157] If a record is marked as seriously inconsistent, its value can be penalized with a penalty coefficient η ∈ (0, 1) in subsequent operations to form an adjusted result

[0158]

[0159] Among them, γ is a safety benchmark value (such as the weighted average of adjacent moments or the balance value set by the service). The smaller η is, the more serious the conflict is, and the lower the trust level of the record;

[0160] For minor inconsistencies, only monitoring or recording is required, and the weight does not need to be immediately reduced. The final validity or correction value of each record will be updated, and a conflict mark or adjustment information will be attached to the metadata;

[0161] Cross-source consistency verification can quickly identify the unreasonable points between data from different sources, providing higher data credibility for data analysis and return on investment prediction. By assigning a penalty coefficient to severely inconsistent records, while retaining the original data, the impact on subsequent calculations can be reduced to a reasonable range without completely discarding the suspicious data; when severe conflicts are found, instead of simply discarding the relevant records, a penalty coefficient is used to reduce its impact in subsequent calculations: by performing weighted fusion with the safety benchmark value γ, the most basic time series continuity is retained.

[0162] Step 4: After receiving the imputed data without severe contradictions, use the cross-correlation function of an external high-precision reference signal or an internal trigger sequence, and combine adjustable parameters to fine-tune the second-level or millisecond-level offset and record the adjustment amount. If the correction amplitude exceeds the safety boundary, it is marked as fine-tuned but suspicious, and a precise time series with a finer final time series that can meet high-frequency or low-tolerance scenarios is output;

[0163] The content of the above-mentioned step 4 includes the following:

[0164] Step 401: Microscopic precision difference identification and external benchmark acquisition

[0165] The refinement of the timestamp is only based on cross-source comparison and external benchmarks, and the numerical fields obtained by the third-step imputation are not modified again. Even if the imputation result affects the inference of the event trigger order, the previously imputed values will not be reversely changed;

[0166] Combined with the metadata M i The recorded time precision information is used to detect whether there is a data source S i whose actual acquisition precision is lower than the requirement. If a certain data source is marked with millisecond precision, but it requires microsecond-level alignment in actual applications and the precision does not meet the requirement, then a precision warning or a low-precision mark is added at this stage;

[0167] If there is an external synchronization signal, such as a GPS clock timing server or a high-precision benchmark of a high-frequency trading platform, the time of this external signal is recorded as T ext (t); At this time, it can be compared with its own time If the deviation between the two exceeds the preset deviation threshold Δ sys, the fine calibration algorithm will be triggered at the next moment to provide a basis for subsequent fine-tuning; within the preset deviation threshold Δ sys If it is within the range, no major calibration will be performed to save overhead or avoid interference; when the external signal is unavailable or unstable, this link will be skipped and subsequent calibration will be completed only based on internal correlation analysis; record the low-precision mark and the high-precision benchmark of the available external time signal in the metadata.

[0168] It can identify in advance which data sources have low time accuracy or obvious uncertainties, so that such sources can be focused on during subsequent calibration. If there is an available external high-precision benchmark, subsequent adjustments will be more accurate and authoritative; otherwise, internal cross-correlation or trigger order can be used for estimation to enhance system adaptability. Combine the availability of the external high-precision signal with the marked precision status, make differential processing in advance, and provide multiple branch options for the subsequent fine-tuning logic, rather than relying on a single alignment method forcefully.

[0169] Step 402, Deviation Estimation and Synchronization Order Inference

[0170] Perform cross-evaluation on different data sources within the same time window or triggered by the same event, and calculate the correlation degree by means of the marked data (such as the key variables after interpolation, the cross-source consistency results). For example, if it is known that multiple sources record the same market event within a certain second, then their recorded standard timestamps should be close or equal. Next, construct a metric function γ(τ; S i , S j ) with an integral kernel;

[0171] Within a specific time window, perform an integral operation on the time difference relationship between source S i and source S i , and measure the similarity with the kernel function, so as to take into account both globality and differentiability; let W be the time window for alignment evaluation (such as the joint sample interval of the last few seconds or minutes), then it can be defined as:

[0172]

[0173] In the formula: and are unrefined timestamps (or consistent timestamps after preliminary interpolation); K(·) is the kernel function used to characterize the contribution of the timestamp difference to the alignment degree, and a Gaussian kernel, an exponential kernel or a custom fast-decaying kernel function can be selected. For example:

[0174] K(x) = exp(-α|x| β ), α > 0, β ≥ 1

[0175] Here, α and β can be set in combination with the data characteristics to control the attenuation speed and the degree of nonlinearity;

[0176] ω(t) is an optional weight function that allows different importance levels to be assigned to different moments within a time window. For example, a larger weight can be set during peak trading hours to increase the alignment priority during this period. If this requirement does not exist, ω(t) can be set to 1;

[0177] By integrating over the entire time interval, the metric function γ(τ; S i , S j ) will accumulate higher values when the time stamp difference between the two is small, thus reflecting a better fit. Maximizing this integral can find the optimal τ for estimating the source S i and S j The best alignment offset between them globally. If a high-precision reference is available, first perform an overall calibration of the current clock based on the external signal, and then perform the above cross-correlation calculation on the local drift between different data sources; if the external signal is missing or unstable, directly estimate the offset between the sources based on the metric function γ(τ; S i , S j );

[0178] Store the calculation results in the deviation mapping table offset mapping table, recording the optimal lag of each pair of data sources (S i , S j ); Record the deviation related to the external reference as According to all the optimal lags in the offset mapping table Use graph theory or sorting algorithms to find the synchronization order that can globally align the vast majority of sources most closely. For example, a complete graph G can be constructed, with its nodes corresponding to data sources, and the edge weights can be set to the relative positions;

[0179] Determine a synchronization path through the minimum spanning tree or minimum cycle cover method, and fine-tune the time stamps of each source in sequence to avoid creating new conflicts caused by simultaneously changing all sources at once;

[0180] Output the deviation mapping table offset mapping table and the synchronization path or priority list; finely evaluate the time misalignment between different sources in high-frequency or complex data scenarios. When cooperating with an external high-precision reference, global calibration + local alignment can be performed simultaneously. Using graph algorithms or topological sorting to infer the synchronization order can avoid the chaos caused by circular adjustment or multi-source mutual restraint, and improve the reliability of a single alignment; for multi-source scenarios, a comprehensive scheme of cross-correlation function + graph theory synchronization order is proposed, getting rid of the limitations of only using a fixed reference source or manually setting priorities, and having both flexibility and global optimality.

[0181] Step 403, Fine-grained time stamp correction and output

[0182] According to the synchronization path, for each data source Si The timestamps are corrected according to their optimal deviation amounts relative to the high-precision reference and adjacent sources as follows: Let the standard timestamp represent the preliminary standard timestamp of the current source S i at time t, and construct the following refined model in the form of integral correction

[0183]

[0184] where: Err i (s) represents the time deviation function measured for source S i at time s, which may come from the aforementioned cross-source alignment analysis or external reference signals, for example:

[0185]

[0186] is a set of multiple deviation factors, characterizing the degree of global and local inconsistencies;

[0187] Φ i (Err i (s), s) is the correction kernel function used to perform cumulative correction according to the deviation conditions at different times. The kernel idea similar to that used in γ(·) before can be selected, or a multi-variable input and non-linear output form can be defined, such as:

[0188] Φ i (x, s) = δ i (s) exp(-γ|x| v ), γ > 0, v ≥ 1

[0189] where: δ i (s) can be regarded as a time-varying weight (similar to ω(t)), or the data integrity, service priority, etc. of source S i at time s can be additionally considered; Λ i is the overall scaling coefficient, determining the influence strength of the integral term on the final timestamp. If Λ i has a large value, it indicates a greater willingness to significantly correct the historical deviation accumulation; if the value is small, more of the preliminary timestamp will be retained;

[0190] To avoid overcorrection during the fine-tuning process, a safety boundary can be set for each record. If the correction amount exceeds this boundary, manual intervention or additional recording of over-adjustment is required to prevent accidental abnormal data sources from distorting the overall alignment result;

[0191] After completing the fine-tuning of all data sources, indicate in the record metadata that the timestamp has been corrected, and use the new timestamp As the final result; if any special marks (such as slight inconsistencies) appear in the previous interpolation or consistency check, the final weight adjustment can be made with reference to them at this time; when outputting the result data, provide the adjustment log along with it;

[0192] When the fine-tuning amplitude of a certain record is greater than the safety threshold, an event of over-adjustment is automatically issued, and a pause / resume option is provided; Pause: Only review this record separately or retain the original value; Resume: Treat this record as a "potential anomaly", which can be marked in the final result; Mark it as fine-tuned but suspicious, and remind the downstream to refer to its reliability when using. Combined with the actual business, it can be decided whether to automatically correct or terminate the global alignment process after manual confirmation. By gradually integrating the global deviation And the local deviation A more consistent multi-source timestamp at the millisecond or even microsecond level can be obtained. After the correction is completed, the entire data set can be used as a high-precision input for direct invocation by real-time return on investment analysis or other applications with high time sensitivity;

[0193] The external global deviation and the adjacent source local deviation are combined non-linearly instead of simply superimposed, so that the error distribution strategy can be flexibly set for different industries or business scenarios. Introduce a safety boundary check mechanism to prevent over-adjustment and provide more safety redundancy;

[0194] Data at the second or even millisecond level can be aligned to a more refined level, and it remains stable and consistent in multi-source concurrency, network latency, and high-frequency scenarios, providing a solid underlying support for real-time industry return on investment analysis or other calculations that require precise timing guarantee. Through multi-link collaboration, not only the time unification and interpolation correction of multi-source heterogeneous data are realized, but also a stable balance point is found between high precision and security.

[0195] Step Five: After completing the refined timestamp alignment, distribute the corrected data to multiple nodes according to the sharding mapping strategy, and maintain the sharding status for each node. Once a node failure or abnormal load is detected and exceeds the failure judgment threshold, trigger the fault tolerance consensus mechanism, perform shard migration and difference repair, and continuously maintain the high-availability real-time data output after alignment in a multi-node environment.

[0196] The content of the above Step Five includes the following:

[0197] Step 501: Multi-node sharding deployment and clock synchronization broadcast

[0198] At the high-precision timestamp And the corresponding data stream Df i Both have been ensured to be refined and can be used as a globally available data set. According to the time alignment result, use the sharding mapping table II (shard)Slice these data according to business dimensions or time periods and distribute them to different computing nodes {N1, N2, …, N m};

[0199] After receiving the allocated data slices, each node will continuously obtain synchronous calibration information from an external high-precision signal high-precision reference to maintain the time offset function Δ n (t) locally, where n ∈ {1, …, m} represents the node number and t represents the actual processing time;

[0200] To ensure that the time offset function Δ n (t) is adaptive to network fluctuations or the node's own clock error, the following differentiable kernel integral form can be used:

[0201]

[0202] where: κ n is the scaling coefficient of node N n , which is used to determine the intensity of the calibration for correcting the local clock; Ω is the size of the sliding window, which limits the integration interval within the recent Ω length; φ n (u) represents the transient clock deviation between node N n and the high-precision reference at time u; is the kernel function. Combining the exponential or logarithmic decay form can smoothly fuse the drift rate of the node in the recent period, thereby dynamically correcting the node's local clock and maintaining high-precision alignment;

[0203] When any node detects that the load is approaching the upper limit during the communication peak period, it will send a slice migration request to the central scheduler or other nodes, triggering the fault-tolerant consensus process;

[0204] By allocating data to multiple nodes under the slice mapping Ⅱ (shard) framework, the single-point computing pressure is greatly reduced, and the throughput is still maintained under the concurrent massive data streams. At the same time, each node synchronizes with the high-precision reference adaptively through the time offset function Δ n (t), effectively absorbing network delay jitter and node clock drift, and ensuring that the cross-node time scale is still connected to the global refinement result in the fourth step;

[0205] Using the Δ n (t) function combined with kernel integration and sliding window enables the local clock correction of each node to rely on both historical error accumulation and rapid response to the current network state. Compared with traditional fixed offset or simple linear compensation, it has higher flexibility and is suitable for fine-grained distributed collaborative deployment in high-frequency scenarios.

[0206] Step 502, Fault-Tolerant Consensus and Fault Switching

[0207] When any node Node xWhen an irrecoverable failure or a high load exceeding the tolerance threshold occurs, the shard status associated with node Node x detects an anomaly through the shard status record, and sends a failure alarm to other available nodes and the scheduling center;

[0208] After receiving the failure alarm, other nodes trigger a fault-tolerant consensus process by using the broadcast time offset function Δ n (t) and the shard mapping information to determine the optimal shard takeover plan;

[0209] If node Node x recovers within a short time, set its shard status to temporarily offline, and other nodes save a small amount of key incremental data for it first; if it does not recover after exceeding the fault judgment threshold ε fail duration, then officially migrate its shards to the available node Node y , and node Node y performs time alignment and status continuation on the migrated data based on the time offset function Δ y (t) and the external calibration reference;

[0210] After the data rejoins the team (i.e., node Node x comes back online or node Node y completes the takeover), differential patching is performed on the tiny amount of data lost or missed during the fault:

[0211] If the local timestamp of some data records is found through the log then these data records are re-incorporated into the shard and merged into the node alignment queue Queue y ;

[0212] The fault-tolerant consensus module avoids global stagnation or data loss caused by a single node failure, and can still maintain continuous updates to the time series chain during node offline; even if the node rejoins, it can quickly complete smooth docking and missing window compensation based on the previous time offset function Δ n (t) and the information of the high-precision benchmark; the dynamic load switching mechanism enables the system to automatically schedule data shards to idle or low-load nodes, smooth out peaks and valleys and maintain the overall real-time processing efficiency. In the distributed fault-tolerant scenario, deeply integrating the refined timestamp calibration time offset function Δ n (t) with the shard takeover process enables the newly taken-over node to correct the timestamp after a fault switch, truly ensuring the global consistency of multi-node alignment rather than simply doing simple master-backup replication: by performing differential patching on the failure window, the data integrity after fault tolerance recovery is further improved.

[0213] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0214] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0215] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0216] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0217] As described above, only the specific implementation manners of this application are provided, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A real-time acquisition method for the industry return on investment based on multi-source heterogeneous data, characterized in that: Including, After receiving multi-source heterogeneous data, based on the identifiers assigned to each data source and the timestamp information, with the help of adjustable gain coefficients and power exponents, perform preliminary normalization conversion on time zones and precision, and mark the converted time as the standard timestamp; When it is detected that the data marked with the standard timestamp is officially processed in a streaming manner, create a buffer queue for it and determine the tolerance delay range according to the watermark threshold, and perform chronological reordering and expiration marking on the out-of-order records within a given window; When null value fields or cross-source data conflicts are found, perform missing value calculations on the missing parts, and reduce the weights of abnormal records based on the conflict rules, and output the complete chronological data after consistency correction and retain the imputation marks; Use an external high-precision reference signal or the cross-correlation function of the internal trigger sequence, combined with adjustable parameters, to fine-tune the second-level or millisecond-level offset and record the adjustment amount. If the correction amplitude exceeds the expectation, mark it as fine-tuned but suspicious, and output the precise time series; Allocate the corrected data to multiple nodes according to the sharding mapping strategy, and maintain the sharding status for each node. Once it is detected that a node fails or the load is abnormal and exceeds the fault judgment threshold, trigger the fault tolerance consensus mechanism, perform sharding migration and difference patching, and output real-time data.

2. The method for real-time obtaining of the industry investment return rate according to claim 1, wherein: Assign a unique identifier to each data source, and establish metadata in the database or cache, including: the geographical location or time zone label of the data source, the description of the time precision of the data reported by the data source. If the source does not have an accurate timestamp locally, it is necessary to indicate the local time here and store it in the index table; For the normal reporting of local time, perform the most preliminary conversion according to the time zone difference or precision hint, map it to the unified reference axis for benchmarking processing, obtain the preliminary standard timestamp, and replace the local timestamp or placeholder timestamp in the original record with it; The data records that have completed benchmarking will have the updated standard timestamp and be written into the index record table in sequence; if it is detected that there are several pieces of data from the same source or multiple sources under the same standard timestamp, add a batch mark to the index item.

3. The method for real-time obtaining of the industry investment return rate according to claim 1, wherein: Establish a unified streaming access point for each record that has completed the first-step processing, and temporarily store it in a data structure named buffer queue, which is divided into a single queue mode or a multi-queue mode according to business requirements; Set an upper limit on the capacity of the buffer queue, and this upper limit can be determined according to the available memory or the estimated traffic; Introduce an adaptive control function, and automatically make decisions on expansion or contraction according to the deviation between the current queue length and the expected capacity during operation.

4. The method for real-time obtaining of the industry investment return rate according to claim 3, wherein: In the real-time stream processing framework, define a watermark function to represent the estimation of the maximum event time of the arrived data, and expand it into a form that includes a non-linear delay margin, and its input is the current time and the allowed delay error; When acquiring data records, the standard timestamp of the record is compared with the current watermark in real time. If the standard timestamp is less than the current watermark, the record is regarded as late data, a delay mark is added to it, and it is marked in the metadata; on the contrary, it is regarded as arriving normally and no late processing is done, and the output includes the updated queue status.

5. The real-time acquisition method of the industry investment return rate according to claim 4, characterized in that: Now, according to the data record information in the queue, all normal records are sorted in ascending order of the standard timestamp and packaged into a logical batch batch unit; After the sorting is completed and a certain batch output condition is reached, missing data interpolation and consistency correction are performed on this batch; For records with correct delay marks, a delayed output strategy and direct mark discard can be adopted; An internal event is triggered after each re-sorting is completed and a batch unit is generated to inform the subsequent links that operations such as missing interpolation can be performed, and the late-arriving or discarded data is also recorded in the flow monitoring day.

6. The real-time acquisition method of the industry investment return rate according to claim 5, characterized in that: Retrieve each record from the batch data batch unit or the delayed waiting area (delayed area). For data in the discarded or expired state, only the minimum reference information is retained; For data with normal or delayed output, it is detected one by one whether the required keyword fields exist; If it is found that a data record lacks a numerical value field or its value is marked as a null value, it is determined that the missing value can be interpolated. If a data source has no data arrival at a certain time period, it is also classified as missing value that can be interpolated; According to the data source metadata, this category information is attached to the record, and the missing values can be divided into critical missing values and general missing values.

7. The real-time acquisition method of the industry investment return rate according to claim 6, characterized in that: For data with general missing values, forward value extension and interval interpolation are used for rapid filling. For critical missing fields, machine learning or rule algorithms are introduced to predict the missing values with high precision; After the interpolation is completed, a filled mark is added to the corresponding record, and the specific interpolation method is written into the metadata; if the interpolation fails, an unfilled mark is added and a null value is retained in the result set; If a record has both a delay mark and a null value field at the same time, it is preferentially regarded as missing data and can enter the interpolation algorithm; if the record is lost, no interpolation process is triggered anymore. In this way, a priority table can be established for all marks, indicating the compatibility or mutual exclusion relationship between each other to avoid conflicts.

8. The real-time acquisition method of the industry investment return rate according to claim 7, characterized in that: For values from different data sources within the same time window or adjacent time windows, if there are logical conflicts, they should be identified and adjusted; based on the constraints or association rules predefined in the metadata, the interpolated or original value records are checked. If a conflict is found, it is marked as slightly inconsistent and severely inconsistent; If a record is marked as severely inconsistent, its value can be applied with a penalty coefficient in subsequent operations to form an adjusted result; For slightly inconsistent records, only monitoring or recording is performed, the final validity or corrected value of each record is updated, and a conflict mark or adjustment information is attached to the metadata.

9. The method for real-time obtaining of the industry investment return rate according to claim 8, characterized in that: Combined with the time precision information recorded by the metadata, it is detected whether the actual acquisition precision of the data source is lower than the requirement. If a certain data source is marked with millisecond precision but the precision does not meet the requirement, a precision alarm is given or a low-precision mark is added at this stage; If there is an external synchronization signal, it can be compared with its own time at this time. If the deviation between the two exceeds the preset deviation threshold, a fine calibration algorithm will be triggered at the next moment; No large-scale calibration is performed within the preset deviation threshold. When the external signal is unavailable or unstable, this link is skipped and subsequent correction is completed only based on the internal correlation analysis.

10. The method for real-time obtaining of the industry investment return rate according to claim 9, characterized in that: Cross-evaluate different data sources within the same time window or triggered by the same event, calculate the correlation degree with the help of the marked data, and construct a metric function with an integral kernel; Make an overall correction to the current clock according to the external signal, perform the above-mentioned cross-correlation calculation on the local drift amount between different data sources. If the external signal is missing or unstable, the offset between the sources is estimated based on the metric function; Store the calculation results in the deviation amount mapping table offset mapping table, record the optimal lag of each pair of data sources, determine a synchronization path through the minimum spanning tree or minimum loop covering method, and output the deviation amount mapping table offset mapping table and the synchronization path or priority list.

11. The method for real-time obtaining of the industry investment return rate according to claim 10, characterized in that: According to the synchronization path, correct the timestamps of each data source according to their optimal deviation amounts relative to the high-precision reference and adjacent sources: Set a safety boundary for each record. If the correction amount exceeds this boundary, manual intervention is required or excessive adjustment is recorded separately; After completing the fine-tuning of all data sources, note in the record metadata that the timestamp has been corrected, use the new timestamp as the final result, and provide an adjustment log when outputting the result data.

12. The method for real-time obtaining of the industry investment return rate according to claim 11, characterized in that: According to the time alignment result, use the shard mapping table to perform sharding according to the business dimension or time period and allocate it to different computing nodes; to ensure that the time offset function is adaptive to network fluctuations or node own clock errors; When any node detects that the load is close to the upper limit during the communication peak period, it requests shard migration to the central scheduler or other nodes, triggering fault-tolerant consensus processing.

13. The method for real-time obtaining of the industry investment return rate according to claim 12, characterized in that: When any node has an irrecoverable failure or a high load exceeding the tolerance threshold, an abnormality is detected through the shard status record associated with the node and a failure alarm is sent to other available nodes and the scheduling center; After receiving the failure alarm, other nodes trigger a fault-tolerant consensus process using the broadcast time offset function and shard mapping information to determine the optimal shard takeover plan; If the node recovers within a short time, set its shard status to temporarily offline, and other nodes save a small amount of key incremental data for it first; If it has not been restored after exceeding the fault judgment threshold duration, its shard will be officially migrated to an available node.

Citation Information

Patent Citations

  • Real-time data processing system and method based on Flink

    CN116932598A

  • Data sorting method and system for massive industrial time series data calculation

    CN117390075A

  • Coal mine dust diffusion intelligent monitoring method based on distributed sensor network

    CN119559763A

  • System for issuing time stamp certificate and its system program

    JP2004038378A

  • Managing timestamps in a sequential update stream recording changes to a database partition

    US11314779B1

Cited By

  • Data fusion method and processing system based on industrial internet

    CN120524437A

  • An industrial internet-based data fusion method and processing system

    CN120524437B

  • Multi-source heterogeneous bus data parallel acquisition and recording method and system

    CN120561187A

  • Intelligent assessment method for dynamic return rate of medical project

    CN121258199A

  • Medical project dynamic return rate intelligent evaluation method

    CN121258199B