Computer data integration management method, device and electronic equipment

CN122838231APending Publication Date: 2026-09-29BEIJING GEYUANTENG TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610987848.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]为此,本发明提供一种计算机数据集成管理方法、装置及电子设备,用以克服现有技术中,仅依托单一指标开展节点监控,缺少对通信链条动态运行状态的分层评估,降低数据集成系统的故障预判能力与运行可靠性的问题

Benefits of technology

[0053]与现有技术相比,本发明通过获取计算机对应各异构数据源以及各数据节点的运行数据,构建运行态势画像;提取若干通信链条对应的处理匹配特征,评估通信链条的匹配失调度,以对通信链条进行标注;基于标注结果,适应性地对通信链条的数据通信进行评估分析。本发明通过构建多层次、递进式的通信链条运行健康度评估体系,实现了对异构数据集成场景下通信链条运行状态的精细化感知、异常精准识别与调度策略的自适应优化,从而提升了数据集成系统的运行可靠性、运维效率和自动化水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838231A_ABST
    Figure CN122838231A_ABST
Patent Text Reader

Abstract

The present application relates to the field of data integration management, and more particularly to a computer data integration management method and device and electronic equipment, the present application constructs a running situation image by obtaining the running data of each heterogeneous data source and each data node corresponding to the computer, extracts the processing matching features corresponding to a plurality of communication chains, evaluates the matching disorder of the communication chains, and labels the communication chains, and based on the labeling result, adaptively evaluates and analyzes the data communication of the communication chains. The present application improves the running reliability, operation and maintenance efficiency and automation level of the data integration system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data integration management, and more particularly to a computer data integration management method, apparatus, and electronic device. Background Technology

[0002] With the rapid deployment of big data and multi-source heterogeneous data services, enterprise data integration platforms need to connect to various heterogeneous data sources such as databases and real-time streams. Massive amounts of data rely on multi-level data nodes and chain-like transmission links to complete aggregation, cleaning, transformation, and storage. As the core carrier of data flow, the communication chain's upstream and downstream processing matching degree, traffic stability, and fault self-healing capability directly determine the throughput, latency, and data quality of the overall data integration system. Therefore, refined monitoring, quantitative evaluation, and autonomous scheduling optimization of the communication chain are key to improving the stability and automated operation and maintenance capabilities of heterogeneous data integration platforms.

[0003] For example, Chinese Patent Publication No. CN113190517A discloses a data integration method, apparatus, electronic device, and computer-readable medium. One specific embodiment of the method includes: periodically acquiring source data indicated by a target data source from multiple data sources; selecting a data mapping method corresponding to the type of the target data source to map the currently acquired source data, obtaining current mapped data; performing a data structure comparison analysis between the current mapped data and historical mapped data, wherein the historical mapped data is the mapped data obtained by mapping the source data indicated by the previously acquired target data source; and generating a change prompt message in response to differences in data structure. This embodiment achieves data integration from different data sources and can dynamically acquire changes in the data structure of the source data.

[0004] However, the following problems still exist in the existing technology.

[0005] Relying solely on a single indicator for node monitoring lacks a layered assessment of the dynamic operating status of the communication chain, reducing the fault prediction capability and operational reliability of the data integration system. Summary of the Invention

[0006] To address this, the present invention provides a computer data integration management method, apparatus, and electronic device to overcome the problems in the prior art, which rely solely on a single indicator for node monitoring, lack hierarchical evaluation of the dynamic operating status of the communication chain, and reduce the fault prediction capability and operational reliability of the data integration system.

[0007] To achieve the above objectives, the present invention provides a computer data integration and management method, comprising:

[0008] Obtain operational data from various heterogeneous data sources and data nodes corresponding to the computer, and construct an operational status profile;

[0009] Extract processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains.

[0010] Based on the annotation results, the data communication of the communication chain is evaluated and analyzed, including,

[0011] Identify whether there are abnormal communication time domain segments within the communication chain, and assess whether there is a logic overload based on the communication mutation characteristics of the abnormal communication time domain segments;

[0012] Based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window, assess whether the communication chain has disturbance self-recovery capability;

[0013] The operational deviation of the real-time operational status profile is determined, and the scheduling adjustment intensity of the communication chain is adaptively adjusted.

[0014] Furthermore, the process of evaluating the mismatch scheduling of the communication link includes:

[0015] Calculate the data processing rate ratio and the lag ratio respectively;

[0016] The sum of the data processing rate ratio adjusted according to the first influence ratio and the lag ratio adjusted according to the second influence ratio is taken as the matching misscheduling;

[0017] Among them, the matching features include the offset of the ratio of data processing rates between upstream and downstream nodes and the lag in data processing between upstream and downstream nodes.

[0018] Furthermore, the communication chain is labeled, including:

[0019] If the mismatch of a communication chain is greater than or equal to a preset mismatch, then the communication chain is marked.

[0020] If a communication chain is marked, the data communication of that communication chain is evaluated and analyzed.

[0021] Furthermore, the process of identifying abnormal communication time domain segments includes:

[0022] The source-end data of the communication chain is traversed within a predetermined sliding time window to produce time-series data of the output rate.

[0023] Calculate the ratio of the average rate within the sliding time window to the average rate of adjacent time periods outside the window;

[0024] When the ratio exceeds the preset burst multiplier and the duration reaches the preset burst duration, the corresponding time domain segment is identified as the communication abnormal time domain segment.

[0025] Furthermore, assess whether there is a logic overload, including:

[0026] If the null value rate of a data field is greater than the preset null value rate, and the unexpected conversion frequency of the data type is greater than the preset conversion frequency, then a logical overload is determined to exist.

[0027] Among them, communication mutation characteristics include the null value rate of data fields and the frequency of unexpected data type conversions.

[0028] Furthermore, the process of assessing whether the communication chain has disturbance self-recovery capability includes:

[0029] Record the retry trigger events of processes within the communication chain and the corresponding retry success rate;

[0030] Construct a curve showing the retry success rate as a function of the number of retries, and fit the decay slope of the curve.

[0031] If the attenuation slope is greater than the preset attenuation slope, it is determined that there is a retry failure.

[0032] If the difference in waiting time between two consecutive retry trigger events is consistently positive, and the increase in waiting time exceeds the preset increase, then a retry waiting anomaly is determined to exist.

[0033] If there are no retry activation exceptions and no retry waiting exceptions at the same time, it is determined that the communication chain has disturbance self-recovery capability.

[0034] Furthermore, the process of determining the operational deviation of the real-time operational status profile includes:

[0035] Compare the real-time operational status profile with the baseline operational envelope;

[0036] The absolute value of the minimum difference between several data items in the statistical operation data and the corresponding baseline data item range;

[0037] The absolute value of the minimum difference is taken as the running deviation.

[0038] The reference operating envelope is the allowed range of values ​​for the data of the communication chain under normal operating conditions.

[0039] Furthermore, the scheduling adjustment intensity of the communication chain is adaptively adjusted, including:

[0040] The intensity of scheduling adjustments is positively correlated with the degree of operational deviation.

[0041] Furthermore, an electronic device that applies a computer data integration and management method is also provided, including:

[0042] The data acquisition module is used to acquire the computer's operational data from various heterogeneous data sources and data nodes, and to build an operational status profile.

[0043] The processing evaluation module is used to extract the processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains.

[0044] The communication analysis module is used to evaluate and analyze the data communication of the communication chain based on the annotation results, including:

[0045] Identify whether there are abnormal communication time domain segments within the communication chain, and assess whether there is a logic overload based on the communication mutation characteristics of the abnormal communication time domain segments;

[0046] Based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window, assess whether the communication chain has disturbance self-recovery capability;

[0047] The operational deviation of the real-time operational status profile is determined, and the scheduling adjustment intensity of the communication chain is adaptively adjusted.

[0048] Furthermore, an apparatus is also provided, comprising:

[0049] One or more processors;

[0050] Memory;

[0051] and one or more programs;

[0052] The one or more programs are configured to be executed by one or more processors, and the memory includes a storage medium storing a computer program that, when executed by the processor, can be used to perform a computer data integration management method.

[0053] Compared with existing technologies, this invention constructs an operational status profile by acquiring operational data from various heterogeneous data sources and data nodes of the computer; extracts processing matching features corresponding to several communication chains, evaluates the mis-scheduling of communication chains, and labels the communication chains; based on the labeling results, it adaptively evaluates and analyzes the data communication of the communication chains. This invention, by constructing a multi-level, progressive communication chain operational health assessment system, achieves refined perception of the operational status of communication chains in heterogeneous data integration scenarios, accurate anomaly identification, and adaptive optimization of scheduling strategies, thereby improving the operational reliability, maintenance efficiency, and automation level of the data integration system.

[0054] In particular, this invention assesses and locates problems in upstream and downstream nodes of the communication chain. By observing the transmission and processing of data flow between these nodes, it intuitively reflects the dynamic game between upstream production and downstream consumption, thus enabling the assessment of the communication chain's operational status. On one hand, from a spatial perspective, the ratio of data processing rates between upstream and downstream nodes determines the degree of matching in data processing capabilities, quantifying the balance between upstream production data speed and downstream consumption data speed. If the ratio is greater than 1, it indicates that upstream production speed exceeds downstream processing capacity, causing data to accumulate in the buffer, eventually leading to memory overflow or increased latency, thereby quantifying the risk of data backlog. If the ratio is less than 1, it indicates that downstream processing capacity is excessive, resulting in a prolonged waiting state, suggesting unreasonable resource allocation or upstream bottlenecks, thus quantifying resource waste or starvation. When the ratio approaches 1, it indicates that upstream and downstream nodes are in a synchronous processing state, with a high match between production and consumption capabilities, and the system operates smoothly. Accordingly, this invention sets the offset of the data processing rate ratio between upstream and downstream nodes to determine the degree of mismatch in processing capabilities between them, unifying the two types of problems—backlog risks caused by excessive upstream production and resource waste caused by excessive downstream processing—within the same quantitative framework. A larger offset indicates a worse match in processing capabilities between upstream and downstream nodes, and a greater deviation of the system from equilibrium. On the other hand, from a time perspective, the timeliness of data transmission is measured by the lag in data processing between upstream and downstream nodes, determining the time delay from when data enters the upstream node to when it is successfully processed by the downstream node. A larger lag indicates a larger backlog of data from upstream production to downstream consumption, and a worse system real-time performance. Through the synergistic analysis of the aforementioned two features, this invention determines the mismatch in the communication chain, comprehensively assessing the health of the communication chain and the matching efficiency of the corresponding upstream and downstream nodes, providing reliable data support for subsequent anomaly labeling, bottleneck location, and adaptive scheduling.

[0055] In particular, this invention characterizes system pressure using data quality and captures rate spikes through a sliding window ratio method to identify whether a substantial and non-random, sudden surge in data production at the source has occurred, indicating a communication anomaly. Specifically, the relative surge intensity is determined by the ratio of the average rate within the sliding time window to the average rate of adjacent time periods outside the window. A larger ratio indicates a more drastic fluctuation in the instantaneously injected data traffic from upstream relative to its historical fluctuations. Furthermore, the duration of the abnormally high traffic state is quantified by the duration of the spike, determining the persistence and stability of the traffic anomaly. A longer duration indicates that the anomaly is less of a random network jitter or CPU time slice switching and more of a real, continuous source of backend pressure. Thus, this invention, through dual quantitative constraints of relative surge intensity and anomaly persistence, achieves accurate identification of communication anomaly time domain segments. It can promptly capture sudden anomalies in data production and filter out false alarms caused by instantaneous fluctuations through the duration condition, improving the accuracy and practicality of anomaly detection and providing a highly reliable trigger time domain segment for subsequent logic overload assessment.

[0056] In particular, this invention considers that in real-world scenarios, when a system is subjected to pressure exceeding its design load, the data processing logic will be the first to collapse, manifested as a sharp increase in field missing rates and frequent type conversion failures. Accordingly, this invention quantifies the proportion of records whose fields were not successfully filled or assigned values ​​during data transmission and processing relative to the total number of records processed within a sliding time window by using the data field null value rate. This determines the degree of data integrity loss at the logical level when the system's processing capacity is insufficient. A higher null value rate indicates that, under the current load, more data processing requests fail to complete the complete field assignment process due to timeouts, exception skipping, buffer overflows, logical truncation, etc., thus reflecting the integrity loss of data at the output end. Furthermore, this invention quantifies the number of unexpected forced type conversions or conversion failures that occur during data flow or processing between nodes due to format mismatches, parsing failures, type conflicts, etc., by using the frequency of unexpected data type conversions, determining the degree of data logic disorder. A higher frequency indicates that the system cannot process more data according to the established data type specifications when processing the current data volume. For example, anomalies such as receiving strings or unparseable date formats when the data should be an integer field reflect a lack of standardization in data processing. Therefore, this invention uses the aforementioned dual-indicator collaborative evaluation to determine the confidence level of logical overload from both the processing result and processing process perspectives, thereby improving the accuracy of anomaly detection.

[0057] In particular, this invention constructs a disturbance self-recovery capability assessment mechanism based on retry behavior analysis, characterizing the true state and evolution trend of the system's disturbance self-recovery capability. At the effect level, this invention determines the rate and trend of the retry success rate decrease as the number of retries increases by using the decay slope of the retry success rate, quantifying the marginal benefit change brought by each additional retry. The larger the decay slope, the more drastically the marginal benefit of retrying decreases, the retry behavior itself is failing, and the system may have encountered an unrecoverable logical error or resource exhaustion. At the direction level, this invention determines the correctness of the retry recovery direction by taking the positive value of the difference in waiting time corresponding to two adjacent retry trigger events. The larger the difference in waiting time, the more the system not only fails to accelerate recovery but also continuously prolongs the waiting time. At the severity level, this invention quantifies the severity of the retry strategy by the increase in waiting time. The larger the increase, the more the system has fallen into a vicious cycle of waiting longer and longer without success, and the worse its self-recovery capability. Furthermore, this invention combines these three elements to form a progressive evaluation logic that assesses the effectiveness of retrying, the correctness of the recovery direction, and the extent of the error. This constructs a complete self-healing capability diagnostic link to evaluate whether the communication chain possesses the ability to self-recover from disturbances. This enables a high-confidence determination of whether the communication chain can recover on its own after disturbances and whether the recovery is efficient, thereby improving the self-healing capability and operational efficiency of large-scale data integration systems in complex operating environments.

[0058] In particular, this invention compares the real-time operational status profile with the baseline operational envelope to quantify the severity of the system's current overall operational state deviating from the normal healthy baseline. A greater deviation indicates that at least one indicator has approached or even exceeded the permissible range boundary, severely degrading the system's health. Furthermore, this invention establishes a mapping relationship between the intensity of scheduling intervention and the severity of system anomalies; a greater deviation indicates a higher anomaly risk, requiring greater intervention intensity. Accordingly, this invention uses operational deviation as the final integrated output of the preceding detection and evaluation stages, maximizing resource utilization efficiency while ensuring system stability, enabling large-scale data integration systems to possess autonomy and self-healing capabilities in dynamic and complex environments. Attached Figure Description

[0059] Figure 1 This is a schematic diagram illustrating the steps of a computer data integration and management method according to an embodiment of the invention;

[0060] Figure 2 A logic block diagram for identifying communication anomalies in the time domain, as shown in an embodiment of the invention;

[0061] Figure 3 A logic block diagram illustrating whether the communication chain in an embodiment of the invention has the ability to self-recover from disturbances;

[0062] Figure 4This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation

[0063] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0064] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0065] Please see Figure 1 The diagram illustrates the steps of a computer data integration and management method according to an embodiment of the present invention. The method includes:

[0066] Step S1: Obtain the operating data of each heterogeneous data source and each data node corresponding to the computer, and construct an operating status profile;

[0067] Step S2: Extract the processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains.

[0068] Step S3, based on the annotation results, evaluate and analyze the data communication of the communication chain, including,

[0069] Step S3001 identifies whether there is a communication abnormal time domain segment within the communication chain, and assesses whether there is a logic overload based on the communication mutation characteristics of the communication abnormal time domain segment;

[0070] Step S3002 evaluates whether the communication chain has disturbance self-recovery capability based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window;

[0071] Step S3003 determines the operational deviation of the real-time operational status profile and adaptively adjusts the scheduling adjustment intensity of the communication chain.

[0072] Specifically, the operational data includes processing matching features, communication mutation features, self-recovery features, and operational deviation.

[0073] Specifically, the method for collecting and acquiring the relevant features corresponding to the operational data is not specifically limited. Several preferred implementation methods are provided below:

[0074] The data processing rate of upstream and downstream nodes can be obtained in the following ways: First, by embedding data points within the nodes and using counters to count the number of data records or the amount of data successfully processed per unit time; second, by deploying log collectors to parse and report processing rate information from the operation logs of each node. Then, the ratio of the data processing rate of the upstream node to that of the downstream node can be calculated to obtain the data processing rate ratio.

[0075] Specifically, the data processing lag between upstream and downstream nodes refers to the total amount of data that has been produced by upstream producer nodes but not yet processed by downstream consumer nodes. Taking Kafka message middleware as an example, the total records-lag of consumer nodes can be obtained through Kafka command-line tools or JMX metrics. This value quantifies the total number of messages written to all partitions but not yet committed and confirmed by consumer nodes. In other message queue or buffering scenarios, corresponding backlog metrics can be obtained by analogy.

[0076] To assess the null value rate of data fields, a data quality verification operator can be embedded in the data stream processing pipeline to count the number of null values ​​in each batch of data in real time, and then output the null value rate of data fields within the sliding window.

[0077] The frequency of unexpected data type conversions can be obtained in the following ways: First, count the number of type conversion exception events recorded in the node's runtime log; second, integrate a data validation framework, define field type constraints in the data quality validation rules, and trigger a count when the validation fails.

[0078] The retry success rate can be obtained by relying on log event aggregation. Each retry event is recorded in the logs of each node, including the retry sequence number and success / failure status. The log aggregation system counts the total number of retries and the number of successes by time window, and the ratio of the number of successes to the total number of retries is used as the retry success rate.

[0079] Specifically, the waiting time between two consecutive retry trigger events refers to the time interval between the previous retry failure and the initiation of the current retry. For example, this can be obtained by setting callback hooks at the wait start event and retry initiation event in the retry library, recording the timestamps of the two events, and the difference between them being the waiting time; alternatively, it can be obtained by extracting the timestamps from the retry logic within the log records, specifically the timestamps when entering the waiting state and when each retry is initiated, and calculating the timestamp difference.

[0080] The increase in waiting time refers to the magnitude of the increase between two consecutive retry waiting times. For example, the difference between the current retry waiting time and the previous retry waiting time can be calculated, and the ratio of this difference to the previous retry waiting time can be used as the increase. Alternatively, all retry waiting times can be written sequentially into a time-series database, and the rate of change and increase of adjacent data points can be calculated using query statements.

[0081] For the operational deviation, the system first obtains the real-time values ​​of each data item, then compares and calculates them with the baseline operational envelope. The output value after normalizing the absolute value of the minimum difference between the real-time value and the corresponding allowable value range is used as the operational deviation. If the real-time values ​​of multiple data items are not within the allowable value range, the sum of the normalized output values ​​of each data item is used as the operational deviation.

[0082] Among them, the historical operating data of the system under known normal conditions is collected, and the allowable value range of each data item is determined by statistical modeling or manual calibration to establish a benchmark operating envelope. The benchmark envelope is dynamically adjusted as the system operating environment changes by using a sliding window or periodic recalculation method.

[0083] Specifically, the communication chain refers to a series of orderly data transmission and processing links involved in the complete process from the generation of data at the source node to its successful reception and processing by the downstream node in a computer data integration environment.

[0084] In this embodiment, a predetermined time window is set to collect the self-recovery characteristics of the communication chain after identifying a communication anomaly time domain segment and determining that there is a logical overload, in order to evaluate whether it has the ability to self-recover from disturbances. For example, the window length can be calculated backward from the maximum fault recovery time agreed in the service SLA, and should be less than the recovery time specified in the SLA to reserve operational space for scheduling intervention; alternatively, the window length can be set to the 95th percentile of historical successful self-recovery times based on statistical analysis of historical self-recovery times, which will not be elaborated further.

[0085] The self-recovery features include retry success rate, waiting time corresponding to two adjacent retry trigger events, and the increase in waiting time.

[0086] Specifically, the process of constructing an operational status profile includes:

[0087] The system collects operational data from various heterogeneous data sources and data nodes, and preprocesses and merges the operational data.

[0088] Several data items are identified from the preprocessed and fused data, and the data items are classified and organized to form the operational status profile;

[0089] The preprocessing includes data cleaning (removing noisy data, handling missing values, and standardizing timestamp formats) and data standardization (converting data from different sources into a unified format and unit).

[0090] The fusion involves integrating relevant data from different data sources according to spatiotemporal relationships or other association rules to form a more complete information view. For example, aligning the CPU utilization of a node with the amount of data it processes over time can accurately reflect the overall load status of that node at a specific moment.

[0091] The data items include general operating metrics (CPU utilization, memory usage, data processing rate, throughput, and response time, etc.) and related features corresponding to the operating data (processing matching features, communication mutation features, self-recovery features, and operating deviation, etc.).

[0092] It is understood that the aforementioned operational status profile provides a structured data foundation for subsequent comparisons between the real-time operational status profile and the baseline operational envelope, thereby driving adaptive scheduling adjustments. Furthermore, once constructed, the operational status profile should be continuously updated as the system operates to reflect the real-time system operational status.

[0093] Specifically, the process of assessing the mismatch of communication links includes:

[0094] Calculate the data processing rate ratio and the lag ratio respectively;

[0095] The sum of the data processing rate ratio adjusted according to the first influence ratio and the lag ratio adjusted according to the second influence ratio is taken as the matching misscheduling;

[0096] The data processing rate ratio is the ratio of the offset of the data processing rate ratio between upstream and downstream nodes to a preset offset;

[0097] The lag ratio is the ratio of the lag in data processing between upstream and downstream nodes to a preset lag.

[0098] Among them, the matching features include the offset of the ratio of data processing rates between upstream and downstream nodes and the lag in data processing between upstream and downstream nodes.

[0099] Specifically, the sum of the first influence ratio and the second influence ratio is 1. Initially, the first influence ratio and the second influence ratio can be set equally, i.e., both are set to 0.5, to reflect the initial equal importance of data processing rate matching and data lag in the matching misscheduling assessment. Subsequently, adaptive adjustments can be made according to the sensitivity requirements of data throughput efficiency and data real-time performance in actual business scenarios. If the scenario focuses more on data processing throughput and matching efficiency, the value of the first influence ratio should be increased; if the scenario focuses more on the timeliness of data transmission and backlog risk control, the value of the second influence ratio should be increased. In addition, those skilled in the art can determine the specific values ​​through on-site calibration or historical data fitting, which will not be elaborated here.

[0100] Specifically, the data processing rate ratio between upstream and downstream nodes is used to determine the degree of matching in data processing capabilities, quantifying the balance between the upstream data production speed and the downstream data consumption speed. If the ratio is greater than 1, it indicates that the upstream production speed exceeds the downstream processing capacity, causing data to accumulate in the buffer, eventually leading to memory overflow or increased latency, thus quantifying the risk of data backlog. If the ratio is less than 1, it indicates that the downstream processing capacity is excessive, resulting in a prolonged waiting state, suggesting unreasonable resource allocation or upstream bottlenecks, thus quantifying resource waste or starvation. When the ratio approaches 1, it indicates that upstream and downstream nodes are processing synchronously, with a high match between production and consumption capabilities, and the system operates smoothly. Based on this, a deviation in the data processing rate ratio between upstream and downstream nodes is set to determine the degree of mismatch in processing capabilities between them, unifying the backlog risk caused by excessive upstream production and the resource waste caused by excessive downstream processing into the same quantitative framework. The larger the deviation, the worse the matching of processing capabilities between upstream and downstream nodes, and the further the system deviates from equilibrium. The offset is referenced to a preset processing rate ratio, and represents the degree of deviation of the data processing rate ratio of upstream and downstream nodes from the benchmark. The offset is obtained by calculating the absolute value of the difference between the data processing rate ratio of upstream and downstream nodes and the preset processing rate ratio.

[0101] In this embodiment, a preset processing rate ratio is set to quantify the ideal target rate matching state between upstream and downstream nodes, serving as a reference benchmark for evaluating the upstream and downstream data processing rate ratio. For example, the preset processing rate ratio can be set to 1 to represent the ideal equilibrium state where the upstream production rate and downstream consumption rate are perfectly matched. The degree of deviation of the real-time rate ratio from 1 reflects the degree to which the system deviates from equilibrium. Furthermore, those skilled in the art can also make adaptive adjustments based on the differences in hardware configuration between upstream and downstream nodes. If the upstream node's hardware configuration is higher than that of the downstream node, it indicates that the upstream has a higher theoretical output ceiling, and the preset processing rate ratio can be moderately increased around 1; conversely, if the downstream node has a higher configuration, the preset processing rate ratio can be moderately decreased.

[0102] In this embodiment, a preset offset is set to quantify the system's upper limit of tolerance for mismatch in upstream and downstream processing capabilities. For example, it can be set according to system design tolerances; for instance, if a fluctuation of ±20% is allowed, the preset offset is set to 0.2. Alternatively, it can be based on statistical analysis of historical normal operation data, taking the high quantile of historical offsets, such as 95% or 99%, as the preset offset. It can also be set in conjunction with differences in hardware configuration and performance redundancy between upstream and downstream nodes, or deduced from business SLA requirements. Furthermore, the specific value of the preset offset can be set and adjusted by those skilled in the art through on-site calibration or historical data fitting based on the actual system architecture, business scenario, and operating environment; this will not be elaborated further.

[0103] Specifically, the timeliness of data transmission is measured by the lag in data processing between upstream and downstream nodes, determining the length of time it takes for data to travel from the upstream node to its successful processing by the downstream node. A larger lag indicates a greater backlog of data from upstream production to downstream consumption, and consequently, poorer system real-time performance.

[0104] In this embodiment, a preset hysteresis is set to quantify the system's tolerance for data backlog. For example, the maximum allowable backlog can be calculated based on the service's latency requirements, such as second-level response time, as the preset hysteresis; alternatively, a safety threshold can be set based on the memory size of the downstream node's buffer, such as not exceeding 80% of the total buffer capacity.

[0105] Specifically, labeling the communication chain includes:

[0106] If the mismatch of a communication chain is greater than or equal to a preset mismatch, then the communication chain is marked.

[0107] If a communication chain is marked, the data communication of that communication chain is evaluated and analyzed.

[0108] In this embodiment, the upper limit of the system's tolerance for upstream and downstream mismatches is determined by a preset mismatch schedule. For example, a historical period of stable and fault-free system operation can be selected, the corresponding mismatch schedule can be calculated, and the high quantile, such as 90% or 95%, can be taken as the preset mismatch schedule.

[0109] Please see Figure 2 The diagram shown is a logical block diagram for identifying communication anomaly time domain segments according to an embodiment of the present invention. The process of identifying communication anomaly time domain segments includes:

[0110] The source-end data of the communication chain is traversed within a predetermined sliding time window to produce time-series data of the output rate.

[0111] Calculate the ratio of the average rate within the sliding time window to the average rate of adjacent time periods outside the window;

[0112] When the ratio exceeds the preset burst multiplier and the duration reaches the preset burst duration, the corresponding time domain segment is identified as the communication abnormal time domain segment.

[0113] The duration of the adjacent time periods is equal to the duration of the sliding time window.

[0114] Specifically, the relative surge intensity of the data is determined by the ratio of the average rate within the sliding time window to the average rate of adjacent time periods outside the window. The larger this ratio, the more drastic the instantaneous data flow injected from upstream relative to its own historical fluctuations.

[0115] In this embodiment, a preset burst multiplier is set to distinguish between normal fluctuations and sudden surges. For example, the coefficient of variation (Cv) of historical rate data can be calculated through historical fluctuation analysis, and the preset burst multiplier can be set to 1 + n*CV. Here, n is an adjustable parameter, typically taking the value of 2 or 3, corresponding to approximately 95% or 99.7% confidence intervals for normal fluctuations, respectively.

[0116] Specifically, the duration of abnormally high traffic is quantified by the burst duration to determine the persistence and stability of the traffic anomaly. The longer the duration, the less likely the anomaly is an instantaneous fluctuation caused by accidental network jitter or CPU time slice switching, but rather a real and continuous source of backend pressure.

[0117] In this embodiment, a preset burst duration is set to filter out transient spikes and confirm persistent anomalies. For example, the preset burst duration can be based on the 95th percentile of the duration distribution of historical abnormal events; it can also be set according to the system sampling period, typically from several seconds to several minutes, such as 30 seconds or 2 minutes, to eliminate single-point jitter interference.

[0118] Specifically, the predetermined sliding time window can be set according to the sampling period of the source data output rate. In this embodiment, it is set to 3 to 10 times the sampling period to ensure that the window contains a sufficient number of sampling points to reflect the statistical characteristics of the rate. Of course, those skilled in the art can adjust it flexibly according to the characteristics of the traffic flow. For systems with stable traffic changes, the window can be appropriately shortened to improve detection sensitivity; for systems with large traffic fluctuations, the window should be appropriately extended to reduce false alarms caused by normal fluctuations.

[0119] Specifically, assessing for logical overload includes:

[0120] If the null value rate of a data field is greater than the preset null value rate, and the unexpected conversion frequency of the data type is greater than the preset conversion frequency, then a logical overload is determined to exist.

[0121] Among them, communication mutation characteristics include the null value rate of data fields and the frequency of unexpected data type conversions.

[0122] Specifically, the null value rate quantifies the proportion of fields that are not successfully filled or assigned values ​​during data transmission and processing. This determines the degree of data integrity loss at the logical level when the system's processing capacity is insufficient. A higher null value rate indicates that, under the current load, more data processing requests fail to complete the complete field assignment process due to timeouts, exception skipping, cache overflows, logical truncation, etc., thus reflecting the integrity loss of data at the output end.

[0123] In this embodiment, a preset null value rate is set to assess the tolerance for data integrity. For example, the null value rate of the data source under normal conditions can be analyzed, and separate settings can be made for key fields and non-key fields, taking the higher percentile, such as 95% or 99%, as the preset null value rate; for key business fields, such as primary keys and business identifiers, the preset null value rate should be close to 0.

[0124] Specifically, the frequency of unexpected data type conversions quantifies the number of times unexpected forced type conversions or conversion failures occur during data flow or processing between nodes due to reasons such as format mismatch, parsing failure, and type conflicts, thus determining the degree of disorder in the data logic. A higher frequency indicates that the system cannot execute data according to established data type specifications when processing the current data volume. For example, receiving exceptions such as strings or unparseable date formats when the data should be an integer field reflects a lack of adherence to data processing standards at the data flow end.

[0125] In this embodiment, a preset conversion frequency is set to assess the tolerance for data standardization. For example,

[0126] It can analyze the null value rate of the data source under normal conditions, distinguish between key fields and non-key fields and set them separately, taking the higher percentile, such as 95%, as the preset conversion frequency; for core data types, the preset conversion frequency should be significantly lower than that for non-core types.

[0127] Please see Figure 3 The diagram shown is a logic block diagram for determining whether a communication chain has disturbance self-recovery capability according to an embodiment of the present invention. The process of evaluating whether the communication chain has disturbance self-recovery capability includes:

[0128] Record the retry trigger events of processes within the communication chain and the corresponding retry success rate;

[0129] Construct a curve showing the retry success rate as a function of the number of retries, and fit the decay slope of the curve.

[0130] If the attenuation slope is greater than the preset attenuation slope, it is determined that there is a retry failure.

[0131] If the difference in waiting time between two consecutive retry trigger events is consistently positive, and the increase in waiting time exceeds the preset increase, then a retry waiting anomaly is determined to exist.

[0132] If there are no retry activation exceptions and no retry waiting exceptions at the same time, it is determined that the communication chain has disturbance self-recovery capability.

[0133] Specifically, the process of constructing the curve of the retry success rate as a function of the number of retries includes:

[0134] Record each retry trigger event initiated by each process in the communication chain. Each retry trigger event log shall contain at least the retry trigger event identifier, the corresponding retry sequence number, and the status indicator of whether the retry was successful.

[0135] Within the predetermined time window, all retry records in all retry events are grouped by retry number, with those having the same retry number grouped together, and the total number of retries and the number of successful retries are counted in each group.

[0136] For each retry number k, the ratio of the number of successful retryes within the corresponding group to the total number of retryes is taken as the success rate of the kth retry.

[0137] Construct a scatter dataset with the retry sequence number as the x-axis and the corresponding retry success rate as the y-axis;

[0138] The least squares method was used to perform linear fitting on the scattered dataset to obtain the fitted line R(k)=a×k+b, where a is the decay slope, b is the intercept, and k is the retry number.

[0139] The retry trigger event is a series of retry operations triggered after a data processing failure. The retry sequence number is used to identify the order of each retry operation within a single retry trigger event. The decay slope is evaluated and analyzed by taking the absolute value.

[0140] Specifically, the rate and trend of the retry success rate decline as the number of retries increases are determined by the decay slope of the retry success rate, quantifying the change in marginal revenue brought by each additional retry. The larger the decay slope, the more drastically the marginal revenue of retrying decreases, the less effective the retrying behavior itself is, and the system may have encountered an unrecoverable logical error or resource exhaustion.

[0141] In this embodiment, a preset decay slope is set to determine whether the retry effect deteriorates rapidly. For example, the decay slope under ideal conditions can be calculated based on theoretical models such as exponential backoff as the preset decay slope; alternatively, the historical retry success rate decay curve of the system under no real fault conditions can be collected, its slope distribution can be fitted, and the 95th percentile can be taken as the preset decay slope.

[0142] Specifically, assuming the difference in waiting time is a positive value, the severity of the retry strategy is quantified by the increase in waiting time. The larger the increase, the more the system has fallen into a vicious cycle of waiting longer and longer without success, and the worse its self-recovery ability is.

[0143] In this embodiment, a preset increase is set to determine whether the retry strategy is out of control. For example, the maximum increase in waiting time can be calculated based on the maximum recovery time allowed by the business SLA as the preset increase; alternatively, the preset increase can be based on the backoff strategy specifications adopted by the system, such as exponential backoff where the waiting time doubles each time, and the theoretical increase multiple, such as 2 times, combined with a certain tolerance margin, such as 2.5 times or 3 times, can be used as the preset increase.

[0144] It is understood that the specific values ​​of each relevant preset value can be set and adjusted by those skilled in the art through on-site calibration or historical data fitting based on the actual system architecture, business scenario and operating environment; and during system operation, the relevant preset values ​​can be dynamically corrected by using a sliding window or periodic recalculation method to adapt to changes in the system operating status, which will not be elaborated here.

[0145] Specifically, the process of determining the operational deviation of the real-time operational status profile includes:

[0146] Compare the real-time operational status profile with the baseline operational envelope;

[0147] The absolute value of the minimum difference between several data items in the statistical operation data and the corresponding baseline data item range;

[0148] The absolute value of the minimum difference is taken as the running deviation.

[0149] The reference operating envelope is the allowed range of values ​​for the data of the communication chain under normal operating conditions.

[0150] Specifically, adaptively adjusting the scheduling intensity of the communication chain includes:

[0151] The intensity of scheduling adjustments is positively correlated with the degree of operational deviation.

[0152] It is understandable that scheduling adjustment refers to a series of operations by which the system intervenes in and corrects the data transmission rhythm, resource allocation, or processing strategy of the communication chain when the overall operating state of the data integration system deviates.

[0153] In this embodiment, no specific limitation is made on the adjustment method corresponding to the scheduling adjustment intensity. The scheduling adjustment intensity is positively correlated with the operational deviation. The greater the operational deviation, the further the current operating state of the system deviates from the normal healthy baseline, the higher the risk of anomaly, and the greater the intensity of intervention required. Conversely, the smaller the operational deviation, the closer the operating state of the system is to the normal envelope range, and the required scheduling adjustment intensity is correspondingly reduced.

[0154] For example, the adjustment method may include, but is not limited to, one or more combinations of the following: adjusting the parallelism parameters of key conversion nodes to change the concurrent processing capability of data processing tasks; adjusting the memory allocation upper limit and connection timeout threshold of each integrated component in the data pipeline to optimize resource allocation strategies; triggering the elastic scaling up or down process or traffic degradation process of the data pipeline to achieve dynamic scaling up or down of system capacity or proactive reduction of non-core traffic.

[0155] It is understandable that those skilled in the art can flexibly configure and dynamically adjust the specific selection and combination strategies of the above adjustment methods based on the actual system architecture, business importance level, current operational deviation, and real-time load status, so as to ensure that scheduling intervention actions are precisely matched with the actual needs of the system.

[0156] Specifically, electronic devices that apply computer data integration and management methods are also provided, including:

[0157] The data acquisition module is used to acquire the computer's operational data from various heterogeneous data sources and data nodes, and to build an operational status profile.

[0158] The processing evaluation module, which is connected to the data acquisition module, is used to extract processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains.

[0159] A communication analysis module, connected to the processing and evaluation module, is used to evaluate and analyze the data communication of the communication chain based on the annotation results, including:

[0160] Identify whether there are abnormal communication time domain segments within the communication chain, and assess whether there is a logic overload based on the communication mutation characteristics of the abnormal communication time domain segments;

[0161] Based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window, assess whether the communication chain has disturbance self-recovery capability;

[0162] The operational deviation of the real-time operational status profile is determined, and the scheduling adjustment intensity of the communication chain is adaptively adjusted.

[0163] It should be noted that the multiple functional modules involved in this application are only a logical division based on the functions implemented according to the present invention, and are not a strict limitation on the physical structure; in practical applications, the above functional modules can be implemented by one or more integrated circuits, a processor executing program code in memory, or a combination of the above devices.

[0164] Specifically, an apparatus is also provided, comprising:

[0165] One or more processors;

[0166] Memory;

[0167] and one or more programs;

[0168] The one or more programs are configured to be executed by one or more processors, and the memory includes a storage medium storing a computer program that, when executed by the processor, can be used to perform a computer data integration management method.

[0169] Please see Figure 4 As shown, it is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Figure 4 The computer system shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0170] like Figure 4 As shown, a computer system may include processing units (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). RAM also stores various programs and data required for the operation of the computer system. The processing units, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0171] Typically, the following devices can be connected to the I / O interface: input devices such as touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices such as liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices such as magnetic tapes, hard drives, etc.; and communication devices. Communication devices allow the computer system to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 A computer system with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0172] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A computer data integration and management method, characterized in that, include: Obtain operational data from various heterogeneous data sources and data nodes corresponding to the computer, and construct an operational status profile; Extract processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains. Based on the annotation results, the data communication of the communication chain is evaluated and analyzed. include, Identify whether there are abnormal communication time domain segments within the communication chain, and assess whether there is a logic overload based on the communication mutation characteristics of the abnormal communication time domain segments; Based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window, assess whether the communication chain has disturbance self-recovery capability; The operational deviation of the real-time operational status profile is determined, and the scheduling adjustment intensity of the communication chain is adaptively adjusted.

2. The computer data integration and management method according to claim 1, characterized in that, The process of assessing the mismatch of a communication link includes: Calculate the data processing rate ratio and the lag ratio respectively; The sum of the data processing rate ratio adjusted according to the first influence ratio and the lag ratio adjusted according to the second influence ratio is taken as the matching misscheduling; Among them, the matching features include the offset of the ratio of data processing rates between upstream and downstream nodes and the lag in data processing between upstream and downstream nodes.

3. The computer data integration and management method according to claim 2, characterized in that, The communication chain is labeled, including: If the mismatch of a communication chain is greater than or equal to a preset mismatch, then the communication chain is marked. If a communication chain is marked, the data communication of that communication chain is evaluated and analyzed.

4. The computer data integration and management method according to claim 1, characterized in that, The process of identifying communication anomalies in the time domain includes: The source-end data of the communication chain is traversed within a predetermined sliding time window to produce time-series data of the output rate. Calculate the ratio of the average rate within the sliding time window to the average rate of adjacent time periods outside the window; When the ratio exceeds the preset burst multiplier and the duration reaches the preset burst duration, the corresponding time domain segment is identified as the communication abnormal time domain segment.

5. The computer data integration and management method according to claim 4, characterized in that, Assess for logical overload, including: If the null value rate of a data field is greater than the preset null value rate, and the unexpected conversion frequency of the data type is greater than the preset conversion frequency, then a logical overload is determined to exist. Among them, communication mutation characteristics include the null value rate of data fields and the frequency of unexpected data type conversions.

6. The computer data integration and management method according to claim 1, characterized in that, The process of assessing whether the communication chain has disturbance self-recovery capability includes: Record the retry trigger events of processes within the communication chain and the corresponding retry success rate; Construct a curve showing the retry success rate as a function of the number of retries, and fit the decay slope of the curve. If the attenuation slope is greater than the preset attenuation slope, it is determined that there is a retry failure. If the difference in waiting time between two consecutive retry trigger events is consistently positive, and the increase in waiting time exceeds the preset increase, then a retry waiting anomaly is determined to exist. If there are no retry activation exceptions and no retry waiting exceptions at the same time, it is determined that the communication chain has disturbance self-recovery capability.

7. The computer data integration and management method according to claim 1, characterized in that, The process of determining the operational deviation of the real-time operational status profile includes: Compare the real-time operational status profile with the baseline operational envelope; The absolute value of the minimum difference between several data items in the statistical operation data and the corresponding baseline data item range; The absolute value of the minimum difference is taken as the running deviation. The reference operating envelope is the allowed range of values ​​for the data of the communication chain under normal operating conditions.

8. The computer data integration and management method according to claim 7, characterized in that, Adaptive adjustment of the scheduling intensity of the communication chain includes: The intensity of scheduling adjustments is positively correlated with the degree of operational deviation.

9. An electronic device employing the computer data integration and management method according to any one of claims 1-8, characterized in that, include: The data acquisition module is used to acquire the computer's operational data from various heterogeneous data sources and data nodes, and to build an operational status profile. The processing evaluation module is used to extract the processing matching features corresponding to several communication chains, evaluate the mismatch of the communication chains, and label the communication chains. The communication analysis module is used to evaluate and analyze the data communication of the communication chain based on the annotation results. include, Identify whether there are abnormal communication time domain segments within the communication chain, and assess whether there is a logic overload based on the communication mutation characteristics of the abnormal communication time domain segments; Based on the self-recovery characteristics of the communication chain within a subsequent predetermined time window, assess whether the communication chain has disturbance self-recovery capability; The operational deviation of the real-time operational status profile is determined, and the scheduling adjustment intensity of the communication chain is adaptively adjusted.

10. An apparatus, characterized in that, include: One or more processors; Memory; and one or more programs; The one or more programs are configured to be executed by one or more processors, and the memory includes a storage medium storing a computer program that, when executed by the processor, can be used to perform the computer data integration management method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Data integration method and device, electronic equipment and computer readable medium

    CN113190517A