Data processing method and device, equipment, storage medium and program product
By using the training model to predict the data transmission time and dynamically adjust the window delay time, the problem of difficult to dynamically adjust the window delay time setting in the prior art is solved, and real-time and complete data transmission in massive real-time data processing is achieved.
Patent Information
- Application Number
- CN202411599135.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the process of massive real-time data processing, streaming data may experience delay and disorderly order due to network delay and insufficient computing resources. The existing technology realizes real-time and integrity of data through window + water level line, but the setting of window delay time is difficult to dynamically adjust, affecting the real-time and integrity of data.
By obtaining the operating status information of the node server and the network status information between nodes, the training model is used to predict the data transmission time, and the window delay time is dynamically adjusted to meet the transmission needs of different streams of data and ensure the real-time and integrity of data transmission.
Dynamic adjustment of window delay time is realized, avoiding the problem of too large or too small window delay time setting, and ensuring the real-time and completeness of data transmission.
Smart Images

Figure CN119996228A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of communication technology, and in particular, relates to a data processing method, device, equipment, storage medium and program product. Background Art
[0002] With the development and popularization of the Internet industry, various data sources continue to emerge. At the same time, massive data contains huge value, which is urgently needed for enterprises to collect, analyze and mine. In the process of massive real-time data processing, stream data is usually processed according to time sequence. However, due to network delays, lack of computing resources and other reasons, stream data delays and disorder may occur. The existing technology uses the window + watermark method to achieve data real-time and integrity, but the watermark is determined by the window delay time. If the window delay time is set too large, it will affect the real-time nature of the data, and if it is set too small, it will affect the integrity. Summary of the invention
[0003] The embodiments of the present application provide a data processing method, apparatus, device, storage medium and program product, which can dynamically adjust the window delay time, meet the transmission requirements of different stream data, and ensure the real-time and integrity of data transmission.
[0004] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a cluster service, wherein the cluster service includes a plurality of node servers, and the method includes:
[0005] Obtaining the operation status information of each of the node servers and the network status information between the nodes under the preset initial delay time;
[0006] Inputting the running status information of the first node server, the running status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and outputting the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among the multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training based on the running status information of multiple source node servers, the running status information of the processing node servers corresponding to each of the source node servers, and the inter-node network status information between each of the source node servers and the corresponding processing node servers;
[0007] According to the data predicted transmission time, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server.
[0008] In a second aspect, an embodiment of the present application provides a data processing device, which is applied to a cluster service, wherein the cluster service includes a plurality of node servers, and the device includes:
[0009] A first acquisition module is used to acquire the operation status information of each node server and the network status information between nodes under a preset initial delay time;
[0010] a processing module, used for inputting the operation status information of the first node server, the operation status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and outputting the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among the multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training the operation status information of multiple source node servers, the operation status information of the processing node servers corresponding to each of the source node servers, and the inter-node network status information between each of the source node servers and the corresponding processing node servers;
[0011] An adjustment module is used to adjust the initial delay time of the first node server according to the data prediction transmission time, and determine the window delay time of the first node server.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the data processing method as described in any one of the above is implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the data processing method as described in any one of the above is implemented.
[0014] In a fifth aspect, an embodiment of the present application provides a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes any one of the data processing methods described above.
[0015] The data processing method, apparatus, device, storage medium and program product of the embodiments of the present application can obtain the operating status information and inter-node network status information of each node server in the cluster service under a preset initial delay time; and input the operating status information of the first node server, the operating status information of the second node server and the inter-node network status information of the first node server and the second node server into a preset training model, and output the data predicted transmission time of the first node server, the first node server is a node server that sends data among multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is trained by the operating status information of multiple source node servers, the operating status information of the corresponding processing node servers of each source node server, and the inter-node network status information between each source node server and the corresponding processing node server; finally, according to the data predicted transmission time, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server. In an embodiment of the present application, the data predicted transmission time of the first node server can be predicted based on the operating status information of the first node server, the operating status information of the second node server, and the node network status information between the first node server and the second node server, so as to dynamically adjust the window delay time of the first node server, so that the window delay time is not set too large or too small, and can meet the transmission requirements of different stream data, thereby ensuring the real-time and integrity of data transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 It is a schematic diagram of a window and a water level line provided in an embodiment of the present application;
[0018] Figure 2 It is a flowchart of a data processing method provided in an embodiment of the present application;
[0019] Figure 3 It is a principle diagram of a real-time streaming data fault tolerance method provided by an embodiment of the present application;
[0020] Figure 4 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0023] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0024] With the development and popularization of the Internet industry, various data sources continue to emerge. At the same time, massive data contains huge value, which is urgently needed for enterprises to collect, analyze and mine. In the process of massive real-time data processing, stream data is usually processed according to time sequence. However, due to network delays, lack of computing resources and other reasons, stream data may be delayed and disordered. Existing technologies use the window + watermark method to achieve data real-time and integrity, such as Figure 1 As shown in the figure, assuming that the allowed window delay time set by the system is 3 minutes, when event D arrives at the window, since the event time is the current maximum event time entering the window, the watermark = 9:11-3 minutes = 9:08, the watermark is within the window, so the window calculation is not triggered. When event C arrives, event C is not the current maximum event time, and the window calculation is not triggered. When time B enters, since time B is the current maximum time, the watermark = 9:15-3 minutes = 9:12, at this time the watermark is outside the window, and the window calculation triggering rule is met: the watermark >= window end time, the window immediately triggers the calculation, event A is discarded and the window is destroyed. Through the above process analysis, the watermark is determined by the window delay time. Setting the window delay time too large will affect the real-time performance of the data, and setting it too small will affect the integrity.
[0025] In order to solve the problems of the prior art, the embodiments of the present application provide a data processing method, device, equipment, storage medium and program product. The data processing method provided by the embodiments of the present application is first introduced below.
[0026] Figure 2 FIG. 1 is a flow chart of a data processing method provided by an embodiment of the present application. Figure 2 As shown, a data processing method is applied to a cluster service, and the cluster service may include multiple node servers. The method may include the following steps S201 to S203:
[0027] S201, obtaining the operation status information of each node server under the preset initial delay time and the network status information between nodes;
[0028] S202, inputting the running status information of the first node server, the running status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and outputting the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training the running status information of multiple source node servers, the running status information of the processing node servers corresponding to each source node server, and the inter-node network status information between each source node server and the corresponding processing node server;
[0029] S203: According to the data predicted transmission time, adjust the initial delay time of the first node server and determine the window delay time of the first node server.
[0030] The data processing method of the embodiment of the present application can obtain the operating status information and inter-node network status information of each node server in the cluster service under a preset initial delay time; and input the operating status information of the first node server, the operating status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and output the data predicted transmission time of the first node server. The first node server is a node server that sends data among multiple node servers, and the second node server is a node server that receives data sent by the first node server. The training model is trained by the operating status information of multiple source node servers, the operating status information of the corresponding processing node servers of each source node server, and the inter-node network status information between each source node server and the corresponding processing node server; finally, according to the data predicted transmission time, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server. In an embodiment of the present application, the data predicted transmission time of the first node server can be predicted based on the operating status information of the first node server, the operating status information of the second node server, and the node network status information between the first node server and the second node server, so as to dynamically adjust the window delay time of the first node server, so that the window delay time is not set too large or too small, and can meet the transmission requirements of different stream data, thereby ensuring the real-time and integrity of data transmission.
[0031] In S201, the above-mentioned operation status information may include information such as CPU usage, memory usage, hard disk usage and network bandwidth.
[0032] The above-mentioned inter-node network status information is the information of the network status between the source node server and the corresponding processing node server.
[0033] The above-mentioned operation status information of each node server and the network status information between nodes under the preset initial delay time are obtained. For example, the operation status information of each node server and the network status information between nodes can be obtained through monitoring software Zabbix and Nagios.
[0034] In S202, the first node server is a node server that sends data among multiple node servers.
[0035] The second node server is a node server that receives data sent by the first node server.
[0036] The above-mentioned data predicted transmission time is the data transmission time predicted by the preset training model based on the operating status information of the first node server, the operating status information of the second node server, and the node network status information between the first node server and the second node server.
[0037] The above training model is trained based on the running status information of multiple source node servers, the running status information of the processing node servers corresponding to each source node server, and the inter-node network status information between each source node server and the corresponding processing node server.
[0038] In S203, the initial delay time of the first node server is adjusted according to the data predicted transmission time, and the window delay time of the first node server is determined. Exemplarily, the initial delay time of the first node server is adjusted according to the data predicted transmission time and the following formula, and the window delay time of the first node server is determined.
[0039] td1=td0+(tp0-td0)*υ
[0040] Wherein, td0 is the initial delay time of the first node server, tp0 is the data prediction transmission time, υ is the preset moving step, and td1 is the window delay time of the first node server.
[0041] As an implementation of the present application, in order to accurately predict the data prediction transmission time of the first node server, before the above S202, the above method may further include:
[0042] Acquire a data sample set, the data sample set includes multiple data samples, each data sample includes the running status information of each source node server, the running status information of the processing node server corresponding to each source node server, the inter-node network status information between each source node server and the corresponding processing node server, and the data transmission time;
[0043] The running status information of each source node server, the running status information of each processing node server, and the network status information between each node in multiple data samples are used as independent variables of the preset model, and each data transmission time is used as the dependent variable of the preset model to train the preset model;
[0044] When the parameter error of the preset model meets the preset conditions, a training model is generated.
[0045] The data sample set includes multiple data samples, each of which includes the operating status information of each source node server, the operating status information of the processing node server corresponding to each source node server, the inter-node network status information between each source node server and the corresponding processing node server, and the data transmission time. The source node server is a node server that sends data, and the processing node server is a node server that receives data sent by the source node server.
[0046] The above data sample set, illustratively, can be as follows.
[0047]
[0048] The above-mentioned operation status information of each source node server, the operation status information of each processing node server, and the network status information between each node in the multiple data samples are used as the independent variables of the preset model, and each data transmission time is used as the dependent variable of the preset model to train the preset model. For example, the data sample set may have n data samples, and the data sample set may be represented as X, where x nm It represents the mth independent variable of the nth sample, and the parameter of the preset model is λ.
[0049]
[0050] According to the running status information of the source node server, the running status information of the processing node server and the network status information between the two nodes, there is a strong correlation with the data transmission time. Based on this, the objective function between the independent variable and the dependent variable is constructed. Assume that the independent variable x = {x1, x2, ..., x m}, the predicted value of the dependent variable is The objective function is as follows:
[0051]
[0052] Model prediction value The calculation method is the product of the independent variable matrix X and λ, which is calculated as follows:
[0053]
[0054] The parameter error of the above preset model meets the preset conditions. For example, it can be when the true value y of the sample is equal to the model prediction value When the error between the two is the smallest, the model is optimal. Find the minimum value of J(λ),
[0055]
[0056] J(λ)=(yX·λ) T (yX·λ)
[0057] Take the derivative of J(λ) so that Finally, λ is calculated according to the following formula to form a training model.
[0058] λ=(X T X) -1 X T y............. (Formula 6)
[0059] In the embodiment of the present application, the operating status information of the source node server, the operating status information of the processing node server and the network status information between the two nodes are strongly correlated with the data transmission time. The operating status information of each source node server, the operating status information of each processing node server and the network status information between the two nodes are used as independent variables of the preset model, and each data transmission time is used as the dependent variable of the preset model. The training model is obtained by training, which can accurately predict the data transmission time of the first node server.
[0060] As another implementation of the present application, in order to prevent overfitting in model training, the above-mentioned operating status information may include central processing unit usage, memory usage, hard disk usage and network bandwidth. In the above-mentioned using the operating status information of each source node server, the operating status information of each processing node server, and the network status information between each node in the multiple data samples as independent variables of the preset model, and each data transmission time as the dependent variable of the preset model, before training the preset model, the above-mentioned method may also include:
[0061] Calculate the correlation coefficient between any two parameters among the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the network status information between each node;
[0062] When the correlation coefficient is greater than a preset first threshold, any one of the two parameters corresponding to the correlation coefficient is determined as a target parameter;
[0063] The above-mentioned using the operation status information of each source node server, the operation status information of each processing node server, and the network status information between each node in the multiple data samples as the independent variables of the preset model, and each data transmission time as the dependent variable of the preset model, and training the preset model may specifically include:
[0064] The first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the target parameters in the network status information between each node are used as independent variables of the preset model, and each data transmission time is used as the dependent variable of the preset model to train the preset model.
[0065] In the above, the correlation coefficient between any two parameters is calculated in the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the network status information between each node. For example, p and q are the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and any two independent variables in the network status information between each node. The correlation coefficient ρ is calculated by the following formula pq :
[0066]
[0067] Among them, D(p) is the variance of the pth independent variable, D(q) is the variance of the qth independent variable, and Cov(p,q) is the covariance matrix between the pth and qth parameters.
[0068] The above-mentioned preset first threshold value, illustratively, may be 0.9. Of course, in the embodiment of the present application, the first threshold value is not limited thereto and may also be set according to actual needs of the user, which is not specifically limited here.
[0069] In an embodiment of the present application, the correlation coefficient between any two parameters is calculated based on the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the network status information between each node. When the correlation coefficient is greater than a preset first threshold, any parameter in the corresponding two parameters is determined as a target parameter. The larger the correlation coefficient, the stronger the correlation between the two parameters. Then, the preset model is trained based on the target parameter to prevent overfitting in the model training.
[0070] In some embodiments, the above S203 may specifically include:
[0071] According to the data prediction transmission time and the following formula, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server.
[0072] td1=td0+(tp0-td0)*υ
[0073] Wherein, td0 is the initial delay time of the first node server, tp0 is the data prediction transmission time, υ is the preset moving step, and td1 is the window delay time of the first node server.
[0074] In the embodiment of the present application, when the predicted value tp0 is greater than the existing window delay time td0, it indicates that the current data source, data processing node, and network operation status are poor, and the window delay time td0 is appropriately extended to improve the integrity of the stream data. When the predicted value tp0 is less than the existing window delay time td0, it indicates that the current data source, data processing node, and network operation status are excellent, and the window delay td0 is appropriately reduced to improve the real-time performance of the data while ensuring integrity.
[0075] As another implementation of the present application, in order to improve the fault tolerance recovery efficiency, the above method may further include:
[0076] Obtain the average data transmission time of multiple first node servers in the cluster service in the i-th window and the associated k-1 historical data average transmission time, where the k-1 data average transmission time is the data of the first node server in the k-1 windows before the i-th window, i is a positive integer, and k is a positive integer greater than 1;
[0077] Sorting average data transmission times of multiple first-node servers in the i-th window to obtain a sorted queue;
[0078] According to the average data transmission time of each first node server in the i-th window and the average transmission time of k-1 historical data, a normal distribution function of each first node server is constructed, and the normal distribution function is used to characterize the probability density of different average data transmission times taking target values;
[0079] When the average data transmission time of the first node server in the i-th window is within the preset interval of the sorting queue, or the probability density of the average data transmission time in the i-th window in the normal distribution function taking the target value is less than a preset second threshold, the first node server is determined to be a potential risk node server;
[0080] Obtain the average transmission time of j data in the i+1th to i+jth windows of each potential risk node server, where j is a positive integer greater than or equal to 1;
[0081] When the average transmission times of j data of the potential risk node server are all within the preset interval of the sorting queue, and the probability density of the target value of each of the j average transmission times of data in the normal distribution function is less than the preset second threshold, the potential risk node server is determined to be a fault risk node server, and the data of the fault risk node server and the corresponding second node server are backed up.
[0082] The preset interval of the above-mentioned sorting queue can be, for example, an interval of α (0<α<1) after all data sources in the queue sequence from small to large, wherein α can be set according to actual user needs and is not specifically limited here.
[0083] The above-mentioned situation where the average data transmission time in the i-th window of the first node server is within the preset interval of the sorting queue, or the probability density of the average data transmission time in the i-th window in the normal distribution function taking the target value is less than the preset second threshold means that the first node server has potential abnormal risk.
[0084] The above j is a positive integer greater than or equal to 1, and illustratively, can be 3 or 5, which is not specifically limited here.
[0085] The above-mentioned j average data transmission times of the potential risk node server are all within the preset interval of the sorting queue, and the probability density of the target value of each of the j average data transmission times in the normal distribution function is less than the preset second threshold, which means that the potential risk node server is likely to have a failure risk.
[0086] In an embodiment of the present application, potential abnormal risks are identified through the average data transmission time of the first node server in the i-th window and the associated k-1 average historical data transmission times, and when the average data transmission time of the first node server in the i-th window is within the preset interval of the sorting queue, or the probability density of the average data transmission time in the i-th window in the normal distribution function taking a target value is less than a preset second threshold, the failure risk is further identified based on the average transmission time of j data in the i+1 to i+j windows of the potential risk node server, so as to achieve micro-backup in advance before the failure occurs, thereby improving the fault-tolerant recovery efficiency.
[0087] As another implementation of the present application, in order to improve the fault tolerance recovery efficiency, the above method may further include:
[0088] Acquire multiple trend coefficients of each second node server in the cluster service in multiple consecutive windows, each trend coefficient being used to characterize a change trend of an average data transmission time of each second node server in a corresponding window;
[0089] When multiple trend coefficients of each second node server are greater than a preset third threshold, the second node server is determined to be a failure risk second node server, and data of the failure risk second node server and the corresponding first node server are backed up.
[0090] The above-mentioned preset third threshold value can be 0, for example. In the embodiment of the present application, it is not limited to this and can also be set according to the actual needs of the user, and no specific limitation is made here.
[0091] The above trend coefficient can be used to characterize the change trend of the average data transmission time of the second node server in the corresponding window. For example, if the trend coefficient is greater than 0, it means that the data transmission time trend of the current data processing node is increasing. If the trend coefficient is less than 0, it means that the data transmission time trend of the current data processing node is decreasing.
[0092] In an embodiment of the present application, by obtaining multiple trend coefficients of each second node server in a cluster service in multiple consecutive windows, and judging that the multiple trend coefficients of the second node server are all greater than a preset third threshold, when the trend coefficient indicates that the trend of increasing average data transmission time of the second node server occurs multiple times in a row, it indicates that the second node server has a greater probability of failure, and micro-backup is performed in advance to improve fault-tolerant recovery efficiency.
[0093] In some embodiments, the obtaining of multiple trend coefficients of each second node server in the cluster service in multiple consecutive windows may specifically include:
[0094] Obtain the average data transmission time of each second node server in the nth window and the average transmission time of m-1 historical data in the cluster service, where the m-1 historical data average transmission time is the average historical data transmission time of the second node server in the m-1 windows before the nth window, where n is a positive integer and m is a positive integer greater than 1;
[0095] According to the average data transmission time of the nth window of each second node server and the average transmission time of m-1 historical data, the trend coefficient of each second node server in the nth window is calculated according to the following formula:
[0096]
[0097] Where G is the trend coefficient, is the average data transmission time of the nth window and the average transmission time of m-1 historical data, and w represents the second node server.
[0098] In an embodiment of the present application, the trend coefficient of the second node server in the nth window can be accurately obtained through the average data transmission time of the second node server in the nth window and the average transmission time of m-1 historical data, thereby improving the accuracy of second node server failure risk identification.
[0099] In order to facilitate the understanding of the data processing method in the embodiment of the present application, the actual application process of this data processing method is described as follows:
[0100] (I) A method for dynamically adjusting window delay time based on data source differentiation
[0101] This application aims to address the shortcomings of the lack of scientific evaluation and setting of window delay time, and the inability to adaptively adjust the window delay time according to the real-time operation status of the data source. A method for dynamically adjusting the window delay time based on data source differentiation is proposed. The specific method is as follows:
[0102] 1.1. Select independent and dependent variables
[0103] Taking the data source, the running status of the data processing node and the network status between the two as independent variables, and the data transmission time (i.e., the event transmission time below) as the dependent variable, the model is trained for each data source and data processing node separately, and the window delay time of each data source is set differently.
[0104] (1) First, the node operation status is obtained through the monitoring software, as shown in Table 1:
[0105] Table 1: Training data samples
[0106]
[0107]
[0108] (2) To prevent overfitting during model training, a correlation analysis is performed between any two parameters in Table 1, and the correlation coefficient of any parameter is calculated. Assume that p and q are any two independent variables, as shown in Formula 1:
[0109]
[0110] Among them, as shown in Table 1, D(p) is the variance of the pth independent variable, D(q) is the variance of the qth independent variable, and Cov(p,q) is the covariance matrix between the pth and qth parameters.
[0111] According to formula 1, the correlation between any two parameters is calculated and the correlation coefficient ρ pq Randomly select one of the parameters greater than 0.9, ρ pq The larger the value, the stronger the correlation between the parameters. m independent variables (equivalent to the target parameters mentioned above) are selected from the 9 parameters.
[0112] 1.2 Training Model
[0113] This application believes that the operating status of the data source, the operating status of the data processing node, and the network status between the two are strongly correlated with the time transmission time. Based on this, the objective function between the independent variable and the dependent variable is constructed. Assume that the independent variable x = {x1, x2, ..., x m}, the predicted value of the dependent variable is The objective function is as shown in formula 2:
[0114]
[0115] Assume that the model is trained with n data samples, and the data sample set can be represented as X, where x nm It represents the mth independent variable of the nth sample, and the model parameter is λ.
[0116]
[0117] Model prediction value The calculation method is the product of the independent variable matrix X and λ, as shown in Formula 3:
[0118]
[0119] When the error between the true value y of the sample and the model prediction value is the smallest, the model is optimal, which is converted into finding the minimum value of J(λ).
[0120]
[0121] J(λ)=(yX·λ) T (yX·λ)........... (Formula 5)
[0122] Take the derivative of J(λ) so that λ is calculated according to formula 6 to form a model.
[0123] λ=(X T X) -1 X T y............. (Formula 6)
[0124] 1.3. Predicting data arrival time
[0125] Substitute the operating status of the data source (equivalent to the above-mentioned first node server) of the current window, the operating status of the data processing node (equivalent to the above-mentioned second node server) and the network conditions between the two into Formula 3 to obtain the predicted data transmission time.
[0126] 1.4. Adaptive adjustment of window delay time
[0127] Assume that the initial window delay (equivalent to the above initial delay time) is td0, the predicted value according to the current state is tp0, and the moving step is υ (can be set by yourself). When the predicted value is greater than the existing window delay time, it indicates that the current data source, data processing node, and network operation status are poor. The window delay time should be appropriately extended to improve the integrity of the stream data. When the predicted value is less than the existing window delay time, it indicates that the current data source, data processing node, and network operation status are excellent. The window delay should be appropriately reduced to ensure the integrity while improving the real-time performance of the data. The specific adaptive iterative algorithm is as shown in Formula 7:
[0128] td1=td0+(tp0-td0)*υ......(Formula 7)
[0129] According to Formula 7, when each window is calculated, the window delay time is adaptively adjusted to obtain the optimal window delay time for the data source.
[0130] 1.5. Differentiated calculation window delay time
[0131] Repeat steps 1, 2, 3, and 4 above. According to the situation between each data source and the data processing node, each data source maintains a separate window delay time. Different data sources have differentiated window delay times. Customize the window calculation trigger conditions for different data sources to improve the real-time and integrity of real-time stream data processing.
[0132] In an embodiment of the present application, a method for dynamically adjusting the window delay time for data sources can be proposed in view of the shortcomings of the window delay time setting of the prior art, which lacks scientific evaluation and cannot be adaptively adjusted according to the real-time operating status of the data source. First, for different data sources and data processing nodes, the data source, the operating status of the data processing node and the network status between the two are used as independent variables, and the data transmission time is used as the dependent variable to train the model. Secondly, the current data source, the state of the data processing node and the network status between the two are substituted into the model to predict the window delay time. Finally, according to the prediction results, the window delay time of different data sources is adaptively and dynamically adjusted to solve the shortcomings of the traditional method that cannot adaptively adjust the window delay time. Ensure that when the computing and network resources are sufficient, the window delay time is reduced to improve the real-time performance of the data. When the computing and network resources are limited, the window delay time is appropriately increased to ensure the integrity of the streaming data.
[0133] (II) A method for improving stream processing fault tolerance through micro-snapshots
[0134] Another requirement for massive real-time data processing is strong fault tolerance and fast fault recovery when a fault occurs. Existing methods for real-time streaming data fault tolerance: When processing real-time streaming data, a distributed snapshot checkpoint mark is periodically generated and inserted into the data stream. When the data processing node detects the mark, the distributed snapshot is immediately backed up. Figure 3As shown in the figure, when the data processing node receives checkpoint n-1, it backs up and generates a snapshot of data A, B, and C on the right side of checkpoint n-1. When a fault occurs at point E, the data is restored to the state at checkpoint n-1, and incremental processing is performed based on this snapshot state to reprocess D, E, and F. However, if the backup cycle is too long, the fault-tolerant recovery efficiency will be low if the backup cycle is too short. If the backup cycle is too short, computing power and storage resources will be wasted. There is a lack of a method to analyze potential faults and back up in advance based on the data source and data processing node operation trends to improve fault-tolerant efficiency.
[0135] For the existing fault-tolerant scheme for periodic backup of distributed snapshots, if the backup cycle is too long, the fault-tolerant recovery efficiency will be low. If the backup cycle is too short, computing power and storage resources will be wasted. There is a lack of a method based on the data source and data processing node operation trend to analyze potential failures and back up in advance. This application proposes a fault-tolerant method that combines periodic + micro-snapshot to improve the efficiency of fault-tolerant recovery. During the stream processing process, before the task node fails, the data transmission processing time will usually be abnormal. To address this phenomenon, this application detects anomalies from two aspects: data source and data processing node. When the trigger conditions are met, a micro-backup is performed in advance before the failure occurs, thereby improving the efficiency of fault-tolerant recovery. The specific anomaly detection method is as follows:
[0136] 2.1. Data source anomaly detection
[0137] In terms of data source anomaly detection, it is measured from two dimensions. The first is the comparison of data transmission time between multiple data sources. When the average data transmission time of a certain data source (equivalent to the first node server mentioned above) ranks behind all data sources by α (0<α<1), it indicates that the data source has potential abnormality risks. The second is the comparison of data transmission time of the data source time series window. When the data transmission time of a certain window is at the tail β of the normal distribution compared with the previous k-1 windows (0<β<1), it also has potential abnormality risks. Combining the above two dimensions, when a certain data source is compared with the latter α and compared with its own previous k-1 time series windows, it is at the tail β of the normal distribution, and three windows appear in succession, it indicates that there is a high probability of failure risk and micro-snapshot backup is performed in advance. The specific process is as follows:
[0138] (1) Assume that the system has l data sources and the average data transmission time in the i-th window is tm i ={tm i1 ,tm i2 ,...,tm il}, when the data transmission time is ranked after α, it has potential abnormal risk.
[0139] (2) Assuming a data source j, the average transmission time between time window i and the previous k-1 windows is expressed as According to tr jConstruct a normal distribution function, when P(tr j >γ)<β, it indicates that the data source has potential abnormal risk.
[0140] Combining the above two aspects, when a data source is ranked at the back α compared with other data sources and is located at the tail β of the normal distribution compared with its previous k-1 time series windows, this situation occurs in three consecutive windows, indicating that the data source is likely to fail and a micro-snapshot backup is performed in advance.
[0141] 2.2. Data processing node anomaly detection
[0142] In terms of data processing node fault detection, within the past m-1 window, if the trend of data transmission time increases for n consecutive times, it indicates that the data processing node will have a certain probability of failure. The specific detection method is as follows:
[0143] Assume that the average transmission time of data from a data source w in m-1 windows before time window n is
[0144]
[0145] Assumptions Where n≥s>d>nm
[0146] calculate
[0147]
[0148] If G>0, it means that the data transmission time trend of the current data processing node is increasing.
[0149] If G<0, it means that the data transmission time trend of the current data processing node is decreasing.
[0150] When the data transmission time increases for η (adjustable) consecutive times, it indicates that the data processing node has a high probability of failure. Micro-backup is performed in advance to improve the efficiency of fault-tolerant recovery.
[0151] In the embodiment of the present application, a fault-tolerant method combining periodic snapshots and micro-snapshots can be proposed for the low efficiency of existing real-time data fault tolerance and the lack of a method based on the data source and the running trend of the data processing node to analyze potential faults and back up in advance. In terms of data source detection, on the one hand, from the comparison of its own time series window dimension, when the data transmission time ranks behind the normal distribution by β, the data source is suspected to be abnormal. On the other hand, when the data source is compared horizontally, when the data transmission time is relatively behind α among the data sources, an abnormality is suspected. When three consecutive windows meet the above two conditions at the same time, micro-backup is performed in advance. In terms of data processing node abnormality detection, within the past m-1 windows, if the trend of data transmission time increases for η times in a row, it indicates that the data processing node may be abnormal. When the trend increases for η times in a row, micro-backup is performed in advance. Through the above method of combining periodic + micro-backup, on the basis of periodic snapshot backup, combined with micro-backup based on fault prediction, the efficiency of fault-tolerant recovery can be greatly improved.
[0152] Based on the data processing method provided in the above embodiment, the present application also provides a specific implementation of a data processing device. Please refer to the following embodiment.
[0153] like Figure 4 As shown, the data processing device 400 provided in the embodiment of the present application is applied to a cluster service, and the cluster service includes multiple node servers. The device 400 may include the following modules: a first acquisition module 401, a processing module 402 and an adjustment module 403.
[0154] The first acquisition module 401 is used to acquire the operation status information of each node server and the network status information between nodes under the preset initial delay time;
[0155] The processing module 402 is used to input the operation status information of the first node server, the operation status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and output the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training the operation status information of multiple source node servers, the operation status information of the processing node servers corresponding to each source node server, and the inter-node network status information between each source node server and the corresponding processing node server;
[0156] The adjustment module 403 is used to adjust the initial delay time of the first node server according to the data prediction transmission time, and determine the window delay time of the first node server.
[0157] The data processing device of the embodiment of the present application can obtain the operating status information and inter-node network status information of each node server in the cluster service under a preset initial delay time; and input the operating status information of the first node server, the operating status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and output the data predicted transmission time of the first node server. The first node server is a node server that sends data among multiple node servers, and the second node server is a node server that receives data sent by the first node server. The training model is trained by the operating status information of multiple source node servers, the operating status information of the corresponding processing node servers of each source node server, and the inter-node network status information between each source node server and the corresponding processing node server; finally, according to the data predicted transmission time, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server. In an embodiment of the present application, the data predicted transmission time of the first node server can be predicted based on the operating status information of the first node server, the operating status information of the second node server, and the node network status information between the first node server and the second node server, so as to dynamically adjust the window delay time of the first node server, so that the window delay time is not set too large or too small, and can meet the transmission requirements of different stream data, thereby ensuring the real-time and integrity of data transmission.
[0158] As an implementation of the present application, in order to accurately predict the data prediction transmission time of the first node server, the above-mentioned device 400 may also include:
[0159] A second acquisition module is used to acquire a data sample set, the data sample set includes multiple data samples, each data sample includes the operation status information of each source node server, the operation status information of the processing node server corresponding to each source node server, the inter-node network status information between each source node server and the corresponding processing node server, and the data transmission time;
[0160] A training module, used to train the preset model by taking the operation status information of each source node server, the operation status information of each processing node server, and the network status information between each node in multiple data samples as independent variables of the preset model, and each data transmission time as the dependent variable of the preset model;
[0161] The generation module is used to generate a training model when the parameter error of the preset model meets the preset conditions.
[0162] As another implementation of the present application, in order to prevent overfitting in model training, the above-mentioned operation status information may include central processing unit usage, memory usage, hard disk usage and network bandwidth, and the above-mentioned device 400 may also include:
[0163] A calculation module, used to calculate the correlation coefficient between any two parameters among the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the network status information between each node;
[0164] A first determination module, configured to determine any one of the two parameters corresponding to the correlation coefficient as a target parameter when the correlation coefficient is greater than a preset first threshold;
[0165] The above-mentioned training module is also used to train the preset model by taking the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the target parameters in the network status information between each node as independent variables of the preset model, and each data transmission time as the dependent variable of the preset model.
[0166] In some embodiments, the adjustment module 403 is specifically used to adjust the initial delay time of the first node server according to the data prediction transmission time and the following formula to determine the window delay time of the first node server.
[0167] td1=td0+(tp0-td0)*υ
[0168] Wherein, td0 is the initial delay time of the first node server, tp0 is the data prediction transmission time, υ is the preset moving step, and td1 is the window delay time of the first node server.
[0169] As another implementation of the present application, in order to improve the fault tolerance recovery efficiency, the above-mentioned device 400 may further include:
[0170] The third acquisition module is used to obtain the average data transmission time of multiple first node servers in the cluster service in the i-th window and the average transmission time of k-1 historical data, where the k-1 historical data average transmission time is the average data transmission time of the first node server in the k-1 windows before the i-th window, i is a positive integer, and k is a positive integer greater than 1;
[0171] A sorting module, used to sort the average data transmission time of multiple first node servers in the i-th window to obtain a sorting queue;
[0172] A construction module is used to construct a normal distribution function of each first node server according to the average data transmission time of each first node server in the i-th window and the average transmission time of k-1 historical data, and the normal distribution function is used to characterize the probability density of different average data transmission times taking target values;
[0173] A second determination module is used to determine that the first node server is a potential risk node server when the average data transmission time of the first node server in the i-th window is within a preset interval of the sorting queue, or the probability density of the average data transmission time in the i-th window in the normal distribution function taking a target value is less than a preset second threshold;
[0174] A fourth acquisition module is used to obtain the average transmission time of j data in the i+1th to i+jth windows of each potential risk node server, where j is a positive integer greater than or equal to 1;
[0175] The first backup module is used to determine that the potential risk node server is a fault risk node server and back up the data of the fault risk node server and the corresponding second node server when the average transmission times of j data of the potential risk node server are all within the preset interval of the sorting queue and the probability density of the average transmission time of each of the j data in the normal distribution function taking the target value is less than the preset second threshold.
[0176] As another implementation of the present application, in order to improve the fault tolerance recovery efficiency, the above-mentioned device 400 may further include:
[0177] A fifth acquisition module is used to acquire multiple trend coefficients of each second node server in the cluster service in multiple consecutive windows, each trend coefficient is used to characterize the change trend of the average data transmission time of each second node server in the corresponding window;
[0178] The second backup module is used to determine that the second node server is a fault-risk second node server when multiple trend coefficients of each second node server are greater than a preset third threshold, and to back up data of the fault-risk second node server and the corresponding first node server.
[0179] In some embodiments, the fifth acquisition module may specifically include:
[0180] An acquisition unit, used to acquire the average data transmission time of each second node server in the nth window and the average transmission time of m-1 historical data in the cluster service, where the m-1 historical data average transmission time is the average data transmission time of the second node server in the m-1 windows before the nth window, where n is a positive integer and m is a positive integer greater than 1;
[0181] The calculation unit is used to calculate the trend coefficient of each second node server in the nth window according to the average data transmission time of the nth window of each second node server and the average transmission time of m-1 historical data according to the following formula:
[0182]
[0183] Where G is the trend coefficient, is the average data transmission time of the nth window and the average transmission time of m-1 historical data, and w represents the second node server.
[0184] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0185] The electronic device may include a processor 501 and a memory 502 storing computer program instructions.
[0186] Specifically, the processor 501 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0187] The memory 502 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 502 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 502 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 502 is a non-volatile solid-state memory.
[0188] In certain embodiments, the memory 502 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0189] The processor 501 implements any one of the data processing methods in the above embodiments by reading and executing computer program instructions stored in the memory 502 .
[0190] In one example, the electronic device may further include a communication interface 503 and a bus 510. Figure 5 As shown, the processor 501, the memory 502, and the communication interface 503 are connected via a bus 510 and communicate with each other.
[0191] The communication interface 503 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0192] Bus 510 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 510 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0193] The electronic device can execute the data processing method in the embodiment of the present application, thereby realizing the combination Figure 2 and Figure 4 The data processing method and device described.
[0194] In addition, in combination with the data processing method in the above embodiments, the present application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the data processing methods in the above embodiments is implemented.
[0195] In combination with the data processing method in the above embodiments, an embodiment of the present application may provide a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes any one of the data processing methods above.
[0196] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0197] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0198] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0199] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0200] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A data processing method, characterized in that: Applied to a cluster service, the cluster service includes a plurality of node servers, and the method includes: Obtaining the operation status information of each of the node servers and the network status information between the nodes under the preset initial delay time; Inputting the running status information of the first node server, the running status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and outputting the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among the multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training based on the running status information of multiple source node servers, the running status information of the processing node servers corresponding to each of the source node servers, and the inter-node network status information between each of the source node servers and the corresponding processing node servers; According to the data predicted transmission time, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server.
2. The method according to claim 1, characterized in that Before inputting the operation status information of the first node server, the operation status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model and outputting the data predicted transmission time of the first node server, the method further includes: Acquire a data sample set, the data sample set comprising a plurality of data samples, each of the data samples comprising operation status information of each of the source node servers, operation status information of a processing node server corresponding to each of the source node servers, inter-node network status information between each of the source node servers and the corresponding processing node servers, and data transmission time; The running status information of each source node server, the running status information of each processing node server, and the network status information between each node in the multiple data samples are used as independent variables of the preset model, and each data transmission time is used as the dependent variable of the preset model to train the preset model; When the parameter error of the preset model meets the preset conditions, the training model is generated.
3. The method according to claim 2, characterized in that The operation status information includes the CPU usage rate, the memory usage rate, the hard disk usage rate and the network bandwidth. Before taking the operation status information of each source node server, the operation status information of each processing node server and the network status information between each node in the multiple data samples as the independent variables of the preset model and each data transmission time as the dependent variable of the preset model, the method further includes: Calculate the correlation coefficient between any two parameters among the first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each source node server, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each processing node server, and the network status information between each node; When the correlation coefficient is greater than a preset first threshold, determining any one of the two parameters corresponding to the correlation coefficient as a target parameter; The method uses the operation status information of each source node server, the operation status information of each processing node server, and the network status information between each node in the multiple data samples as independent variables of the preset model, and each data transmission time as the dependent variable of the preset model to train the preset model, including: The first CPU usage rate, the first memory usage rate, the first hard disk usage rate and the first network bandwidth of each of the source node servers, the second CPU usage rate, the second memory usage rate, the second hard disk usage rate and the second network bandwidth of each of the processing node servers, and the target parameters in the network status information between each of the nodes are used as independent variables of the preset model, and each of the data transmission times is used as the dependent variable of the preset model to train the preset model.
4. The method according to claim 1, characterized in that: The adjusting the initial delay time of the first node server according to the data predicted transmission time and determining the window delay time of the first node server includes: According to the data predicted transmission time and the following formula, the initial delay time of the first node server is adjusted to determine the window delay time of the first node server, td1=td0+(tp0-td0)*υ Among them, td0 is the initial delay time of the first node server, tp0 is the data prediction transmission time, υ is the preset moving step, and td1 is the window delay time of the first node server.
5. The method according to claim 1, characterized in that The method further comprises: Obtain the average data transmission time of the first node servers in the cluster service in the i-th window and the average transmission time of k-1 historical data, wherein the k-1 historical data average transmission time is the average data transmission time of the first node server in k-1 windows before the i-th window, where i is a positive integer and k is a positive integer greater than 1; sorting the average data transmission times of the plurality of first node servers in the i-th window to obtain a sorted queue; According to the average data transmission time of each first node server in the i-th window and the average transmission time of the k-1 historical data, a normal distribution function of each first node server is constructed, and the normal distribution function is used to characterize the probability density of different average data transmission times taking target values; When the average data transmission time of the first node server in the i-th window is within the preset interval of the sorting queue, or the probability density of the average data transmission time in the i-th window in the normal distribution function taking the target value is less than a preset second threshold, the first node server is determined to be a potential risk node server; Obtaining the average transmission time of j data in the i+1th to i+jth windows of each of the potential risk node servers, where j is a positive integer greater than or equal to 1; When the j average data transmission times of the potential risk node server are all within the preset interval of the sorting queue, and the probability density of each of the j average data transmission times in the normal distribution function taking the target value is less than the preset second threshold, the potential risk node server is determined to be a fault risk node server, and the data of the fault risk node server and the corresponding second node server are backed up.
6. The method according to claim 1, characterized in that The method further comprises: Acquire multiple trend coefficients of each second-node server in the cluster service in multiple consecutive windows, each trend coefficient being used to characterize a change trend of an average data transmission time of each second-node server in a corresponding window; When multiple trend coefficients of each second node server are greater than a preset third threshold, the second node server is determined to be a failure-risk second node server, and data of the failure-risk second node server and the corresponding first node server are backed up.
7. The method according to claim 6, characterized in that The obtaining of multiple trend coefficients of each second node server in the cluster service in multiple consecutive windows includes: Obtain the average data transmission time of each second node server in the cluster service in the nth window and the average transmission time of m-1 historical data, where the m-1 historical data average transmission time is the average data transmission time of the second node server in the m-1 windows before the nth window, where n is a positive integer and m is a positive integer greater than 1; According to the average data transmission time of the nth window of each second node server and the average transmission time of the m-1 historical data, the trend coefficient of each second node server in the nth window is calculated according to the following formula: Wherein, G is the trend coefficient, is the average data transmission time of the nth window and the average transmission time of the m-1 historical data, and w represents the second node server.
8. A data processing device, characterized in that: Applied to a cluster service, the cluster service includes a plurality of node servers, and the device includes: A first acquisition module is used to acquire the operation status information of each node server and the network status information between nodes under a preset initial delay time; a processing module, used for inputting the operation status information of the first node server, the operation status information of the second node server, and the inter-node network status information of the first node server and the second node server into a preset training model, and outputting the predicted data transmission time of the first node server, wherein the first node server is a node server that sends data among the multiple node servers, and the second node server is a node server that receives data sent by the first node server, and the training model is obtained by training the operation status information of multiple source node servers, the operation status information of the processing node servers corresponding to each of the source node servers, and the inter-node network status information between each of the source node servers and the corresponding processing node servers; An adjustment module is used to adjust the initial delay time of the first node server according to the data prediction transmission time, and determine the window delay time of the first node server.
9. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the data processing method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.
11. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data collaboration processing method, system, device and equipment and storage medium
CN114860426A
Equipment state detection method, computer equipment and storage medium
CN115935193A
Data processing method, device and equipment and computer storage medium
CN117135085A
Dynamic load balancing method of server and related equipment
CN118540326A
Task allocation method and device, electronic equipment and computer program
CN118796441A