Method and system for analyzing interference among multiple computing tasks based on network communication behaviors

By collecting network communication traffic through a bypass, a baseline of network performance indicators and correlation coefficients are constructed, and interference between tasks is quantified. This solves the problems of unstable job performance and unfair resources in multi-user environments, realizes automatic identification and quantification of interference, and supports resource fairness assessment.

CN121967280APending Publication Date: 2026-05-01BEIJING WANGSHEN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING WANGSHEN TECH CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In multi-user computing environments, existing technologies struggle to identify and quantify interference between computing jobs, leading to unstable job performance and unfair resource utilization. Traditional monitoring methods are unable to pinpoint the causes of performance degradation.

Method used

By collecting network communication traffic through bypass, parsing data packet information, constructing historical baselines for network performance indicators, calculating indicator deviations and comprehensive correlation coefficients, quantifying the interference intensity and index of computational tasks, and identifying interference between computational tasks.

Benefits of technology

It enables automatic identification of multi-user job interference and quantification of interference levels without intruding on computing nodes, supports resource fairness assessment and scheduling optimization, and reduces the difficulty of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967280A_ABST
    Figure CN121967280A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for analyzing interference among multiple computing tasks based on network communication behaviors, and belongs to the field of network management and high-performance computing control. The method comprises the following steps of: carrying out bypass acquisition and analysis on network communication flow in a multi-user computing task environment, and carrying out computing task identification; counting a network performance index (NPM) of the network communication flow; constructing a historical baseline of the network performance indexes, calculating network performance index deviations of the corresponding calculation tasks based on the historical baseline, and constructing an index deviation time sequence; based on the index deviation time sequence, performing synchronism analysis on the calculation tasks, calculating a comprehensive correlation coefficient, and calculating an interference intensity item of each calculation task; and calculating the interference index of each calculation task according to the comprehensive correlation coefficient and the interference intensity item. According to the method, the interference degree of different computing tasks in the multi-user computing environment is quantitatively expressed, and fairness evaluation and abnormal access behavior recognition are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

A Method and System for Interference Analysis among Multiple Computational Tasks Based on Network Communication Behavior Technical Field

[0001] This invention belongs to the field of network management and high-performance computing control, and specifically relates to a method and system for interference analysis between multiple computing tasks based on network communication behavior. Background Technology

[0002] With the widespread deployment of high-performance computing clusters, artificial intelligence training platforms, and cloud computing resources, multiple users can share the same computing and network infrastructure, saving resources and computing costs. However, while sharing is achieved, different users' computing jobs often run concurrently within the same time period. When sharing network resources such as switches, links, and buffer queues, problems such as unstable job performance, difficulty in locating job interference, and lack of quantitative and fairness evaluation mechanisms often occur. For example, during the execution of some user jobs, although the utilization of node CPU, GPU, and memory is not significantly abnormal, unstable phenomena such as sudden increases in communication phase time, longer synchronization waiting time, and slower distributed training iteration may occur. At the same time, since multiple users share network paths, when some jobs generate sudden or micro-bursts of traffic, it may lead to a decrease in the communication quality of other users' jobs, but existing monitoring methods are difficult to determine whether performance interference is caused by other users' jobs. In addition, traditional network monitoring is mostly based on port bandwidth, packet loss counters, or link utilization, which cannot clearly attribute performance degradation to specific users, nor can it quantify and evaluate the fairness of resource use in multi-user environments. Summary of the Invention

[0003] In view of the above-mentioned defects or deficiencies in the prior art, the present invention aims to provide a method and system for interference analysis between multiple computing tasks based on network communication behavior. In a multi-user computing environment, by passively collecting network communication data, the communication behavior of different user computing jobs is identified, correlated and the degree of interference is quantified. There is no need to deploy probes on computing nodes. Multi-user job mutual interference can be identified based solely on bypass traffic, and the degree of interference can be quantified, thereby realizing the fairness assessment of multi-user job operation and the identification of abnormal access behavior.

[0004] To achieve the above objectives, the embodiments of the present invention adopt the following technical solution: Firstly, the embodiments of the present invention provide an interference analysis method for multiple computing tasks based on network communication behavior. The method includes: bypassing the collection of network communication traffic in a multi-user computing task environment; parsing the collected network communication traffic and identifying computing tasks based on the parsed information; statistically analyzing the network performance index (NPM) of the network communication traffic corresponding to each computing task; constructing a historical baseline for the network performance index of each computing task, calculating the deviation of the network performance index of the corresponding computing task based on the historical baseline, and constructing a time series of the index deviation; performing synchronization analysis on the computing tasks based on the index deviation time series and calculating a comprehensive correlation coefficient; calculating the interference intensity term of each computing task based on the index deviation time series; and calculating the interference index of each computing task based on the comprehensive correlation coefficient and the interference intensity term.

[0005] As a preferred embodiment of the present invention, the identification of computing tasks based on parsed information specifically includes: parsing network communication traffic to obtain data packet timestamps, source addresses, destination addresses, and transport layer session identifier information; dividing network traffic based on communication sessions according to the parsed source addresses, destination addresses, and transport layer session identifiers; determining the user and computing task to which each communication session belongs based on the communication topology characteristics or scheduling mapping information of each communication session; and aggregating communication sessions generated by the same computing task to obtain the network communication traffic corresponding to each computing task.

[0006] As a preferred embodiment of the present invention, the network performance index NPM includes round-trip time (RTT), packet loss probability, retransmission strength, retransmission packet ratio, duplicate sequence number occurrence ratio, and abnormal confirmation or out-of-order behavior ratio.

[0007] As a preferred embodiment of the present invention, the statistical network performance index (NPM) specifically includes: setting a preset time window and dividing the communication traffic of computing tasks into multiple consecutive time windows W; and statistically analyzing the NPM of all computing tasks in each time window = { }

[0008] As a preferred embodiment of the present invention, when the network computing task is round-trip time (RTT), the median or high quantile of the RTT within the time window W, or the RTT jitter amplitude, is used for characterization.

[0009] In a preferred embodiment of the present invention, the calculation formula for the network performance index deviation of the corresponding computing task based on the historical baseline is as follows: In equation (1), It is the j-th network performance metric value of task t within the time window W; It is the historical baseline of the j-th network performance metric for computation task t within the time window W; It is the deviation of the network performance metric for task t within the time window W; the deviation of the network performance metric across all time windows W. The time series of index deviations constituting each computational task { }

[0010] In a preferred embodiment of the present invention, the step of performing synchronization analysis on computing tasks and calculating a comprehensive correlation coefficient specifically includes: constructing a reference object; calculating the correlation coefficient for each indicator based on the reference object; performing directional constraint processing on the correlation coefficients to retain positive correlations that deteriorate synchronously; and fusing the correlation coefficients of all indicators for each computing task to obtain a comprehensive correlation coefficient for the computing task.

[0011] As a preferred embodiment of the present invention, the calculation of the interference intensity term for each calculation task specifically includes: performing directional constraint processing on the index deviation to obtain a positive index deviation; normalizing the magnitude of the positive index deviation; and fusing the calculation of the interference intensity term.

[0012] In a preferred embodiment of the present invention, the interference index of each computational task is calculated using the following formula: In equation (9), It represents the comprehensive correlation coefficient between the computation task t and the reference object in terms of network performance index deviation, and is used to reflect whether network performance anomalies occur synchronously in the time dimension; This represents the overall intensity of network performance degradation for computation task t within the current time window; This represents the network interference index of task t within the current time period; clip(·) represents the range constraint function, used to limit the calculation results to a preset range.

[0013] Secondly, embodiments of the present invention provide an interference analysis system for multiple computing tasks based on network communication behavior. The system includes: a data acquisition module, a task identification module, an NPM statistics module, an index deviation calculation module, a correlation coefficient calculation module, an interference intensity term calculation module, and an interference index calculation module. Specifically, the data acquisition module is used to collect network communication traffic in a multi-user computing task environment via bypass; the task identification module is used to parse the collected network communication traffic and identify computing tasks based on the parsed information; the NPM statistics module is used to calculate the network performance index (NPM) of the network communication traffic corresponding to each computing task; the index deviation calculation module is used to construct a historical baseline of the network performance index for each computing task, calculate the network performance index deviation of the corresponding computing task based on the historical baseline, and construct an index deviation time series; the correlation coefficient calculation module is used to perform synchronization analysis on the computing tasks based on the index deviation time series and calculate a comprehensive correlation coefficient; the interference intensity term calculation module is used to calculate the interference intensity term for each computing task based on the index deviation time series; and the interference index calculation module is used to calculate the interference index for each computing task based on the comprehensive correlation coefficient and the interference intensity term.

[0014] The technical solution provided by the embodiments of the present invention has the following beneficial effects: the interference analysis method and system for multiple computing tasks based on network communication behavior does not require the deployment of probes on computing nodes and is applicable to large-scale HPC clusters; it can identify low-proportion, short-term but significantly impactful mutual interference behaviors; it transforms mutual interference from empirical judgment into quantifiable indicators; it supports resource fairness assessment and scheduling optimization decisions in multi-user environments; and it provides clear evidence of job mutual interference for operation and maintenance personnel, reducing the difficulty of troubleshooting.

[0015] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 is a flowchart of the interference analysis method between multiple computing tasks based on network communication behavior according to an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can also be combined with each other.

[0019] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the embodiments of the present invention, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. In addition, sometimes a subscript such as W1 may be written in a non-subscript form such as W1, and their meanings are consistent unless the distinction is emphasized.

[0020] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0021] This invention provides a method and system for interference analysis between multiple computing tasks based on network communication behavior. Based on bypass-collected network communication data, the system statistically analyzes the round-trip latency and packet loss behavior of different user jobs according to time windows, compares this data with historical baselines, and further analyzes the correlation between round-trip latency deviation and packet loss deviation over time. A computing task interference index is constructed to characterize the degree of network interference between multiple computing tasks. This allows for the identification, correlation, and quantification of the interference level of communication behaviors of different user computing tasks within the same time period. Without intruding on computing nodes or relying on job logs or probes, the system achieves automatic identification of multi-user computing tasks, quantitative description of abnormal network communication performance, identification of mutual interference behavior among computing tasks in a multi-user environment, and quantitative assessment of the fairness of multi-user resource usage through passive analysis of network communication traffic.

[0022] As shown in Figure 1, the interference analysis method between multiple computing tasks based on network communication behavior includes the following steps: Step S1, bypassing the network communication traffic in a multi-user computing task environment.

[0023] In this step, the multi-user computing environment generally refers to a computing cluster network, commonly used by multiple users to share computing and network resources to complete their own computing tasks or jobs. When performing bypass collection of communication traffic on this network, bypass collection of network communication traffic within the cluster is performed at the core switching node, aggregation switching node, or computing node access layer via port mirroring, optical splitters, or traffic replication devices. This bypass collection does not affect the computing node's operating environment and does not require the deployment of any probe programs within the computing nodes.

[0024] The collected network communication traffic, including packet timestamps, source addresses, destination addresses, and transport layer session identifiers, can be used to identify the request and response characteristics of round-trip communication relationships.

[0025] Step S2: Analyze the collected network communication traffic and identify the computational task based on the analyzed information.

[0026] In this step, the computational task identification is performed based on the parsed information, specifically including: Step S21, parsing the network communication traffic to obtain information such as data packet timestamps, source addresses, destination addresses, and transport layer session identifiers; Step S22, dividing the network traffic based on the parsed source addresses, destination addresses, and transport layer session identifiers, according to the communication sessions; Step S23, determining the user and computational task to which each communication session belongs based on the communication topology characteristics or scheduling mapping information of each communication session; Step S24, aggregating the communication sessions generated by the same computational task to obtain the network communication traffic corresponding to each computational task.

[0027] Step S3: Calculate the network performance metric (NPM) for the network communication traffic corresponding to each computing task, which is used to characterize the network communication status exhibited by the computing task during operation.

[0028] In this step, the network performance metric NPM can be set to... The metrics should include at least Round-Trip Time (RTT), packet loss probability, retransmission strength, retransmission rate, duplicate sequence number occurrence rate, and abnormal acknowledgment or out-of-order behavior rate. A time window should be preset before performing metric statistics. The current computing task Divided into several (e.g.) (number) time windows fragments; based on and Statistics are collected in each preset time window. The index value within. Specifically, it includes the following steps: Step S31, preset the time window and divide the communication traffic of the computing task into multiple consecutive time windows.

[0029] In this step, a fixed-length time window is preset before performing network performance metric statistics. and will compute tasks The corresponding network communication traffic is divided into multiple consecutive time windows in chronological order. All subsequent network performance metrics will be calculated using the computation task plus time window as the statistical granularity.

[0030] Step S32, statistically calculate all NPMs for the task in each time window. }

[0031] In this step, the NPM includes RTT, packet loss probability, retransmission strength, retransmission packet ratio, duplicate sequence number occurrence ratio, and abnormal acknowledgment or out-of-order behavior ratio. Among these, packet loss probability, retransmission strength, retransmission packet ratio, duplicate sequence number occurrence ratio, and abnormal acknowledgment or out-of-order behavior ratio are all network abnormal behaviors. For example, packet loss probability describes the proportion of data packets that fail to reach the other end during network transmission; retransmission strength describes the frequency of retransmission behavior triggered by the transport layer protocol to compensate for data loss or abnormality.

[0032] When calculating the Round-Trip Time (RTT) metric, for request-response communication sessions belonging to the same computational task, the RTT within the time window W is extracted, and the RTT statistic for computational task t within the time window W is recorded as follows: . The RTT can be characterized using any of the following statistical methods: the median or high quantile (e.g., p90, p99) of the RTT within the time window W, or the RTT jitter amplitude, such as the difference between the high and low quantiles. The RTT is used to reflect the network latency level and stability of the computing task within the corresponding time window.

[0033] It should be noted that while packet loss probability and retransmission intensity may be numerically correlated, they reflect different technical meanings, corresponding to network layer loss behavior and transport layer compensation behavior, respectively. Their statistical methods and manifestations may differ under different network protocols or communication scenarios.

[0034] In practice, the above-mentioned network anomaly behavior indicators can be used individually or in combination to reflect the degree of network anomalies in computing tasks.

[0035] Step S4 involves constructing a historical baseline for the network performance metrics of each computing task, calculating the deviation of the corresponding network performance metrics based on the historical baseline, and constructing a time series of the metric deviations. In this step, to eliminate the inherent differences caused by different jobs and different communication paths, a historical baseline for the corresponding network performance metrics is constructed for each computing task. For example, for network performance metrics such as Round Trip Time (RTT) and Loss (Network Abnormal Behavior), corresponding historical baselines are constructed, namely the RTT baseline. and Loss baseline When constructing historical baselines, network performance metrics for the same type of historical tasks by the same user, historical statistical results for the same communication topology and similar time periods, or network performance metrics during the initial stable operation phase of the current computing task can be used.

[0036] Based on the established historical baseline, the index deviation is calculated using the following formula: In equation (1), It is the j-th network performance metric value of task t within the time window W; It is the historical baseline of the j-th network performance metric for computation task t within the time window W; It is the deviation of the network performance metric for the calculation task t within the time window W for the j-th network performance metric.

[0037] Network performance metrics deviations across all time windows The time series of index deviations constituting each computational task { }

[0038] For example, when calculating the RTT bias, the formula for calculating the RTT bias of the computation task within the current time window W is as follows: In equation (2), It represents the RTT deviation of the computation task t within the time window W, reflecting the degree of change in job communication latency relative to its normal state.

[0039] The time series of the index deviation is used to characterize the dynamic process of network performance changes in computing tasks.

[0040] Step S5: Based on the time series of index deviations, perform synchronization analysis on the computing tasks and calculate the comprehensive correlation coefficient to determine whether contention for shared network resources or changes in the overall network environment have caused abnormal network performance of the computing tasks.

[0041] This step specifically includes the following steps: Step S51, constructing a reference object.

[0042] Centered on the computation task t to be analyzed, at least one reference object is constructed to characterize the network influencing factors outside the computation task. The reference object includes, but is not limited to: 1) Task pair object: forming a task pair between computation task t and another computation task; 2) Task set object: forming a reference set of multiple computation tasks other than computation task t; 3) Network environment object: constructing a network environment reference sequence based on the index deviation of multiple computation tasks within the same time window to characterize the overall network operating status.

[0043] By constructing a reference object, we can determine whether the network performance degradation of a certain computing task occurs synchronously with other tasks or the overall network environment in the time dimension. If we only look at the deviation of the indicators of a single task, we can only know "whether it has deteriorated", but we cannot determine "why it has deteriorated". Therefore, we analyze the synchronicity by constructing a reference object.

[0044] Among them, the time series of the index deviation of the reference object { Based on the deviations of various computational task indicators, the system is constructed as needed during the synchronization analysis process, according to the rules of the reference object.

[0045] Step S52: Calculate the correlation coefficient for each indicator based on the reference object.

[0046] In this step, for the j-th network performance index, the time series correlation coefficient of the index deviation between the computational task t to be analyzed and the reference object on that index is calculated. : In equation (3), To calculate the time series of the deviation of the j-th network performance metric for task t, Let ref be the time series of the deviation of the j-th network performance metric corresponding to the reference object ref.

[0047] Step S53: Perform directional constraint processing on the correlation coefficients to retain the positive correlations that deteriorate synchronously.

[0048] This step uses directional constraints to ensure that the synchronicity analysis only retains changes that are meaningful to network interference, automatically removing improvements, reverse changes, and random noise. To ensure that the correlation results only reflect the synchronous deterioration of network performance, directional constraints are applied to the correlation coefficients of each indicator: In equation (4), It represents the positive correlation coefficient of the j-th network performance index under the meaning of synchronous deterioration.

[0049] Step S54: Combine the correlation coefficients of all indicators for each computation task to obtain the comprehensive correlation coefficient of the computation task.

[0050] In this step, the most significant evidence of synchronous deterioration is extracted from multiple network performance metrics and fused to obtain a comprehensive positive correlation coefficient: In equation (5), The comprehensive synchronization correlation coefficient between task t and the reference object is used to characterize the degree of synchronous deterioration of their network performance deviations, and serves as the input for subsequent calculations of the interference intensity term and interference index.

[0051] Step S6: Calculate the interference intensity term for each computation task based on the index deviation time series.

[0052] In this step, to further quantify the severity of network performance degradation of the computation task within the current time period, the deviation magnitude of the network performance indicators of the computation task itself is analyzed, and an interference strength term is calculated to characterize the degree of network performance degradation corresponding to the computation task. Specifically, this includes the following steps: Step S61, applying directional constraints to the indicator deviation to obtain a positive indicator deviation.

[0053] This step addresses the deviations of the computational task across various network performance metrics. To retain only the bias indicating the direction of network performance degradation, the index bias is subjected to directional constraint processing, resulting in a positive index bias: In equation (6), This represents the positive deviation of the j-th network performance metric relative to the historical baseline within the time window W for calculation task t. When the metric deviation is less than or equal to zero, it indicates that the network performance is in a normal or improved state and is not included in the interference intensity calculation.

[0054] By using the above-mentioned directional constraint processing, the impact of normal fluctuations or improvements in network performance on interference intensity assessment can be effectively eliminated, allowing subsequent analysis to focus only on performance changes that adversely affect computing tasks.

[0055] Step S62: Normalize the magnitude of the positive index deviation.

[0056] Because different network performance metrics differ in numerical scale, units, and ranges—for example, round-trip time, packet loss probability, and retransmission intensity are difficult to directly compare or integrate—the positive deviation of each network performance metric is normalized. In equation (7), norm(·) represents the normalization function, which is used to map the positive deviation of different indicators to a unified dimensional range. The normalization process can be implemented based on historical statistical range, maximum and minimum values, quantiles or other preset methods, and the specific implementation method is not limited.

[0057] Normalization ensures that different network performance indicators are comparable in interference intensity calculations, avoiding unreasonable amplification or suppression of calculation results due to differences in the dimensions of a single indicator.

[0058] Step S63: Calculate the interference intensity term.

[0059] After normalizing the positive deviations of each network performance indicator, the normalized deviations of multiple key network performance indicators are fused to obtain the interference intensity term for the computation task within the current time window: In equation (8), This represents the overall intensity of network performance degradation for computation task t within the current time period.

[0060] By employing a product-based fusion approach, the interference intensity term only increases significantly when multiple key network performance indicators deteriorate significantly at the same time; when only a single indicator deteriorates slightly or other indicators are in a normal state, the interference intensity term remains at a low level, thereby effectively avoiding misjudgments caused by occasional or single indicator anomalies.

[0061] Step S7: Calculate the interference index for each computational task based on the comprehensive correlation coefficient and the interference intensity term.

[0062] This step integrates the synchronization analysis results with the interference intensity, quantifying the degree of network interference experienced by the computational task within the current time period to obtain the interference index. For the computational task t to be analyzed, based on the comprehensive synchronization correlation coefficient... and interference intensity term The formula for calculating the interference index is as follows: In equation (9), It represents the comprehensive correlation coefficient between the computation task t and the reference object in terms of network performance index deviation, and is used to reflect whether network performance anomalies occur synchronously in the time dimension; This represents the overall intensity of network performance degradation for computation task t within the current time window; This represents the network interference index for task t within the current time period; clip(·) represents a range constraint function used to limit the calculation results to a preset range. For example, the interference index, through the range constraint function, is limited to a value between 0 and 1.

[0063] By multiplying the synchronicity correlation coefficient by the interference intensity term, the interference index only reaches a large value when the network performance anomaly of the computation task occurs synchronously with that of the reference object in the time dimension, and the performance degradation of the computation task itself reaches a certain level. Range constraints facilitate a unified comparison and evaluation of interference levels across different computation tasks, time periods, and network environments, thus simplifying subsequent scheduling decisions, fairness analysis, and resource optimization. Network Interference Index The interference index is used to characterize the degree of network interference experienced by a computing task within the current time period. It comprehensively considers the synchronicity and severity of network performance anomalies, providing a direct basis for the identification, quantification, and decision-making regarding network interference between multiple computing tasks. The calculation results can be used for subsequent job interference identification and multi-tenant fairness assessment.

[0064] Further analysis of interference between computational tasks can be performed based on the interference index. For example, by specifying the range constraints of the interference index, the analysis of interference between computational tasks includes: when... When the value is close to 0, it indicates that the computation task has not been significantly affected by network interference in the current time period; when... A value close to 1 indicates that the computational task is experiencing a high degree of network interference during the current time period.

[0065] Based on the same approach, this invention also provides an interference analysis system for multiple computing tasks based on network communication behavior. The system includes: a data acquisition module, a task identification module, an NPM statistics module, an index deviation calculation module, a correlation coefficient calculation module, an interference intensity term calculation module, and an interference index calculation module. Specifically, the data acquisition module is used to collect network communication traffic in a multi-user computing task environment via bypass; the task identification module is used to parse the collected network communication traffic and identify computing tasks based on the parsed information; the NPM statistics module is used to calculate the network performance index (NPM) of the network communication traffic corresponding to each computing task; the index deviation calculation module is used to construct a historical baseline of the network performance index for each computing task, calculate the network performance index deviation of the corresponding computing task based on the historical baseline, and construct an index deviation time series; the correlation coefficient calculation module is used to perform synchronization analysis on the computing tasks based on the index deviation time series and calculate a comprehensive correlation coefficient; the interference intensity term calculation module is used to calculate the interference intensity term for each computing task based on the index deviation time series; and the interference index calculation module is used to calculate the interference index for each computing task based on the comprehensive correlation coefficient and the interference intensity term.

[0066] The system or device for executing the method in this embodiment of the invention can be a terminal or a server. The system includes a processor, a memory, and / or a transceiver, etc., and is connected via a communication bus. Each module can be implemented by a processor, a memory, and / or a transceiver, etc. The processor can be, but is not limited to, one or more microprocessors (MPUs), central processing units (CPUs), network processors (NPs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc., or can be configured to implement one or more integrated circuits of this invention. The processor can perform various functions by running or executing software programs in the memory and calling data in the memory. The memory includes Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Compact Disc Read-Only Memory (CD-ROM), and / or Non-Volatile Memory (NVM), etc. The transceiver is used to communicate with network devices or terminal devices, and includes a receiver and a transmitter. The memory and transceiver can be integrated with the processor or exist independently.

[0067] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0068] It should also be noted that the interference analysis system for multiple computing tasks based on network communication behavior described in this embodiment corresponds to the interference analysis method for multiple computing tasks based on network communication behavior. The description and limitations of the method also apply to the system, and will not be repeated here.

[0069] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0070] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed, and is not intended to limit the scope of the claimed invention, but merely to illustrate preferred embodiments of the invention. Those skilled in the art should understand that the scope of the invention is not limited to the specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

Claims

1. A method for interference analysis among multiple computational tasks based on network communication behavior, characterized in that, The method includes: bypassing the collection of network communication traffic in a multi-user computing task environment; parsing the collected network communication traffic and identifying computing tasks based on the parsed information; statistically analyzing the network performance index (NPM) of the network communication traffic corresponding to each computing task; constructing a historical baseline for the NPM of each computing task, calculating the deviation of the NPM of the corresponding computing task based on the historical baseline, and constructing a time series of the NPM deviation; performing synchronization analysis on the computing tasks based on the time series of the NPM deviation and calculating a comprehensive correlation coefficient; calculating the interference intensity term of each computing task based on the time series of the NPM deviation; and calculating the interference index of each computing task based on the comprehensive correlation coefficient and the interference intensity term.

2. The method according to claim 1, characterized in that, The process of identifying computational tasks based on parsed information includes: parsing network communication traffic to obtain data packet timestamps, source addresses, destination addresses, and transport layer session identifiers; segmenting network traffic based on communication sessions using the parsed source addresses, destination addresses, and transport layer session identifiers; determining the user and computational task to which each communication session belongs based on the communication topology characteristics or scheduling mapping information of each communication session; and aggregating communication sessions generated by the same computational task to obtain the network communication traffic corresponding to each computational task.

3. The method according to claim 1, characterized in that, The network performance metrics (NPM) include round-trip time (RTT), packet loss probability, retransmission strength, retransmission rate, duplicate sequence number occurrence rate, and abnormal acknowledgment or out-of-order behavior rate.

4. The method according to claim 1, characterized in that, The statistical network performance metrics (NPM) specifically include: setting a preset time window and dividing the communication traffic of computation tasks into multiple consecutive time windows W; and statistically analyzing the NPM of all computation tasks within each time window. }。 5. The method according to claim 4, characterized in that, When the network computing task has a round-trip time (RTT), the median or high quantile of the RTT within the time window W, or the RTT jitter amplitude, is used to characterize it.

6. The method according to claim 4, characterized in that, The calculation formula for the network performance index deviation of the corresponding computing task based on the historical baseline is as follows: In equation (1), It is the j-th network performance metric value of task t within the time window W; It is the historical baseline of the j-th network performance metric for computation task t within the time window W; It is the deviation of the network performance metric for task t within the time window W; the deviation of the network performance metric across all time windows W. The time series of index deviations constituting each computational task { }。 7. The method according to claim 1, characterized in that, The process of performing synchronization analysis on computational tasks and calculating the comprehensive correlation coefficient specifically includes: constructing a reference object; calculating the correlation coefficient for each indicator based on the reference object; applying directional constraints to the correlation coefficients to retain positive correlations that deteriorate synchronously; and fusing the correlation coefficients of all indicators for each computational task to obtain the comprehensive correlation coefficient of the computational task.

8. The method according to claim 1, characterized in that, The calculation of the interference intensity term for each calculation task specifically includes: performing directional constraint processing on the index deviation to obtain a positive index deviation; normalizing the magnitude of the positive index deviation; and fusing the calculation of the interference intensity term.

9. The method according to claim 1, characterized in that, The interference index for each computational task is calculated using the following formula: In equation (9), It represents the comprehensive correlation coefficient between the computation task t and the reference object in terms of network performance index deviation, and is used to reflect whether network performance anomalies occur synchronously in the time dimension; This represents the overall intensity of network performance degradation for computation task t within the current time window; This represents the network interference index of task t within the current time period; clip(·) represents the range constraint function, used to limit the calculation results to a preset range.

10. A system for analyzing interference between multiple computational tasks based on network communication behavior, characterized in that, The system includes: a data acquisition module, a task identification module, an NPM statistics module, an indicator deviation calculation module, a correlation coefficient calculation module, an interference intensity calculation module, and an interference index calculation module. Specifically, the data acquisition module is used to collect network communication traffic in a multi-user computing task environment via bypass; the task identification module is used to parse the collected network communication traffic and identify computing tasks based on the parsed information; the NPM statistics module is used to calculate the network performance index (NPM) of the network communication traffic corresponding to each computing task; the indicator deviation calculation module is used to construct a historical baseline for the network performance index of each computing task, calculate the network performance index deviation of the corresponding computing task based on the historical baseline, and construct an indicator deviation time series; the correlation coefficient calculation module is used to perform synchronization analysis on computing tasks based on the indicator deviation time series and calculate a comprehensive correlation coefficient; the interference intensity calculation module is used to calculate the interference intensity item for each computing task based on the indicator deviation time series; and the interference index calculation module is used to calculate the interference index for each computing task based on the comprehensive correlation coefficient and the interference intensity item.