A large file transmission state monitoring method based on multi-service cooperation

By deploying monitoring agents in multinational, multi-branch, heterogeneous networks, collecting transmission telemetry and business policies, and performing time synchronization and consistency of reporting, a multi-service collaborative monitoring model is constructed. This solves the compliance and priority issues in the process of large file transmission, improves the accuracy of monitoring and compliance awareness, and ensures the protection of critical business flows and the availability of upper-layer systems.

CN121530964BActive Publication Date: 2026-05-12BEIJING INTRON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INTRON INFORMATION TECH CO LTD
Filing Date
2025-11-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In multinational, multi-branch, heterogeneous networks, existing monitoring solutions cannot effectively address compliance and priority issues during large file transfers, resulting in high false alarm/missed alarm rates, failure to promptly detect cross-regional routing or non-compliant links, and difficulty in supporting resource allocation decisions.

Method used

By deploying monitoring agents on branch nodes, data transmission telemetry and business policies are collected, time synchronization and standardization are achieved, a multi-business collaborative monitoring model is built, a composite health score and compliance risk level are generated, and model parameters are optimized through online updates and federated learning to provide prioritized alarms and scheduling suggestions.

Benefits of technology

It improves the monitoring accuracy and compliance awareness of large file transfers in cross-border, multi-branch scenarios, ensures the protection of critical business flows and the availability of upper-layer systems, and reduces model drift and false alarm/missed alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121530964B_ABST
    Figure CN121530964B_ABST
Patent Text Reader

Abstract

The application discloses a kind of big file transmission state monitoring methods based on multi-service cooperation, it is related to computer network and data transmission technical field, the application is by in branch side uniform collection effective throughput, delay, packet loss and so on Key telemetry is added to service priority, timeliness and regional compliance strategy, first, the consistency of time reference and statistical caliber is completed, comparable transmission flow record is formed, the misjudgment problem caused by the different data caliber and time sequence mismatch of cross-domain monitoring is solved;On this basis, a multi-service cooperative monitoring model is used, the three types of constraints of performance, SLA and compliance are integrated into a single health metric, and the interzone risk is included in the continuous penalty term, so that the monitoring result can reflect both performance degradation and compliance risks, avoiding distortion by only looking at throughput while ignoring compliance paths.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer network and data transmission, and particularly relates to a large file transmission state monitoring method based on multi-service cooperation. BACKGROUND

[0002] In the business of financial consolidation, supply chain reconciliation and compliance audit, multinational enterprises need to synchronize large files among multiple branches, such as objects that need to be fragmented or transmitted in segments. The network conditions in different branches are significantly different, and there is high latency, jitter and bandwidth fluctuation in some areas. At the same time, it is necessary to comply with the data compliance requirements of the local place. For example, GDPR sets compliance conditions for cross-border transmission to third countries, such as adequacy decisions or standard contract terms, and there are data localization requirements in some countries or industries.

[0003] Existing monitoring solutions mostly collect end-to-end telemetry (throughput, latency, error code, etc.) continuously and combine anomaly detection / prediction models to evaluate the health degree. Some systems use federated learning to reduce the risk of exporting raw data, or combine dynamic thresholds to adapt to network fluctuations.

[0004] In the scenario where multiple branches, heterogeneous networks and compliance constraints coexist, in the existing monitoring, the model drifts with the rapid changes of network and business strategy, and the false positive / false negative rate increases. The monitoring indicators do not encode the region and compliance strategy, and cannot timely find cross-zone routing or non-compliant links. The business priority and SLA are not included in the monitoring domain, and the executable priority alarm cannot be provided for key business flows. In addition, the cross-node time reference is inconsistent and the indicator caliber is not unified, which leads to the deviation of the prediction result and the actual transmission state. On the heterogeneous link, only the throughput threshold is used for judgment, and the quality indicators such as effective throughput and retransmission rate are ignored, which is difficult to support resource allocation decisions. Therefore, there is an urgent need for a large file transmission state monitoring method based on multi-service cooperation to solve such problems. SUMMARY

[0005] In view of the above existing problems, the present application is proposed.

[0006] The present application provides a large file transmission state monitoring method based on multi-service cooperation to solve the problem that existing monitoring is difficult to support under the constraint of compliance and priority in the synchronization of large files in multinational multi-branch heterogeneous networks.

[0007] To solve the above technical problems, the present application provides the following technical solutions:

[0008] The embodiment of the present application provides a large file transmission state monitoring method based on multi-service cooperation, which comprises:

[0009] Step S1: Deploy monitoring agents on multiple branch nodes to collect local transmission telemetry and business policies. Transmission telemetry includes effective throughput, round-trip time, packet loss / retransmission rate, error codes, queue length and bandwidth utilization. Business policies include business priority, data timeliness requirements and regional compliance policies.

[0010] Step S2: Synchronize the collected data in time and ensure consistency in terminology to generate a transmission stream record with regional and compliance labels;

[0011] Step S3: Construct a multi-service collaborative monitoring model based on the transmission stream records, and output a composite health score and compliance risk level for each service stream;

[0012] Step S4: Generate monitoring results based on the scores and risk levels, including: prioritized alarms, root cause indications and scheduling suggestions, and provide them to the upper-layer system through the northbound interface;

[0013] Step S5: Update the model parameters and thresholds online based on the historical false alarm / missed alarm statistics;

[0014] The scheduling recommendations include, but are not limited to, speed limits, path isolation, or preferred paths within a region.

[0015] As a preferred embodiment of the large file transmission status monitoring method based on multi-service collaboration described in this invention, the regional compliance policy includes restrictions on the regions traversed by the transmission path, restrictions on data storage / processing locations, and cross-border transmission conditions. The monitoring agent adds a regional identifier and a processing location identifier to each transmission segment.

[0016] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the multi-service collaborative monitoring model performs weighted fusion of performance indicators, SLA constraints and compliance constraints to obtain the composite health score. The composite health score is used to sort the service flow and compare it with the alarm threshold to trigger alarms and suggestions.

[0017] The weighted fusion step of the composite health score includes:

[0018] Step S31: Unify the directional attributes and statistical caliber of effective throughput, round-trip time (RTT), packet loss / retransmission rate, bandwidth utilization, and completion time quantile values; determine the business target value and warning limit for each indicator; and ensure that the sampling window and quantile statistics are consistent with the SLA definition, serving as the baseline for normalization and penalties.

[0019] Step S32: In the scoring phase, perform interval linear mapping on each original indicator and truncate both ends to obtain utility scores.

[0020] ,

[0021] in, Indicates business flow index In the index of indicators Normalization effect on Indicates business flow In terms of indicators The original observations As an indicator The warning boundary, As an indicator The business target value, As a direction factor, It means the bigger the better. It means the smaller the better. Indicates the real number Clip to range ;

[0022] Step S33: Perform a convex combination of the static weights and the learned weights to obtain the final weights:

[0023] ,

[0024] in, Index of indicators The final weight, The fusion coefficient is... To configure weights statically, The weights obtained through learning, and Source identification badge;

[0025] Step S34: Perform dimensionless aggregation on negative deviations exceeding the target.

[0026] ,

[0027] in, Indexing for business flows SLA default aggregate amount To include the set of indicators in the scoring, For absolute values, Subscript of positive part operator ;

[0028] Step S35, map discrete cross-regional risk to Consecutive penalties:

[0029] ,

[0030] in, Indexing for business flows Compliance penalties To comply with the risk level of crossing the boundary, The upper bound of the risk level is a constant subscript. ;

[0031] Step S36: Perform a weighted summation of the utilities, deduct the two types of penalties, and then trim to the nearest whole number. Get the final health rating:

[0032] ,

[0033] in, Indexing for business flows The final composite health score, superscript The status indicator after deduction and cropping. , For indicator weights and normalized utility, This is the SLA penalty coefficient. This represents the aggregate amount of SLA defaults. For compliance penalty coefficients, For compliance penalties;

[0034] Step S37: Stratify by business priority category and construct multi-level alarm thresholds based on the quantiles of historical healthy samples:

[0035] ,

[0036] in, Priority Category Index Alarm Level Index The health threshold at that location Here is the quantile function, and the quantile parameters are: , For composite health, the badge Indicates category hierarchy, superscript Indicates the alarm level.

[0037] As a preferred embodiment of the large file transmission status monitoring method based on multi-service collaboration described in this invention, the monitoring model includes a graph neural network, which models branch nodes as graph nodes and transmission links as graph edges.

[0038] Node features encode business strategy vectors, and edge features encode link states, which are used to estimate path-level degradation probability and impact range.

[0039] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the method includes a compliance path verification step:

[0040] The autonomous region, region, and transit point information of the real-time path are compared with the regional compliance policy, and the risk of crossing the region is marked and incorporated into the compliance risk level.

[0041] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the monitoring results include a priority alarm queue, and each alarm is given an executable handling suggestion type and the scope of the suggestion's effectiveness, suggesting that the network configuration should not be directly modified.

[0042] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the online update includes a threshold adaptive module based on reinforcement learning, which adjusts the alarm threshold and the weight of each indicator according to the false alarm / missed alarm statistics of historical alarms and the SLA satisfaction rate of key services, in order to reduce the overall cost and improve the SLA satisfaction rate.

[0043] The reward function for reinforcement learning is constructed as follows:

[0044] Step S51: Set the offline playback or weak supervision annotation criteria, according to time steps. Within a sliding window, count false alarms, missed alarms, and the achievement of critical business SLAs, and record the number of alarms and the change in adjacent steps as inputs for various rewards;

[0045] Step S52, at each time step, calculate the weighted cost of false positives and false negatives:

[0046] ,

[0047] in, Indicates time step The cost of mistakes, For false alarm weighting coefficients, For underreported weighting coefficients, For false alarm rate, The false negative rate, The number of false alarms For the number of underreported entries, The total number of samples already evaluated is indicated by the superscript. , , For item category subscripts;

[0048] Step S53: Provide positive rewards for the achievement rate of key business tasks:

[0049] , ,

[0050] in, For time step SLA rewards, This is the SLA reward coefficient. To ensure SLA satisfaction rate for key business operations, To meet the critical business sample size requirements of the SLA, The total number of key business samples, superscript , For item category subscripts;

[0051] Step S54, measure scale and jitter using the number of alarms and the difference between adjacent steps:

[0052] ,

[0053] in, As a regularization cost, The coefficient is the quantity regularization factor. This is the jitter regularization coefficient. For time step The number of alarms, This represents the number of alerts from the previous step. For scale normalization constants, superscript , , , For item category subscripts;

[0054] In step S55, at the same time step, the reward and cost are combined into a scalar return:

[0055] ,

[0056] in, For time step Instant rewards For SLA rewards, For the cost of mistakes, For regularization costs;

[0057] Step S56: Encode the recent health distribution, branch load, and time period label into a status:

[0058] ,

[0059] in, For time step The state vector, Composite health score within the window The mean, For the same window standard deviation, For branch load index, This is a one-hot vector for the time period label. For time period index, The histogram ratio of the third interval.

[0060] Step S57, the action consists of a threshold and a fine-tuning step size of the weights, and the constraints are satisfied by projection:

[0061] , ,

[0062] ,

[0063] in, Weight of indicators Step size, Threshold Step size, , Pre-projection measurement For projection onto the probabilistic simplex The operator, For projection onto the set of monotonic thresholds The operator, The upper bound of the step size. This serves as the subscript for the next time step.

[0064] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the monitoring model is updated using a federated learning framework.

[0065] Each branch node trains model parameters locally and reports encrypted gradients to the aggregation end. After the aggregation end completes model aggregation and consistency verification, it issues updates.

[0066] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, wherein: the federated learning introduces a differential privacy mechanism in the aggregation phase to add noise to the aggregation amount.

[0067] As a preferred embodiment of the large file transfer status monitoring method based on multi-service collaboration described in this invention, the time synchronization and consistency of standards include:

[0068] The branch nodes are clock-synchronized, the sampling period and indicator definitions are unified, and the effective throughput and retry amplification factor are calculated on the monitoring agent side.

[0069] The beneficial effects of this invention are as follows: By uniformly collecting key telemetry data such as effective throughput, latency, and packet loss on the branch side and overlaying them with business priority, timeliness, and regional compliance policies, this invention first achieves consistency in time benchmarks and statistical standards, forming comparable transmission flow records and solving the problem of misjudgment caused by inconsistent cross-domain monitoring data standards and timing mismatches. Based on this, a multi-service collaborative monitoring model is adopted, integrating performance, SLA, and compliance constraints into a single health metric, and incorporating cross-regional risk into continuous penalties. This allows monitoring results to simultaneously reflect performance degradation and compliance risks, avoiding the neglect of compliance issues when only considering throughput. The system addresses path distortion; by prioritizing alarms and providing actionable suggestions with execution scope, monitoring signals are directly transformed into actionable rate limiting, path isolation, or regional optimal path suggestions, alleviating resource competition between critical and batch operations and improving the certainty of SLA achievement; the online adaptive module uses false positives / false negatives, SLA fulfillment rate, and alarm scale fluctuations as rewards and costs to automatically adjust thresholds and weights, reducing model drift caused by environmental non-stationarity; federated learning combined with differential privacy enables cross-branch model collaborative updates without leaking original data, balancing privacy and generalization, and adapting to regional differences;

[0070] In summary, this invention improves the accuracy, interpretability, and compliance awareness of large file transfer monitoring in multinational, multi-branch scenarios, and strengthens the protection of critical business flows and the availability support for upper-layer systems. Attached Figure Description

[0071] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.

[0072] Figure 1 This is a flowchart illustrating the large file transfer status monitoring method based on multi-service collaboration in this embodiment. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0074] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0075] For example, the terms “first” and “second” used in this application are only used to distinguish and describe similar objects, to differentiate the first object from another object, and are not used to describe a specific order or sequence, nor should they be interpreted as indicating or implying relative importance.

[0076] This application proposes a method for monitoring the status of large file transfers based on multi-service collaboration, combining... Figure 1 As shown, the method includes:

[0077] Step S1: Deploy monitoring agents on multiple branch nodes to collect local transmission telemetry and business policies. Transmission telemetry includes effective throughput, round-trip time, packet loss / retransmission rate, error codes, queue length and bandwidth utilization. Business policies include business priority, data timeliness requirements and regional compliance policies.

[0078] Step S2: Synchronize the collected data in time and ensure consistency in terminology to generate a transmission stream record with regional and compliance labels;

[0079] Step S3: Construct a multi-service collaborative monitoring model based on transmission flow records, and output a composite health score and compliance risk level for each service flow;

[0080] Step S4: Generate monitoring results based on the score and risk level, including: prioritized alarms, root cause indications and scheduling suggestions, and provide them to the upper layer system through the northbound interface;

[0081] Step S5: Update the model parameters and thresholds online based on the historical false alarm / missed alarm statistics;

[0082] The scheduling recommendations include, but are not limited to, speed limits, path isolation, or preferred paths within a region;

[0083] In this embodiment, the monitoring agent is deployed at the transmission endpoint or edge gateway of the branch node. The northbound interface refers to the service interface that outputs monitoring results and scheduling suggestions to the upper-layer system. The transmission flow record is an observation item at the business flow granularity after time alignment and standardization. The effective throughput is the amount of data successfully delivered per unit time without retransmissions. The queue length is the queuing depth on the sending or receiving side. The default sampling period is 1 second, which can be configured within 200 milliseconds to 5 seconds. The default rolling window for scoring and alarms is five minutes, which can be adjusted between 1 and 15 minutes. Online updates are triggered once every 15 minutes by default, which can be configured between 5 and 30 minutes, depending on the operation and maintenance strategy and the consumption capacity of the upper-layer system. The business target value and warning threshold are derived from the service level agreement and historical stable period statistics. The percentile statistics default to 95%. The priority category is derived from the business policy configuration and is dynamically loaded as the policy changes. Optionally, the queue length is preferentially taken from the sending side. If the receiving side is more representative of the bottleneck, the receiving side is taken as the standard and the source is marked in the record. The northbound interface can reuse the push channel of the existing enterprise service bus or configuration center without changing its permission model. If necessary, if sampling is missing or a short-term failure results in insufficient data within the window, the average of the most recent stable window will be used to backfill the data and mark it as the inferred value. If the business strategy is temporarily unavailable, the monitoring results will be generated with the default priority and the ordinary compliance strategy, and will be updated to compensate when the strategy is restored.

[0084] In one embodiment, the regional compliance policy includes restrictions on the regions traversed by the transmission path, restrictions on the location of data storage / processing, and cross-border transmission conditions. The monitoring agent adds a regional identifier and a processing location identifier to each transmission segment.

[0085] Specifically, the geographic identifier maps the network traversed by the transmission segment to the landing location to a predetermined geographic encoding set, while the processing location identifier is the physical or logical geographic code where the data is stored or processed; both are derived from the merging of endpoint registration information and path detection results. The geographic and processing location mapping is calibrated daily by default and can be adjusted within 6 to 48 hours; the policy rule base is synchronized every 24 hours by default and can be adjusted within 1 to 72 hours, with calibration based on the geographic list and cross-border conditions maintained by the compliance team. If the endpoint registration information and path detection are inconsistent, a more conservative geographic set is used for accounting, triggering a low-level alarm to prompt manual review; if a geographic location cannot be resolved, it is treated as a high-risk situation and automatically reviewed in subsequent calibration cycles. Optionally, when path changes frequently, only the geographic identifier for the changed segments is recalculated to reduce overhead while maintaining the continuity of the transmission stream record.

[0086] In one embodiment, the multi-service collaborative monitoring model weights and fuses performance metrics, SLA constraints, and compliance constraints to obtain a composite health score. The composite health score is used to sort business flows and compare them with alarm thresholds to trigger alarms and recommendations.

[0087] The weighted fusion steps of the composite health score include:

[0088] Step S31: Unify the directional attributes and statistical calibers of effective throughput, round-trip time (RTT), packet loss / retransmission rate, bandwidth utilization, and completion time quantiles, determine the business target value and warning limit for each indicator, and ensure that the sampling window and quantile statistics are consistent with the SLA definition, which are used as the baseline for normalization and penalties.

[0089] Step S32: In the scoring phase, perform interval linear mapping on each original indicator and truncate both ends to obtain utility scores.

[0090] ,

[0091] in, Indicates business flow index In the index of indicators Normalization effect on Indicates business flow In terms of indicators The original observations As an indicator The warning boundary, As an indicator The business target value, As a direction factor, It means the bigger the better. It means the smaller the better. Indicates the real number Clip to range ;

[0092] Step S33: Perform a convex combination of the static weights and the learned weights to obtain the final weights (then perform normalization within the dimension to sum to 1):

[0093] ,

[0094] in, Index of indicators The final weight, The fusion coefficient is... The weights are statically configured (given by business priority and SLA importance). The weights are learned (trained from historical samples or partial order pairs). and Source identification badge;

[0095] Step S34: Perform dimensionless aggregation on negative deviations exceeding the target.

[0096] ,

[0097] in, Indexing for business flows SLA default aggregate amount To include the set of indicators in the scoring, For absolute values, Subscript of positive part operator ;

[0098] Step S35, map discrete cross-regional risk to Consecutive penalties:

[0099] ,

[0100] in, Indexing for business flows Compliance penalties The compliance risk level for crossing regional boundaries is determined by comparing path-regional strategies. The upper bound of the risk level is a constant subscript. ;

[0101] Step S36: Perform a weighted summation of the utilities, deduct the two types of penalties, and then trim to the nearest whole number. Get the final health rating:

[0102] ,

[0103] in, Indexing for business flows The final composite health score, superscript The status indicator after deduction and cropping. , For indicator weights and normalized utility, This is the SLA penalty coefficient. This represents the aggregate amount of SLA defaults. For compliance penalty coefficients, For compliance penalties;

[0104] Step S37: Stratify by business priority category and construct multi-level alarm thresholds based on the quantiles of historical healthy samples:

[0105] ,

[0106] in, Priority Category Index Alarm Level Index The health threshold at that location Here is the quantile function, and the quantile parameters are: , For composite health, the badge Indicates category hierarchy, superscript Indicates the alarm level;

[0107] The trigger rule can be set as: when Trigger level three, Trigger Level 2, Trigger Level 1, higher than This is normal;

[0108] Specifically, the system first defines the direction and statistical scope of indicators to ensure comparability and segmentation of various metrics; normalization uses linear mapping combined with truncation to transform observations of different dimensions into a unified utility scale, facilitating subsequent weight aggregation; the weighting component adopts a fusion of static and learned sources, preserving management preferences from the policy side while introducing adaptive capabilities updated with historical performance, suitable for multi-business parallel scenarios; two types of penalties are connected to SLA deviation and regional compliance violations respectively, continuously incorporating the discrete conclusions of compliance verification into the scoring, enabling health to reflect both performance and compliance dimensions simultaneously; the final score is weighted and penalized with deductions and interval pruning to obtain standardized results that are easy to compare horizontally and display hierarchically; thresholds and levels are constructed using quantiles and hierarchically categorized by business type, which can form differentiated trigger sensitivities under different priorities, and can be fine-tuned online by combining historical alarm effects and key business satisfaction rates, gradually reducing the handling costs caused by false alarms and missed alarms;

[0109] For example, the implementation of the composite health score follows these engineering guidelines: Directional attributes are uniformly categorized as either "the larger the better" or "the smaller the better" before mapping; business target values ​​are derived from service level agreements or the average of historical stable periods plus a safety margin; warning limits are the lower or upper limits of the operational range set around the target; the mapping interval is linearized between the target and the warning limit and truncated at both ends to ensure score stability. The fusion coefficient in weight fusion defaults to 0.5 and can be adjusted between 0.3 and 0.7; static weights are given by business importance, while learning weights are obtained by fitting historical results and verified offline before deployment. Service level penalty coefficients and compliance penalty coefficients default to 1 and 1.5 respectively, and can be configured between 0 and 3 to reflect higher penalties for compliance; compliance risk levels are mapped to discrete levels from zero to three, directly given by path and region comparison. The quantiles of the grading thresholds default to 0.9, 0.8, and 0.65 respectively, calculated independently by priority category; fine-tuning is allowed within the range of 0.5 to 0.95 when the distribution is significantly skewed. Optionally, during the cold start phase when the business target value has not yet converged, only static weights and fixed thresholds are enabled, while learning weights and quantile adaptation are disabled. If the completion time quantile value is unavailable, a combination of effective throughput and latency metrics is used instead, and the original metrics are automatically switched back after subsequent recovery. If necessary, if the number of observed samples within a window is less than the set lower limit, an alarm is triggered with a delay and evaluated again after accumulation in the next window to avoid false alarms caused by occasional jitter.

[0110] In one embodiment, the monitoring model includes a graph neural network, which models branch nodes as graph nodes and transmission links as graph edges;

[0111] Node features encode business strategy vectors, and edge features encode link states, used to estimate path-level degradation probability and impact range;

[0112] Similarly, the node features of the graph neural network include sparse or dense representations of business priority, data timeliness, and regional compliance labels, while the edge features include statistics such as near real-time effective throughput, round-trip latency, packet loss, and bandwidth utilization, along with their rates of change. The output is a path-level degradation probability and an impact range score. Model training uses historical degradation events and manual review results as weak supervision labels, maintaining consistency between training and inference. Parameters are refreshed daily by default and can be adjusted between 6 hours and 7 days, with each training iteration using a data window from the most recent 1-4 weeks. Optionally, in stages with fewer branches or stable links, degradation is implemented using a lightweight configuration based on average neighborhood aggregation. When temporary breaks occur in the graph, inference is performed separately for each connected component, and the maximum value is taken based on the impact range during the merging phase. If necessary, if some edge features are missing, the median of the most recent stable window is used instead, and the confidence level in the output is reduced.

[0113] In one embodiment, the method includes a compliance path verification step:

[0114] The autonomous region, region, and transit point information of the real-time path are compared with the regional compliance policy, and the risk of crossing the region is marked and incorporated into the compliance risk level.

[0115] Optionally, path information is derived from the summary of the nearest reachable paths and transit points reported by endpoints and monitoring agents. Autonomous systems and regions are maintained by a pre-defined mapping table. During comparison, it is first verified whether the path passes through a restricted region, then whether the processing location meets the regional residency requirements, and then it is determined whether additional constraints are needed based on cross-border conditions. The priority for determining the risk level of cross-regional violations is that the processing location violation is higher than the path violation; if both exist simultaneously, the higher level is used. Risk labeling is executed synchronously with each path change by default. If the path fluctuates frequently in a short period of time, a suppression window is used to merge similar items to reduce duplicate alarms. If necessary, when the path or processing location cannot be determined, it is processed at a higher level and recorded as an item pending review. After the policy library is updated, it is automatically reviewed and backfilled.

[0116] In one embodiment, the monitoring results include a prioritized alarm queue, providing each alarm with an actionable handling suggestion type and its effective scope, without directly modifying network configuration. Further, the prioritized alarm queue is sorted by composite health and compliance risk, and within the same priority level, it is stably sorted by business importance and most recent trigger time. Handling suggestion types include rate limiting, path isolation, and regional path optimization, with effective scopes at the connection, endpoint, or branch levels, defaulting to prioritizing smaller effective scopes to reduce impact on other services. The default suggestion validity period is 10-30 minutes, automatically expiring and being regenerated in the next evaluation round; if a higher-level suggestion is received within the validity period for the same service, the old suggestion is overwritten. Optionally, if the upper-layer system reports that a suggestion has been executed or rejected, the monitoring system records the execution result for online updates, but does not directly modify any network configuration. If necessary, if the queue length exceeds the upper-layer system's processing capacity, the queue is trimmed according to business priority and risk level, retaining key entries.

[0117] In one embodiment, the online update includes a threshold adaptation module based on reinforcement learning, which adjusts the alarm threshold and the weight of each indicator according to the false alarm / missed alarm statistics of historical alarms and the SLA satisfaction rate of key businesses, in order to reduce the overall cost and improve the SLA satisfaction rate.

[0118] The reward function for reinforcement learning is constructed as follows:

[0119] Step S51: Set the offline playback or weak supervision annotation criteria, according to time steps. Within a sliding window, count false alarms, missed alarms, and the achievement of critical business SLAs, and record the number of alarms and the change in adjacent steps as inputs for various rewards;

[0120] Step S52, at each time step, calculate the weighted cost of false positives and false negatives:

[0121] ,

[0122] in, Indicates time step The cost of mistakes, For false alarm weighting coefficients, For underreported weighting coefficients, For false alarm rate, The false negative rate, The number of false alarms For the number of underreported entries, The total number of samples already evaluated is indicated by the superscript. , , For item category subscripts;

[0123] Step S53: Provide positive rewards for the achievement rate of key business tasks:

[0124] , ,

[0125] in, For time step SLA rewards, This is the SLA reward coefficient. To ensure SLA satisfaction rate for key business operations, To meet the critical business sample size requirements of the SLA, The total number of key business samples, superscript , For item category subscripts;

[0126] Step S54, measure scale and jitter using the number of alarms and the difference between adjacent steps:

[0127] ,

[0128] in, As a regularization cost, The coefficient is the quantity regularization factor. This is the jitter regularization coefficient. For time step The number of alarms, This represents the number of alerts from the previous step. For scale normalization constants, superscript , , , For item category subscripts;

[0129] In step S55, at the same time step, the reward and cost are combined into a scalar return:

[0130] ,

[0131] in, For time step Instant rewards For SLA rewards, For the cost of mistakes, For regularization costs;

[0132] Step S56: Encode the recent health distribution, branch load, and time period label into a status:

[0133] ,

[0134] in, For time step The state vector, Composite health score within the window The mean, For the same window standard deviation, The branch load index is a weighted average of the utilization rates of each branch link, then normalized. This is a one-hot vector for the time period label. For time period index, The histogram ratio of the third interval.

[0135] Corresponding intervals , , ;

[0136] Step S57, the action consists of a threshold and a fine-tuning step size of the weights, and the constraints are satisfied by projection:

[0137] , ,

[0138] ,

[0139] in, Weight of indicators Step size, Threshold Step size, , Pre-projection measurement For projection onto the probabilistic simplex The operator, For projection onto the set of monotonic thresholds The operator, The upper bound of the step size. This serves as the index for the next time step.

[0140] Specifically, false alarms and false negatives are factored into the cost in a weighted linear form, and the weight settings can reflect the differences in tolerance for different error types. The SLA achievement rate of key businesses generates a separate positive reward, allowing the strategy to maintain a preference for core businesses while reducing error costs. The number of alarms and the difference between adjacent steps together form a regularization term; the former limits the spillover of scale, while the latter suppresses jitter in the time series, which is conducive to the stable consumption of alarm streams by upper-level systems. The state design uses the mean and variance of the recent health distribution to provide central tendency and dispersion, combined with branch load and time period labels to characterize the predictable periodicity and capacity pressure, and further supplemented by a three-segment histogram ratio to improve the sensitivity to the distribution pattern. The action space is composed of the fine-tuning step size of thresholds and weights. By projecting the probabilistic simplex and the set of monotonic thresholds, it satisfies engineering constraints such as summation, non-negativity, and hierarchical monotonicity, thereby avoiding unexecutable solutions. The overall structure is integrated with the existing scoring system and is suitable for direct invocation in the online update module.

[0141] In this embodiment, the online update training cycle is 15 minutes by default and can be configured between 5 and 30 minutes; offline playback is used for periodic retraining and is executed once a day by default. The criteria for determining false positives and false negatives are based on manual review or the results of upper-level system handling; the service achievement statistics of key businesses are based on business-side reports and aggregated according to a unified time window. The weights of rewards and costs are set by default as follows: the weight of false positives is slightly lower than that of false negatives, the service achievement reward is not less than half of the sum of the two, and the quantity regularization and jitter regularization account for a small proportion without affecting key businesses; the alarm scale normalization constant is set according to the throughput capacity of the upper-level system and is calibrated by stress testing before going live. The constraints after action projection include non-negative weights that sum to one, and the thresholds within the same priority category maintaining a monotonic relationship from strict to lenient; the upper bound of the step size is small by default to avoid oscillation. Optionally, if the recent distribution is stable, the training cycle can be extended and the step size reduced; if the monitoring environment changes suddenly, the training cycle is shortened and the upper limit of the step size is increased to accelerate convergence. If necessary, if any component of the state vector is missing, the most recent stable window statistics will be used instead, and the update magnitude of this round will be reduced.

[0142] In one embodiment, the monitoring model is updated using a federated learning framework:

[0143] Each branch node trains model parameters locally and reports encrypted gradients to the aggregation end. The aggregation end then sends out updates after completing model aggregation and consistency verification.

[0144] Specifically, the federated learning collaborative process adopts a combination of fixed-round and trigger-based methods: fixed-round rounds are triggered once every 12-24 hours by default, while trigger-based rounds are additionally triggered when there is a significant drift in the distribution of monitoring metrics or major changes in the policy library; the number of rounds and batch size for local training are adaptively determined based on the node's computing power and data volume and reported in the metadata. Consistency verification includes checking the model version, loss reduction, and non-degradation of core metrics; only after verification can the model be distributed; if some nodes are offline or refuse to participate, the current round of aggregation continues and a list of non-participating nodes is recorded. Optionally, when there are too few participating nodes, distribution can be paused, and only the local model is retained until the next round of aggregation is completed. If necessary, if the encryption gradient decryption fails or the integrity verification fails, the contribution of that node is discarded and an alarm is issued, without affecting the aggregation of other nodes.

[0145] In one embodiment, federated learning introduces a differential privacy mechanism during the aggregation phase, adding noise to the aggregation amount to suppress back-inference of single-branch business strategies.

[0146] For example, differential privacy noise is generated independently at the aggregation end according to the aggregation quantity dimension and added to the aggregation quantity in the same shape. The noise intensity is preset according to the organization's privacy strategy, typically ranging from 1% to 5% of the average amplitude of the aggregation quantity. To reduce accuracy loss, lower noise can be used for critical layers or critical channels, while higher noise can be used for non-critical channels. Optionally, the noise intensity can be reduced to maintain model convergence when the number of nodes is large, and appropriately increased to enhance privacy when the number of nodes is small. When the number of participating nodes in the current round is significantly reduced, the noise intensity can be temporarily increased and restored in the next round. If necessary, if the consistency verification fails after noise injection, the model can be rolled back to the previous version and the noise intensity reduced before re-aggregating.

[0147] In one embodiment, time synchronization and caliber consistency include:

[0148] Clock synchronization is performed on branch nodes, sampling period and indicator definition are unified, and effective throughput and retry amplification factor are calculated on the monitoring agent side to eliminate the artificially high impact of retransmission on throughput.

[0149] Similarly, clock synchronization is aligned across all branches using a unified time source, with a default maximum allowable deviation of 50 milliseconds. The sampling period defaults to 1 second and is configured between 200 milliseconds and 5 seconds. Metric definitions are uniformly distributed and version numbers are recorded in the configuration center. Effective throughput is calculated using the unique valid byte of the service load as the numerator and the time window length as the denominator, excluding retransmissions and duplicate segments. The retry amplification factor is the ratio of retransmitted bytes to the unique valid byte, used to identify artificially high throughput caused by retries. Optionally, when the endpoint lacks fragmentation and deduplication capabilities, the number of bytes sent initially recorded by the sender is approximated to the unique valid byte, and the confidence level in the result is reduced. When short-term jitter in the link is significant, the time window can be shortened from 5 minutes to 1 minute to improve sensitivity. If necessary, if the clock deviation exceeds the allowable range, it is marked as low confidence in the corresponding window's scoring and alarms, triggering an operational calibration.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0151] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

Claims

1. A method for monitoring the status of large file transfers based on multi-service collaboration, characterized in that, include: Step S1: Deploy monitoring agents on multiple branch nodes to collect local transmission telemetry and business policies. Transmission telemetry includes effective throughput, round-trip time, packet loss / retransmission rate, error codes, queue length and bandwidth utilization. Business policies include business priority, data timeliness requirements and regional compliance policies. Step S2: Synchronize the collected data in time and ensure consistency in terminology to generate a transmission stream record with regional and compliance labels; Step S3: Construct a multi-service collaborative monitoring model based on the transmission stream records, and output a composite health score and compliance risk level for each service stream; Step S4: Generate monitoring results based on the scores and risk levels, including: prioritized alarms, root cause indications and scheduling suggestions, and provide them to the upper-layer system through the northbound interface; Step S5: Update the model parameters and thresholds online based on the historical false alarm / missed alarm statistics; The scheduling recommendations include, but are not limited to, speed limits, path isolation, or preferred paths within a region; The multi-service collaborative monitoring model weights and fuses performance indicators, SLA constraints, and compliance constraints to obtain the composite health score. The composite health score is used to sort the service flow and compare it with the alarm threshold to trigger alarms and suggestions. The weighted fusion step of the composite health score includes: Step S31: Unify the directional attributes and statistical caliber of effective throughput, round-trip time (RTT), packet loss / retransmission rate, bandwidth utilization, and completion time quantile values; determine the business target value and warning limit for each indicator; and ensure that the sampling window and quantile statistics are consistent with the SLA definition, serving as the baseline for normalization and penalties. Step S32: In the scoring phase, perform interval linear mapping on each original indicator and truncate both ends to obtain utility scores. , in, Indicates business flow index In the index of indicators Normalization effect on Represents business flow In indicators The original observations As an indicator The warning boundary, As an indicator The business target value, As a direction factor, It means the bigger the better. It means the smaller the better. Indicates the real number Clip to range ; Step S33: Perform a convex combination of the static weights and the learned weights to obtain the final weights: , in, Index of indicators The final weight, The fusion coefficient is... To configure weights statically, The weights obtained through learning, and Source identification badge; Step S34: Perform dimensionless aggregation on negative deviations exceeding the target. , in, Indexing for business flows SLA default aggregate amount To include the set of indicators in the scoring, For absolute values, Subscript of positive part operator ; Step S35, map discrete cross-regional risk to Consecutive penalties: , in, Indexing for business flows Compliance penalties To comply with the risk level of crossing the boundary, The upper bound of the risk level is a constant subscript. ; Step S36: Perform a weighted summation of the utilities, deduct the two types of penalties, and then trim to the nearest whole number. Get the final health rating: , in, Indexing for business flows The final composite health score, superscript The status indicator after deduction and cropping. , For indicator weights and normalized utility, This is the SLA penalty coefficient. This represents the aggregate amount of SLA defaults. For compliance penalty coefficients, For compliance penalties; Step S37: Stratify by business priority category and construct multi-level alarm thresholds based on the quantiles of historical healthy samples: , in, Priority Category Index Alarm Level Index The health threshold at that location Here is the quantile function, and the quantile parameter is... , For composite health, the badge Indicates category hierarchy, superscript Indicates the alarm level.

2. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, The regional compliance policy includes restrictions on the regions traversed by the transmission path, restrictions on the location of data storage / processing, and conditions for cross-border transmission. The monitoring agent adds a regional identifier and a processing location identifier to each transmission segment.

3. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, The monitoring model includes a graph neural network, which models branch nodes as graph nodes and transmission links as graph edges; Node features encode business strategy vectors, and edge features encode link states, which are used to estimate path-level degradation probability and impact range.

4. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, This method includes a compliance path verification step: The autonomous region, region, and transit point information of the real-time path are compared with the regional compliance policy, and the risk of crossing the region is marked and incorporated into the compliance risk level.

5. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, The monitoring results include a priority alarm queue, and each alarm is given an actionable handling suggestion type and the scope of the suggestion, which suggests not to directly modify the network configuration.

6. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, The online update includes a threshold adaptive module based on reinforcement learning, which adjusts the alarm threshold and the weight of each indicator according to the false alarm / missed alarm statistics of historical alarms and the SLA satisfaction rate of key businesses, in order to reduce the overall cost and improve the SLA satisfaction rate. The reward function for reinforcement learning is constructed as follows: Step S51: Set the offline playback or weak supervision annotation criteria, according to time steps. Within a sliding window, count false alarms, missed alarms, and the achievement of critical business SLAs, and record the number of alarms and the change in adjacent steps as inputs for various rewards; Step S52, at each time step, calculate the weighted cost of false positives and false negatives: , in, Indicates time step The cost of mistakes, For false alarm weighting coefficients, For underreported weighting coefficients, For false alarm rate, The false negative rate, The number of false alarms For the number of underreported entries, The total number of samples already evaluated is indicated by the superscript. , , For item category subscripts; Step S53: Provide positive rewards for the achievement rate of key business tasks: , , in, For time step SLA rewards, This is the SLA reward coefficient. To ensure SLA satisfaction rate for key business operations, To meet the critical business sample size requirements of the SLA, The total number of key business samples, superscript , For item category subscripts; Step S54, measure scale and jitter using the number of alarms and the difference between adjacent steps: , in, As a regularization cost, The coefficient is the quantity regularization factor. This is the jitter regularization coefficient. For time step The number of alarms, This represents the number of alerts from the previous step. For scale normalization constants, superscript , , , For item category subscripts; In step S55, at the same time step, the reward and cost are combined into a scalar return: , in, For time step Instant rewards For SLA rewards, For the cost of mistakes, For regularization costs; Step S56: Encode the recent health distribution, branch load, and time period label into a status: , in, For time step The state vector, Composite health score within the window The mean, For the same window standard deviation, For branch load index, For time period labels, For time period index, The histogram ratio of the third interval. Step S57, the action consists of a threshold and a fine-tuning step size of the weights, and the constraints are satisfied by projection: , , , in, Weight of indicators Step size, Threshold Step size, , Pre-projection measurement For projection onto the probabilistic simplex The operator, For projection onto the monotonic threshold set The operator, The upper bound of the step size. This serves as the subscript for the next time step.

7. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 1, characterized in that, The monitoring model is updated using a federated learning framework: Each branch node trains model parameters locally and reports encrypted gradients to the aggregation end. After the aggregation end completes model aggregation and consistency verification, it issues updates.

8. The method for monitoring the status of large file transfers based on multi-service collaboration as described in claim 7, characterized in that, The federated learning introduces a differential privacy mechanism during the aggregation phase, adding noise to the aggregation amount.

9. The method for monitoring the status of large file transfer based on multi-service collaboration as described in claim 1, characterized in that, The time synchronization and caliber consistency include: The branch nodes are clock-synchronized, the sampling period and indicator definitions are unified, and the effective throughput and retry amplification factor are calculated on the monitoring agent side.