Archive data file synchronous transmission scheduling method and system

By extracting multi-dimensional features and parsing the structure of archival data files, dynamically determining transmission priorities, and constructing a differentiated transmission channel cluster, the problems of latency and low resource utilization in the synchronous transmission of archival data are solved, thereby improving transmission efficiency and throughput performance.

CN121644556APending Publication Date: 2026-03-10GUANGZHOU XIEZHENG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511976652.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to adapt to multi-dimensional differences in the synchronous transmission of archival data files, resulting in delays in the transmission of high-value files, inefficient bandwidth usage, and network congestion. They also lack the ability for refined hierarchical scheduling and global collaboration.

Method used

By extracting multi-dimensional features and parsing the structure of archival data files, a structured feature description is generated. A pre-trained classification model and a business rule engine are used to dynamically determine the transmission priority and build a virtual transmission channel cluster with differentiated service levels. The traffic load distribution strategy is adjusted in real time to achieve synchronous scheduling of archival files.

Benefits of technology

It has achieved refined hierarchical management and global resource optimization of archival data transmission, improving the synchronization efficiency of high-value archives and the overall system throughput performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644556A_ABST
    Figure CN121644556A_ABST
Patent Text Reader

Abstract

The invention discloses an archive data file synchronous transmission scheduling method and system, and the method comprises the steps: carrying out the multi-dimensional feature extraction and structure analysis of an archive data file to be synchronized, and generating a structured feature description; dynamically judging the transmission priority category of each file and estimating the network transmission overhead of the file based on the structural feature description; pre-allocating an initial bandwidth weight and a computing resource quota for each channel according to the transmission priority category and the network transmission overhead; in the data transmission process, the throughput performance and the node load of each virtual transmission channel are monitored in real time, and a cross-channel traffic load distribution strategy is dynamically adjusted according to the channel health degree; and based on the dynamically adjusted flow load distribution strategy, coordinating and executing multi-path parallel and sequential synchronous transmission of the archive data in the virtual transmission channel cluster. By utilizing the embodiment of the invention, refined hierarchical management and global resource optimization of archive data transmission can be realized, and the synchronization efficiency and the overall throughput performance of archives are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data transmission technology, and in particular to a method and system for synchronous transmission and scheduling of archival data files. Background Technology

[0002] In applications such as archival management and cross-domain data sharing, massive and heterogeneous archival data files require efficient and reliable synchronous transmission between distributed nodes. Traditional file synchronization methods often employ first-come, first-served or fixed-priority scheduling strategies, which are ill-suited to the multi-dimensional differences in archival data regarding business attributes, format sensitivity, and data dependencies. Existing technologies typically rely on static bandwidth allocation and unified transmission channels, failing to dynamically allocate resources based on file characteristics and real-time network conditions. This can easily lead to problems such as transmission delays for high-value files, inefficient bandwidth usage, and network congestion. Especially when facing large-scale, multi-type archival synchronization tasks, existing methods lack fine-grained hierarchical scheduling and global coordination capabilities for the transmission process, resulting in a need to improve overall transmission efficiency and resource utilization. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for synchronizing and scheduling archival data file transmission, in order to overcome the shortcomings of the prior art, realize refined hierarchical management and global resource optimization of archival data transmission, and improve the synchronization efficiency of high-value archives and the overall system throughput performance.

[0004] One embodiment of this application provides a method for synchronizing and scheduling the transmission of archive data files, the method comprising: Multi-dimensional feature extraction and structural parsing are performed on the archive data files to be synchronized to generate a structured feature description that includes file business attributes, format sensitivity and data dependency relationships; Based on the structured feature description, the transmission priority category of each file is dynamically determined and its network transmission overhead is estimated through a pre-trained classification model and a business rule engine. Based on the transmission priority category and the network transmission overhead, a virtual transmission channel cluster with differentiated service levels is self-organized and constructed, and an initial bandwidth weight and computing resource quota are pre-allocated to each channel; During data transmission, the throughput performance and node load of each virtual transmission channel are monitored in real time, and the cross-channel traffic load distribution strategy is dynamically adjusted based on the channel health to maintain optimal global transmission efficiency. Based on the dynamically adjusted traffic load distribution strategy, the multi-path parallel and sequential synchronous transmission of archive data is coordinated in the virtual transmission channel cluster to achieve synchronous scheduling of archive files among distributed nodes.

[0005] Another embodiment of this application provides a system for synchronizing and transmitting archival data files, the system comprising: The parsing module is used to extract multi-dimensional features and parse the structure of the archive data files to be synchronized, generating a structured feature description that includes file business attributes, format sensitivity, and data dependency relationships. The determination module is used to dynamically determine the transmission priority category of each file and estimate its network transmission overhead based on the structured feature description, through a pre-trained classification model and a business rule engine. The allocation module is used to self-organize and construct a virtual transmission channel cluster with differentiated service levels according to the transmission priority category and the network transmission overhead, and pre-allocate initial bandwidth weight and computing resource quota for each channel; The adjustment module is used to monitor the throughput performance and node load of each virtual transmission channel in real time during data transmission, and dynamically adjust the cross-channel traffic load distribution strategy based on the channel health to maintain optimal global transmission efficiency. The scheduling module is used to coordinate the multi-path parallel and sequential synchronous transmission of archive data in the virtual transmission channel cluster based on the dynamically adjusted traffic load distribution strategy, so as to realize the synchronous scheduling of archive files among distributed nodes.

[0006] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0007] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0008] Compared with existing technologies, the present invention provides a method for synchronizing and scheduling the transmission of archival data files, which can realize refined hierarchical management and global resource optimization of archival data transmission, and improve the synchronization efficiency of high-value archives and the overall throughput performance of the system. Attached Figure Description

[0009] Figure 1 A hardware structure block diagram of a computer terminal for a method of synchronizing and scheduling the transmission of archive data files, provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a method for synchronizing and scheduling the transmission of archival data files, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an archive data file synchronous transmission scheduling system provided in an embodiment of the present invention. Detailed Implementation

[0010] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0011] This invention first provides a method for synchronizing and scheduling the transmission of archive data files. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0012] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a method of synchronizing and scheduling the transmission of archival data files, provided in an embodiment of the present invention. (See diagram below.) Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0013] See Figure 2 The present invention provides a method for synchronizing and scheduling the transmission of archive data files, which may include the following steps: S201, perform multi-dimensional feature extraction and structure parsing on the archive data files to be synchronized, and generate a structured feature description that includes file business attributes, format sensitivity and data dependency relationships; Specifically, it can scan the set of archive data files to be synchronized, extract the basic metadata of each file, including file size, creation time, modification time, file type and storage path, and generate a basic metadata dataset; The set of archive data files to be synchronized is typically stored in a distributed file system across nodes. The scanning process employs a "distributed traversal + incremental synchronization" strategy to avoid the resource waste caused by a full scan. First, the root directory list of each node is obtained through a file system interface (such as a POSIX-based interface). Then, the files are recursively traversed in the order of "node-directory hierarchy". During the traversal, the unique identifier of each scanned file is recorded (using a combination of "node ID-file inode" to ensure cross-node uniqueness). For files that have been scanned in the past, only the modification time is checked; if it has not been updated, it is skipped, and only the metadata of newly added or modified files is extracted.

[0014] The extraction of basic metadata needs to be optimized for different file system characteristics to ensure data accuracy. File size is measured in bytes, accurate to a single byte. For very large files (≥10GB), the calculation is performed by summarizing the information stored in file system segments to avoid performance loss caused by direct reading.

[0015] The basic metadata dataset is organized in JSON format, with each file corresponding to a metadata record containing seven fields: file unique identifier, file size, creation timestamp, modification timestamp, file type, storage path, and extraction timestamp. After the metadata dataset is generated, CRC-32 verification is used to ensure data integrity. The checksum is appended to the end of the dataset and stored in a distributed metadata management node, supporting shared access across multiple nodes.

[0016] Based on the basic meta-dataset, natural language processing technology is used to parse the content summary and keywords of the file, and the business domain of the file is identified by combining the file type to generate a business attribute feature set; This step is the core of unlocking the business value of files. It uses NLP technology to extract semantic information from the file content and combines this with file type to achieve precise classification within the business domain. This provides a business-dimensional basis for subsequent transmission priority determination. The specific implementation method is as follows: Document content parsing requires differentiated extraction strategies based on file type to ensure the integrity of text information acquisition. For text files (such as .docx, .pdf, .txt), structured text is directly extracted using a file parsing library. For example, paragraph text in .docx files is obtained by parsing their internal XML structure, and copyable text in .pdf files is extracted using a PDF parsing engine, avoiding errors caused by OCR recognition. For image files (such as .jpg, .png, including scanned archival images), Optical Character Recognition (OCR) technology is first used to convert the image into text. The OCR engine's recognition accuracy must be ≥98%. For blurry or handwritten text, an enhanced recognition mode (improving clarity through image preprocessing) is enabled.

[0017] The content summary generation adopts a combination of "TextRank algorithm + key sentence extraction". The TextRank algorithm constructs a graph model of text sentences and calculates the importance score of each sentence. Sentences with higher scores are more likely to be core sentences. The algorithm parameters are set as follows: window size of 3 (i.e., each sentence is associated with the three sentences before and after it), damping coefficient of 0.85 (to control the iteration convergence speed), and 100 iterations. The top three sentences with the highest scores are extracted as core sentences, and then a summary (length controlled between 100-200 words) is generated through logical connections between sentences.

[0018] Keyword extraction employs a strategy of "TF-IDF algorithm + domain-specific thesaurus filtering." The TF-IDF algorithm calculates the term frequency (TF) and inverse document frequency (IDF) of each word. TF is the ratio of the number of times a word appears in the current document to the total number of words, and IDF is the logarithm of the total number of documents in the document set to the number of documents containing that word. TF-IDF value = TF × IDF; a higher value indicates that the word is more representative of the current document. The top 20 words with the highest TF-IDF values ​​are extracted and then filtered through a domain-specific thesaurus (containing over 1000 commonly used words in the field, such as "contract," "archives," "approval," "salary," and "project"), ultimately retaining 5-8 keywords. For example, in a "project progress report," the top 20 words with the highest TF-IDF values, after filtering by the domain-specific thesaurus, are "project progress, R&D, drawing review, deadlines, resource allocation, Zhang San."

[0019] Business domain attribution identification is based on the rule of "keyword matching + document type association". Ten core document business domains are predefined (such as "labor contracts, personnel files, project research and development, financial statements, administrative approvals", etc.), with each domain corresponding to a keyword set and a typical document type. During identification, the matching degree between the document's keywords and the keyword sets of each domain is calculated (number of matching keywords / total number of keywords in the domain). This is combined with the weight of the document type (e.g., .docx has a weight of 0.8 in the "labor contracts" domain and 0.2 in the "financial statements" domain). The domain with the highest overall score is the business domain attribution for that document. For example, the keyword "labor contract - Zhang San.docx" has a matching degree of 0.7 with the "labor contract" domain, a document type weight of 0.8, and an overall score of 0.7 × 0.8 = 0.56, which is higher than other domains, thus the business domain is determined to be "labor contracts".

[0020] The business attribute feature set integrates four categories of information: "content summary, keywords, business domain, and core business elements." Core business elements are extracted differently based on the business domain. For example, in the "labor contract" domain, "Party A, Party B, and contract term" are extracted, while in the "project R&D" domain, "project name, person in charge, and progress nodes" are extracted. The feature set format is associated with the basic metadata dataset, achieving a one-to-one mapping through file_id.

[0021] Based on the business attribute feature set, analyze the complexity and compression characteristics of file formats, calculate format conversion and compatibility indicators, and generate format sensitivity assessment results; The file format complexity assessment is conducted across three dimensions: structural hierarchy, encoding method, and special elements. Each dimension has a quantitative scoring standard (0-10 points, with higher scores indicating higher complexity). The final complexity score is the average of the scores from the three dimensions. Structural hierarchy assesses the internal organizational complexity of the file. Plain text files (.txt) have the simplest structure (1 point), multi-level XML structure files (.docx) score 5 points, and binary files containing nested objects (.dwg, .psd) score 9-10 points. Encoding method assesses the complexity of character encoding. UTF-8 encoded text files (2 points) are lower than GB2312 encoded files (4 points), and files containing mixed encoding of multiple languages ​​(such as Chinese, English, and Japanese) score 7 points. Special elements assess whether the file contains non-standard elements. Files without special elements (3 points) are lower than files containing embedded images or formulas (6 points), and files containing encryption or digital signatures (such as encrypted PDFs) score 10 points.

[0022] Compression characteristic analysis focuses on three core indicators: compression efficiency, decompression time, and compression loss. Data is obtained by performing standard compression tests on files (using the ZIP compression algorithm, with a compression level of 6 to balance compression ratio and speed). Compression efficiency is expressed as the compression ratio, which is calculated as: compression ratio = compressed file size / original file size. The smaller the ratio, the higher the compression efficiency. For example, a 20MB file compressed to 5MB has a compression ratio of 0.25, indicating excellent compression efficiency. Decompression time is the time taken to decompress the file on a standard configuration node (CPU i7-12700H, 16GB RAM), measured in milliseconds. For example, the decompression time for the aforementioned file is 800ms. Compression loss is used for unstructured files (such as images and audio), expressed as the quality loss rate before and after compression. It is calculated by comparing the peak signal-to-noise ratio (PSNR) of the image. When PSNR ≥ 30dB, the loss rate is < 5% (no significant loss), and when PSNR < 20dB, the loss rate is > 20% (severe loss).

[0023] The format conversion metric is used to evaluate the difficulty and quality of converting files between different formats. Three commonly used target formats (e.g., .docx to PDF, JPEG to PNG, .dwg to .dxf) are selected for conversion tests, and three parameters are calculated: conversion success rate, conversion time, and format distortion. Conversion success rate is the ratio of the number of successful conversions to the total number of tests. For example, if a .docx file is successfully converted to PDF, TXT, and HTML, the success rate is 100%. Conversion time is the average time for a single conversion; for example, if .docx to PDF takes 1200ms, the conversion time is 1200ms. Format distortion is determined by comparing the consistency of the core content of the files before and after conversion (character matching rate for text files, PSNR for images). For example, after converting .docx to TXT, the character matching rate is 99.5%, and the distortion is 0.5%. For files that cannot be converted (e.g., encrypted files without decryption permissions), the conversion success rate is 0%, and the distortion is recorded as 100%.

[0024] The compatibility index evaluates the ability to open and read files under different software and hardware environments. Five mainstream file processing software programs and three different node configurations (high-performance, normal, and low-performance) were selected for compatibility testing. Two parameters were calculated: "open success rate" and "reading time". The open success rate is the ratio of the number of software and node combinations that successfully opened the file to the total number of combinations (5 × 3 = 15). For example, if a .pdf file is successfully opened in 14 combinations, the success rate is approximately 14 / 15 ≈ 93.3%. The reading time is the average time it takes for the software to complete content rendering after opening the file. For example, the reading time for this .pdf file on a normal node is 1500ms.

[0025] The format sensitivity assessment result is a comprehensive rating of the above indicators, mapping the four core parameters of "complexity score, compression ratio, conversion power, and compatibility success rate" to sensitivity levels (low sensitivity, medium sensitivity, and high sensitivity). The mapping rules are as follows: Low sensitivity is defined as follows: complexity score < 4, compression ratio < 0.3, conversion power ≥ 90%, and compatibility success rate ≥ 95%; Medium sensitivity is defined as follows: complexity score 4-7, compression ratio 0.3-0.6, conversion power 70%-90%, and compatibility success rate 85%-95%; High sensitivity is defined as follows: complexity score > 7, compression ratio > 0.6, conversion power < 70%, and compatibility success rate < 85%.

[0026] By combining the business attribute feature set and format sensitivity assessment results, the reference relationships and version dependencies between files are queried through a graph database to construct a data dependency graph, and finally, a structured feature description is generated.

[0027] Document dependencies mainly fall into two categories: "reference relationships" and "version dependencies." A reference relationship means that file A explicitly references content from file B (e.g., a chart inserted in a report comes from another Excel file). A version dependency means that file B is a newer version of file A (e.g., "Project Plan V1.0.docx" and "Project Plan V2.0.docx"). This relationship information is usually implicit in the file content (e.g., reference paths, comments), filenames (e.g., containing "V1.0" and "V2.0"), and business process records. It needs to be stored and queried in a structured manner using a graph database. Graph databases use an attribute graph model, defining two types of elements: "file nodes" and "dependency edges." The attributes of a file node are file_id and filename, while the attributes of a dependency edge are relationship type (reference / version dependency), dependency strength (1-10 points, assessed based on reference frequency or version update importance), and associated location (e.g., the specific page number or paragraph of the reference relationship).

[0028] The query for data dependencies employs a triple strategy of "content parsing + metadata matching + business rules". First, it parses the reference information in the file content, such as the "Insert Object" path in a Word file or the "Attachment Reference" link in a PDF file, extracting the filename or storage path of the referenced file. For example, parsing "Reference Object Path: Node-A: / project / data / Progress Chart.xlsx" from "Project Report.docx" confirms that it references "Progress Chart.xlsx". Second, it matches the association information in the file metadata, such as filenames containing the same business identifier (e.g., "Labor Contract - Zhang San" and "Resignation Application - Zhang San" share the "Zhang San" identifier) ​​and the order of creation time (in version dependencies, the newer version was created later). For example, "Project Plan V2.0.docx" was created later than "V1.0", and the filename prefixes are the same, indicating a version dependency. Finally, it applies business rules to supplement implicit relationships. For example, documents in the "Labor Contract" domain often depend on "Employee Basic Information Table.xlsx", and even if not explicitly referenced in the content, the association needs to be established based on business rules.

[0029] The query process for graph databases requires different query statements designed for two types of dependency relationships. Reference relationship queries start with the current file's `file_id`, querying all outgoing edges of the "reference" type to obtain the associated file nodes. Version dependency queries use the current file's business identifier (such as "project plan") as the keyword, querying all file nodes containing that identifier, sorting them by creation time, and then determining the version dependency relationship. For example, when querying the dependency relationship of "Project Report.docx" (file_id: Node-A-78901234), the reference relationship query returns "Progress Chart.xlsx" (file_id: Node-A-78904321), while the version dependency query returns no results (this report was generated for the first time and there are no updated versions). When querying the dependency relationship of "Project Plan V2.0.docx", the version dependency query returns "Project Plan V1.0.docx" (file_id: Node-A-78905678), with a dependency strength of 10 (major update).

[0030] A data dependency graph is a visual and structured representation of query results. The graph centers on file nodes, using edges of different colors to distinguish between reference relationships (blue) and version dependencies (red). The thickness of the edges corresponds to the dependency strength (the stronger the dependency, the thicker the edge). The structured data of the graph is stored in JSON format, containing two parts: a "node list" and an "edge list." The node list records the file_id, filename, and business domain for each file node. The edge list records the start-point file_id, end-point file_id, relationship type, and dependency strength for each edge. For files with multiple dependencies (e.g., referencing three files simultaneously and having two historical versions), it is essential to ensure that all relationships are completely recorded to avoid omissions.

[0031] Structured feature description is the final integration of all preceding feature information. The integration scope includes the basic metadata dataset, business attribute feature set, format sensitivity evaluation results, and data dependency graph. All information is aligned and linked using `file_id` to form a complete file feature record. The record format uses JSON-LD (Related Data Markup Language) to ensure the machine-readable and semantically relevant feature information, facilitating direct access by subsequent pre-trained models and business rule engines. Redundant information must be removed during the integration process. For example, `file_type` in the basic metadata and `business_domain` in the business attributes are not redundant and must be retained in their entirety; while `file_id` in different modules must be consistent to avoid conflicts. This description comprehensively covers the file's basic attributes, business characteristics, format features, and dependencies, providing a comprehensive decision-making basis for subsequent transmission scheduling.

[0032] S202, Based on the structured feature description, the transmission priority category of each file is dynamically determined and its network transmission overhead is estimated through a pre-trained classification model and a business rule engine. Specifically, structured feature descriptions can be input into a pre-trained classification model, which is trained based on historical transmission logs, and outputs preliminary priority scores and transmission time predictions for files to generate model prediction results. Structured feature descriptions encompass various types of unstructured and structured information, including basic file attributes, business features, format characteristics, and dependency relationships. Before inputting these features into the model, standardized preprocessing is required to convert different types of features into numerical vectors recognizable by the model. For textual features (such as the business domain "labor contract" and content summaries), word embedding technology (Word2Vec) is used to convert them into 128-dimensional vectors. The business domain is mapped to a fixed vector through a domain dictionary (e.g., "labor contract" corresponds to vectors [0.12, 0.35, ..., 0.08]). For numerical features (such as file size 20,971,520 bytes and format complexity score 5.2), Min-Max normalization is used to compress them to the [0,1] interval. For Boolean features (such as whether a file is encrypted), they are directly converted to binary values ​​of 0 or 1. For graph structure features (data dependency graphs), a 32-dimensional global feature vector is extracted from the graph using a graph neural network (GCN) to characterize the file's dependency strength and association complexity. Finally, all feature vectors are concatenated to form a unified input feature vector of length 256.

[0033] The pre-trained classification model adopts a hybrid architecture of "BERT-Base + fully connected layer". The BERT-Base model is responsible for capturing semantic relationships and high-order interactions between features (especially suitable for semantically rich features such as business attributes), while the fully connected layer is responsible for mapping the output of BERT to specific prediction metrics. The training data of the model comes from the archive data transmission logs of the past 6 months, which record more than 100,000 file transmission samples. Each sample contains three core fields: "input feature vector, manually labeled priority score (1-100 points), and actual transmission time". The priority score is evaluated by business personnel based on the urgency and business value of the file, and the transmission time is the actual time taken (in seconds) for the file to travel from the source node to the target node. During training, the model adopts a dual-task learning strategy, simultaneously optimizing the "priority score regression task" and the "transmission time prediction task". The loss function is mean squared error (MSE), i.e., Loss=(1 / N)×Σ[(y_pred-y_true)²], where N is the number of samples, y_pred is the model prediction value, and y_true is the true value. The training parameters were set as follows: batch size 32, initial learning rate 5e-5, cosine annealing strategy to dynamically adjust the learning rate, 100 iterations, and training was stopped when the loss function of the validation set no longer decreased for 5 consecutive iterations to avoid overfitting.

[0034] During model prediction, the preprocessed 256-dimensional feature vector is input into the trained model. The BERT layer outputs a 768-dimensional feature representation, which is then mapped through a fully connected layer (containing two hidden layers with 256 and 64 neurons respectively, using ReLU activation function). The final output includes two prediction results: a preliminary priority score (1-100 points, higher scores indicate higher priority) and a predicted transmission time (in seconds). For example, the input feature vector of a file named "Labor Contract - Zhang San.docx" will output a preliminary priority score of 65 and a predicted transmission time of 20 seconds after model processing. To ensure prediction reliability, a prediction confidence score (calculated based on the Softmax probability of the model's output layer) is also appended to the model output. Samples with a confidence score below 0.7 are marked as "low confidence" and will be given special attention during subsequent rule engine correction. The model prediction results are stored in JSON format and associated with the file's file_id.

[0035] The model prediction results are input into the business rule engine, and the preset business rules are applied to correct the priority score and generate the rule-corrected priority. The business rules engine adopts a "rule base + inference engine" architecture. The rule base stores preset business rules, which are formulated by experts in the field of document management based on business processes, compliance requirements, and urgency. These rules are categorized into four types: "urgency rules, business value rules, format-specific rules, and dependency / association rules." Each type of rule includes triggering conditions and correction strategies. Urgency rules target the time characteristics of documents, such as "administrative approval documents created within 24 hours receive an additional 20 points in priority score," and "documents modified more than 30 days ago and without business relevance receive a deduction of 15 points in priority score." Business value rules are based on the business domain of the documents, such as "financial statement documents receive an additional 15 points in priority score," and "temporary draft documents receive a deduction of 10 points in priority score." Format-specific rules target highly sensitive format documents, such as "documents with a highly sensitive format receive an additional 10 points in priority score (to avoid format corruption due to transmission delays)." Dependency / association rules are based on data dependencies, such as "core documents referenced by three or more documents receive an additional 25 points in priority score." Rules have higher priority than model predictions. When multiple rules are triggered simultaneously, a weighted summation method is used for correction. The weight of each rule is set by experts (range 0.8-1.2, core rule weight 1.2, ordinary rule weight 1.0).

[0036] The inference engine's workflow is "rule matching → condition verification → weight calculation → score correction." First, based on the structured feature description associated with the file_id in the model's prediction results, it extracts the key information required for rule matching (such as business domain, creation time, format sensitivity, and citation count). Then, it traverses the rule base, matching all rules that meet the conditions. For example, the business domain of "Labor Contract - Zhang San.docx" is "Labor Contract" (matching the business value rule "Personnel-related documents add 10 points"), the creation time is 2 days ago (not matching the urgency rule), and the format sensitivity is medium sensitivity (not matching the special format rule). The file is cited by one file (not matching a high-citation association rule), and the model prediction confidence is 0.82 (high confidence, no further adjustment needed). Next, the validity of the rule triggering conditions is verified, such as confirming that "labor contract" belongs to the personnel-related field and that the citation count is accurate. Then, a correction score is calculated based on the rule weight. The rule weight triggered by this file is 1.0, and the correction score is 10 points. Finally, the initial priority score is corrected using the formula: Corrected Priority = Initial Priority Score + Σ(Rule Correction Score × Rule Weight). An upper limit of 100 points and a lower limit of 1 point are set to prevent scores from exceeding a reasonable range. For example, if the initial score for this file is 65 points, the corrected score would be 65 + 10 × 1.0 = 75 points.

[0037] For low-confidence samples where the model predicts a confidence level below 0.7, the rule engine will activate an "enhanced correction" mechanism. In addition to regular rule matching, a "manual intervention prompt rule" will be added. This means the corrected priority score will have a ±5-point fluctuation range, and a manual review request will be generated and pushed to the file administrator's client. For example, a "engineering drawing.dwg" file with a model prediction confidence level of 0.65 initially has a priority score of 50. After rule correction, the score becomes 50 + 15 (highly sensitive format) + 20 (core reference file) = 85, and the final output is "85 points (±5 points, manual review required)". The priority result after rule correction needs to record the correction process, including the triggered rule name, correction score, weight, and confidence level handling method, forming a correction log for subsequent traceability and rule optimization. For example, "file_id:Node-A-12345678, correction rule: add 10 points to personnel-related documents (weight 1.0), initial score 65, corrected score 75, confidence level handling: no adjustment for high confidence."

[0038] Based on the priority after rule correction, combined with file size and historical network performance data, a bandwidth estimation algorithm is used to calculate the transmission cost of each file and generate an estimated value of network transmission cost. The core evaluation metrics for transmission overhead include three dimensions: transmission time, bandwidth utilization, and resource consumption cost. Transmission time is fundamental, bandwidth utilization reflects the intensity of network resource usage, and resource consumption cost is a comprehensive quantitative value combining computing and storage resources. Calculating these metrics requires obtaining three types of basic data: rule-corrected priority (affecting the allocation priority of transmission resources and indirectly affecting actual bandwidth), file size (a normalized value that needs to be restored to its original byte count; for example, a normalized value of 0.0019 corresponds to an original size of 20 × 1024 × 1024 = 20,971,520 bytes), and historical network performance data (statistical data on network bandwidth, latency, and packet loss rate between the source and target nodes over the past 7 days, stored in a time-series database).

[0039] The preprocessing of historical network performance data employs a "sliding window averaging method," using a 5-minute time window to calculate the average bandwidth (in Mbps, 1 Mbps = 1024 × 1024 / 8 Byte / s), average latency (in ms), and average packet loss rate (in %) within each window. Then, the historical data that best matches the current transmission time period (e.g., weekday 9:00-10:00) is selected as a reference. For example, if the current time is Monday 9:30, the average data from the past three Mondays from 9:00-10:00 is selected, resulting in an average bandwidth of 100 Mbps, an average latency of 50 ms, and an average packet loss rate of 0.5%. If historical data is insufficient (e.g., for the first transmission), default network parameters are used: bandwidth 50 Mbps, latency 100 ms, and packet loss rate 1%.

[0040] The bandwidth estimation algorithm adopts a "priority-based dynamic bandwidth allocation model". The core idea is that the higher the priority of a file, the closer the actual bandwidth that can be allocated is to the average network bandwidth, while the lower the priority, the more likely the allocated bandwidth will be compressed. The algorithm first defines a priority coefficient α, which is positively correlated with the corrected priority score. The calculation formula is α = corrected priority / 100. For example, a priority score of 75 corresponds to α = 0.75. Then, it calculates the effective bandwidth B_eff, B_eff = B_avg × α × (1 - packet loss rate), where B_avg is the average network bandwidth, and the packet loss rate is converted to a decimal (e.g., 0.5% = 0.005). This formula reflects the impact of priority on bandwidth and the loss due to packet loss. For example, if B_avg = 100 Mbps, α = 0.75, and the packet loss rate is 0.005, the calculated B_eff = 100 × 0.75 × (1 - 0.005) = 74.625 Mbps. Next, it calculates the transmission time T, T = file size / (B_eff × 1024 × 1024 / 8), converting the file size to bytes. The bandwidth is converted to bytes per second. For example, for a file size of 20,971,520 bytes, B_eff = 74.625 Mbps. The calculated time is T = 20971520 / (74.625 × 1024 × 1024 / 8) ≈ 20971520 / 9830400 ≈ 2.13 seconds. This time takes into account the bandwidth allocated by priority and packet loss, and is closer to reality than the model's prediction of 20 seconds. Finally, the bandwidth utilization rate and resource consumption cost are calculated. The bandwidth utilization rate = B_eff / B_avg × 100% = 74.625%, and the resource consumption cost = T × (CPU utilization coefficient + memory utilization coefficient). The CPU utilization coefficient (0.1 yuan / second) and memory utilization coefficient (0.05 yuan / second) are determined by the resource pricing strategy. For example, the cost = 2.13 × (0.1 + 0.05) = 0.32 yuan.

[0041] The estimated network transmission overhead is a comprehensive encapsulation of the above indicators, with "transmission time (seconds), bandwidth utilization (%), resource consumption cost (yuan), and effective bandwidth (Mbps)" as core fields, along with additional calculation basis (such as average network bandwidth and priority coefficient). For very large files (file size > 10GB), additional fragmentation transmission overhead needs to be calculated. The file is divided into 1GB fragments, and the transmission time of each fragment is calculated. The total overhead is the sum of the overhead of each fragment, while also considering the transmission gap between fragments (each fragment adds a 0.1-second gap). For example, a 15GB file is divided into 15 fragments, and the transmission time of each fragment is 5 seconds, so the total time = 15 × 5 + 14 × 0.1 = 76.4 seconds.

[0042] Based on the estimated network transmission overhead and the priority after rule correction, a clustering algorithm is used to divide the files into different transmission priority categories, generating transmission priority category division results.

[0043] The choice of clustering algorithm needs to consider the characteristics of "low feature dimensionality and clear category boundaries." The K-Means clustering algorithm is adopted. This algorithm iteratively divides the data into K clusters, with high similarity among the data within each cluster. The input features for clustering are "priority after rule correction" and "transmission time in the estimated network transmission cost." These two features represent the "business importance" and "transmission cost" of the file, respectively, and are the core basis for classification and scheduling. The input features need to be standardized to eliminate differences in units. Priority (1-100) and transmission time (0.1-1000 seconds) are both standardized using Z-Score.

[0044] The K-value (number of clusters) is determined using the "elbow rule." By calculating the silhouette coefficient for different K values, a silhouette coefficient closer to 1 indicates better clustering. The K-value with the largest silhouette coefficient is selected as the final number of clusters. Considering the actual business needs of archive transmission, priority categories are typically divided into four classes (high priority, medium priority, low priority, and normal). Therefore, we initially assume K=4 and calculate the silhouette coefficient. For example, when clustering 1000 file samples, the silhouette coefficient is 0.82 when K=4, 0.75 when K=3, and 0.78 when K=5. Therefore, K=4 is determined.

[0045] The specific process of K-Means clustering is as follows: First, randomly select four points from the samples as initial cluster centers. For example, select four samples with standardized features of (1.0, -0.5), (0.5, 0.0), (-0.2, 0.5), and (-0.8, 1.0). Then, calculate the Euclidean distance from each sample to the four cluster centers and assign the sample to the nearest cluster. For example, a sample with standardized features of (0.75, -0.98) has a distance of √[(0.75- ] to the first center. [1.0)²+(-0.98+0.5)²]≈√[0.0625+0.2304]≈0.54, which is farther from other centers, so it is assigned to the first cluster; then the mean of features of all samples in each cluster is calculated as the new cluster center, for example, the new center of the first cluster is (0.8,-0.7); the process of distance calculation, sample allocation and center update is repeated until the change in cluster center is less than the preset threshold (such as 0.001) or the number of iterations reaches 100, and then the clustering stops.

[0046] After clustering, each cluster needs semantic naming and feature definition. Based on business requirements, the four clusters are named "High-Priority Low-Consumption," "High-Priority High-Consumption," "Medium-Priority Regular," and "Low-Priority Deferred." The features of each category are derived from the mean features of the samples within the cluster (by inverting the standardized features back to the original features). The characteristics of the High-Priority Low-Consumption category are: corrected priority ≥ 85 points, transmission time ≤ 5 seconds, corresponding to core urgent documents (such as financial statements requiring approval that day); the High-Priority High-Consumption category is: priority ≥ 80 points, transmission time > 30 seconds (mostly very large files such as engineering drawings); the Medium-Priority Regular category is: priority 50-80 points, transmission time 5-30 seconds (such as ordinary labor contracts, project reports); and the Low-Priority Deferred category is: priority < 50 points, arbitrary transmission time (such as historical archive backups, temporary drafts). For example, the corrected priority of "Labor Contract - Zhang San.docx" is 75 points, with a transmission time of 2.13 seconds, and it is assigned to the Medium-Priority Regular category.

[0047] The transmission priority classification results must include the category to which each file belongs, a description of the category characteristics, and scheduling suggestions. The scheduling suggestions provide guidance for subsequent transmission channel configuration. For example, for the "High Priority, Low Consumption" category, it is recommended to "dedicate a high-bandwidth channel and transmit with priority," while for the "Low Priority, Lazy" category, it is recommended to "share a low-bandwidth channel and transmit during idle times." The classification results should be summarized in tabular form (text description), for example: "file_id:Node-A-12345678, Category: Medium Priority, Regular, Category Characteristics: Priority 50-80 points, Transmission Time 5-30 seconds, Scheduling Suggestion: Allocate a medium-bandwidth channel and transmit sequentially; file_id:Node-A-78901234 (Engineering Drawing), Category: High Priority, High Consumption, Scheduling Suggestion: Allocate a dedicated high-bandwidth channel and transmit in fragments in parallel." The classification results must be synchronized to the metadata management node to ensure that subsequent steps can quickly obtain the file priority category information.

[0048] S203, Based on the transmission priority category and the network transmission overhead, a virtual transmission channel cluster with differentiated service levels is self-organized and constructed, and an initial bandwidth weight and computing resource quota are pre-allocated to each channel; Specifically, based on the transmission priority category classification results, the required number and service level of virtual transmission channels can be determined, with each service level corresponding to a priority category, thus generating a virtual transmission channel configuration scheme. The transmission priority classification results typically include four typical categories (high priority low power consumption, high priority high power consumption, medium priority regular, and low priority bufferable). The number of channels is determined according to the principle of "one main channel per category + flexible backup channel"—at least one main channel is allocated to each priority category, and backup channels are configured according to the number of files in that category and the concurrent transmission requirements. The number of backup channels = ceil(number of main channels × concurrency coefficient), where the concurrency coefficient ranges from 0.3 to 0.8. The more files and the higher the transmission frequency, the larger the concurrency coefficient. For example, there are 200 high-priority, low-consumption files, with approximately 30 concurrent transmission requests per minute. One main channel is configured with a concurrency coefficient of 0.6, and one backup channel is required (ceil(1×0.6) = 1). This category ultimately has 2 channels. Although there are only 30 high-priority, high-consumption files, each file transmission consumes a large amount of bandwidth. One main channel is configured with a concurrency coefficient of 0.8, and one backup channel is required. There are 500 medium-priority, regular files, with one main channel, a concurrency coefficient of 0.5, and one backup channel. There are 1000 low-priority, cacheable files with low transmission requirements. Only one main channel is configured, with no backup channels. The total number of channels is ultimately determined to be 2+2+2+1=7.

[0049] The service level definition needs to be based on a differentiated system built around three core indicators: latency, bandwidth guarantee, and packet loss tolerance. Each level corresponds to a set of clear performance thresholds, forming a strict mapping with priority categories: High-priority low-consumption corresponds to "Platinum" service, with core requirements of low latency (transmission latency ≤50ms) and zero packet loss guarantee (packet loss rate ≤0.01%), suitable for small files that are time-sensitive, such as urgent approval documents; High-priority high-consumption corresponds to "Gold" service, with core requirements of high bandwidth guarantee (dedicated bandwidth ≥50Mbps) and low packet loss (packet loss rate ≤0.1%), with latency relaxed to ≤200ms, suitable for large core files such as engineering drawings; Medium-priority regular corresponds to "Silver" service, with indicators of latency ≤500ms, packet loss rate ≤0.5%, and shared bandwidth ≥20Mbps, suitable for ordinary labor contracts, project reports, etc.; Low-priority deferred corresponds to "Bronze" service, with no fixed bandwidth guarantee (shared remaining bandwidth), latency ≤1000ms, and packet loss rate ≤1%, suitable for non-urgent files such as historical archive backups. The service level also needs to clarify the fault recovery mechanism. Platinum and Gold levels require that the time for automatic switching to the backup channel during a fault is ≤1 second, Silver level is ≤3 seconds, and Bronze level has no forced switching requirement and only records fault logs.

[0050] The virtual transmission channel configuration scheme is a structured integration of the number of channels, service levels, and performance indicators, comprising three parts: "Channel Configuration Overview," "Details of Channels by Level," and "Resource Reservation Instructions." The channel configuration overview clearly defines the total number of channels, the priority categories of coverage, and the core design principles. The details of each service level are described individually, including the channel identifier prefix, the number of primary and backup channels, performance indicator thresholds, applicable file types, and typical examples. For example, the Platinum-level channel details are: "Channel identifier prefix: VT-PT, 1 primary channel, 1 backup channel, latency ≤50ms, packet loss rate ≤0.01%, applicable documents: administrative approval documents within 24 hours, example: Node-A: / approval / 20251018-001.docx." The resource reservation instructions specify the proportion of network resources reserved for each channel level: Platinum and Gold levels reserve a combined 60% bandwidth, Silver level 30%, and Bronze level 10%. The configuration scheme is stored in XML format for easy machine parsing during subsequent channel creation.

[0051] Based on the virtual transport channel configuration scheme, virtual channel instances are dynamically created on distributed network nodes, and a unique identifier and initial routing path are assigned to each instance to generate a set of virtual transport channel instances; The virtual transport channel is technically implemented based on a converged architecture of Software-Defined Networking (SDN) and Virtual Private Network (VPN). SDN is responsible for the dynamic creation, routing control, and resource scheduling of the channel, while VPN ensures the security and isolation of data transmission within the channel. The creation process is led by the SDN controller. The controller first parses the virtual transport channel configuration scheme, extracting information such as the number of channels and performance requirements for each level. Then, it sends channel creation instructions to the SDN switches of the distributed nodes. These instructions include parameters such as the channel's service level, encapsulation protocol (using GRE encapsulation to ensure cross-network compatibility), and encryption algorithm (AES-256 encryption for Platinum and Gold levels, and AES-128 encryption for Silver and Bronze levels). For example, when creating a Platinum-level primary channel, the SDN controller sends instructions to the switches of the source node Node-A and the destination node Node-D to configure a GRE tunnel, enable AES-256 encryption, and set a QoS (Quality of Service) policy to ensure latency ≤50ms.

[0052] Each virtual tunnel instance's unique identifier (ID) uses a structured format of "fixed prefix-service class abbreviation-source node ID-serial number" to ensure uniqueness and readability. For example, the ID of the Platinum-level primary tunnel is "VT-PT-NodeA-01", where "VT" is the fixed prefix for the virtual tunnel, "PT" is the abbreviation for Platinum, "NodeA" is the source node identifier, and "01" is the sequence number of the tunnel at that level. The corresponding backup tunnel ID is "VT-PT-NodeA-02". These identifiers are used not only for tunnel management and monitoring but also as identifiers for the VPN tunnel, ensuring that data is accurately routed to the target tunnel.

[0053] The initial routing path planning employs a "multi-constraint path selection algorithm," which uses "minimum transmission delay, maximum link bandwidth, and minimum node load" as constraints to select the optimal path from all available links from the source node to the target node. Before path planning, the SDN controller obtains the topology and real-time link status (bandwidth, delay, load) of the distributed network through a link probing protocol (such as LLDP) and constructs a link state matrix. For example, there are three available links from source node Node-A to destination node Node-D: Node-A→Node-B→Node-D (bandwidth 80Mbps, latency 30ms, Node-B load 40%), Node-A→Node-C→Node-D (bandwidth 100Mbps, latency 45ms, Node-C load 35%), and Node-A→Node-E→Node-D (bandwidth 60Mbps, latency 60ms, Node-E load 50%). Considering the constraint of "latency ≤ 50ms" for Platinum-level channels, the third link is excluded. Taking into account bandwidth and load factors, Node-A→Node-C→Node-D is finally selected as the initial routing path. This path has a latency of 45ms, sufficient bandwidth, and low node load.

[0054] Once the routing path is determined, the SDN controller configures the path information into the flow tables of all switches involved in the channel. Flow table entries include information such as channel ID, source IP, destination IP, and next-hop node, ensuring that data frames within the channel can be forwarded according to the planned path. Simultaneously, a path backup mechanism is configured for each path. When the primary path fails, it automatically switches to the backup path. For example, the primary path for a Platinum-level channel is Node-A→Node-C→Node-D, and the backup path is Node-A→Node-B→Node-D. The switching trigger condition is a link latency exceeding 50ms or a packet loss rate exceeding 0.01%.

[0055] The virtual transport channel instance set is a summary of information for all created channel instances. It includes core information such as each instance's unique ID, service level, source node, destination node, initial routing path, encapsulation protocol, encryption algorithm, and QoS parameters. This information is stored in the SDN controller's database in JSON format for easy real-time querying and management. After the instance set is created, the SDN controller performs connectivity tests on each channel, sending test data packets to verify the smoothness of data transmission and whether performance indicators meet the standards. Channels that fail to meet the standards will be recreated or have their routing paths adjusted.

[0056] Based on the estimated network transmission overhead, an initial bandwidth weight is calculated for each virtual transmission channel instance. The weight is proportional to the priority and transmission overhead of the file served by the channel, and an initial bandwidth weight allocation table is generated. The initial bandwidth weight calculation follows the principle of "priority-driven, transmission overhead-assisted". The weight value is positively correlated with the average priority score of the files served by the channel and also positively correlated with the average transmission time (the core indicator of transmission overhead). The higher the priority, the more bandwidth is needed for the file; the longer the transmission time (usually corresponding to a larger file), the more bandwidth is needed to shorten the transmission time. Before calculation, it is necessary to statistically analyze the files served by each virtual transmission channel instance, obtain the rule-corrected priority score and the estimated network transmission overhead (transmission time) of all files in the channel, and calculate the average priority score (P_avg) and average transmission time (T_avg). For example, the Platinum-level main channel (VT-PT-NodeA-01) serves 200 high-priority, low-consumption files with an average priority score of 92 and an average transmission time of 2.5 seconds; the Gold-level main channel (VT-GD-NodeA-01) serves 30 high-priority, high-consumption files with an average priority score of 83 and an average transmission time of 45 seconds; the Silver-level main channel (VT-SL-NodeA-01) serves 500 medium-priority, regular files with an average priority score of 65 and an average transmission time of 15 seconds; and the Bronze-level channel (VT-BZ-NodeA-01) serves 1000 low-priority, cacheable files with an average priority score of 35 and an average transmission time of 10 seconds.

[0057] The bandwidth weight is calculated using a weighted summation model, with the formula: W=(P_avg / P_max)×α+(T_avg / T_max)×β, where W is the initial weight of the channel, P_max is the maximum average priority score of all channels (92 points here), T_max is the maximum average transmission time of all channels (45 seconds here), α and β are weight coefficients, α=0.6 (emphasizing the dominant role of priority), β=0.4 (transmission overhead plays a supporting role), and the sum of the weight coefficients is 1 to ensure the normalization basis of the calculation results. For example, to calculate the weight of the Platinum-level main channel: W_PT = (92 / 92) × 0.6 + (2.5 / 45) × 0.4 = 1 × 0.6 + 0.0556 × 0.4 ≈ 0.6222; the weight of the Gold-level main channel: W_GD = (83 / 92) × 0.6 + (45 / 45) × 0.4 ≈ 0.8913 × 0.6 + 1 × 0.4 ≈ 0.9348; the weight of the Silver-level main channel: W_SL = (65 / 92) × 0.6 + (15 / 45) × 0.4 ≈ 0.7065 × 0.6 + 0.3333 × 0.4 ≈0.5572; Bronze-level channel weight: W_BZ=(35 / 92)×0.6+(10 / 45)×0.4≈0.3804×0.6+0.2222×0.4≈0.3171; The weight calculation of the backup channel is consistent with that of the corresponding main channel. For example, the weight of the Platinum-level backup channel (VT-PT-NodeA-02) is also 0.6222, the weight of the Gold-level backup channel (VT-GD-NodeA-02) is 0.9348, and the weight of the Silver-level backup channel (VT-SL-NodeA-02) is 0.5572.

[0058] To ensure that the weights can be directly mapped to the bandwidth allocation ratio, the initial weights of all channels need to be normalized. The normalized weight W_norm = W / ΣW, where ΣW is the sum of the initial weights of all channels. First, calculate the total weight: ΣW = 0.6222 (PT primary) + 0.6222 (PT backup) + 0.9348 (GD primary) + 0.9348 (GD backup) + 0.5572 (SL primary) + 0.5572 (SL backup) + 0.3171 (BZ primary) ≈ 4.5455. Then, calculate the normalized weights for each channel: PT primary = 0.6222 / 4.5455 ≈ 0.1369 (13.69%), PT backup = 0.1369, GD primary = 0.9348 / 4.5455 ≈ 0.2057 (20.57%), GD backup = 0.2057, SL primary = 0.5572 / 4.5455 ≈ 0.1226 (12.26%), SL backup = 0.1226, BZ primary = 0.3171 / 4.5455 ≈ 0.0697 (6.97%). The normalized weights directly correspond to the bandwidth allocation ratio of the channels. If the total available bandwidth of the distributed network is 100Mbps, then the allocable bandwidth for the PT primary channel = 100 × 13.69% ≈ 13.69Mbps, the GD primary channel = 100 × 20.57% ≈ 20.57Mbps, and so on.

[0059] The initial bandwidth weight allocation table is a structured presentation of the above calculation results, including fields such as channel ID, service level, file type served, average priority score, average transmission time, initial weight, normalized weight, and estimated allocated bandwidth (Mbps). For example, a table fragment (text description) is as follows: "Channel ID: VT-PT-NodeA-01, Service Level: Platinum, File Type: High Priority Low Consumption, P_avg: 92 points, T_avg: 2.5 seconds, Initial Weight: 0.6222, Normalized Weight: 13.69%, Estimated Bandwidth: 13.69Mbps; Channel ID: VT-GD-NodeA-01, Service Level: Gold, File Type: High Priority High Consumption, P_avg: 83 points, T_avg: 45 seconds, Initial Weight: 0.9348, Normalized Weight: 20.57%, Estimated Bandwidth: 20.57Mbps." After the allocation table is generated, it is synchronized to the SDN controller and bandwidth management module as the direct basis for bandwidth allocation. At the same time, the "initial weight" attribute is marked, indicating that the weight will be dynamically adjusted according to the channel health during transmission.

[0060] Based on the initial bandwidth weight allocation table, computing resources are reserved from the global resource pool for each virtual transmission channel instance, and a computing resource quota allocation table is generated.

[0061] The global resource pool is an aggregation of computing resources from distributed nodes, encompassing the CPU core count, memory capacity, disk I / O bandwidth, and other resources of all participating nodes. It is centrally managed and scheduled by the resource management system. Before reserving resources, the resource management system needs to calculate the total available resources in the global resource pool. For example, the total available CPU cores are 100, the total available memory is 512GB, and the total available disk I / O bandwidth is 5000MB / s. Simultaneously, a resource reservation safety threshold is set (reserving 20% ​​of resources as an emergency buffer to avoid resource exhaustion). Therefore, the actual resources available for allocation are 80 CPU cores, 409.6GB of memory, and 4000MB / s I / O bandwidth.

[0062] The allocation of computing resource quotas is directly linked to bandwidth weights, following the principle that "bandwidth weights determine resource quota ratios." A higher bandwidth weight indicates a more important file being served or a higher transmission overhead, requiring more computing resources to support data encapsulation, encryption, and forwarding operations. The allocated resource types include CPU quotas (for data processing and protocol encapsulation), memory quotas (for data caching to reduce I / O latency), and disk I / O quotas (for file reading and writing to ensure data access efficiency). The quota ratios for these three types of resources are consistent with the channel's normalized weights, while also being fine-tuned based on file characteristics (e.g., high-priority, high-consumption files require more I / O resources).

[0063] The formula for calculating CPU quota is: CPU quota (cores) = total allocable CPUs × normalized weight × file complexity coefficient, where the file complexity coefficient is determined based on the format sensitivity of the files served by the channel - the coefficient is 1.2 for high-sensitivity formats (such as engineering drawings.dwg), 1.0 for medium-sensitivity formats (such as labor contracts.docx), and 0.8 for low-sensitivity formats (such as plain text.txt). For example, the Gold-level main channel (VT-GD-NodeA-01) serves high-priority, high-consumption files (highly sensitive to format), with a normalized weight of 20.57%, and can be allocated 80 CPU cores. The CPU quota = 80 × 20.57% × 1.2 ≈ 80 × 0.2057 × 1.2 ≈ 19.75 cores, rounded up to 20 cores; the Platinum-level main channel (VT-PT-NodeA-01) serves high-priority, low-consumption files (medium sensitive to format), with a quota = 80 × 13.69% × 1.0 ≈ 10.95 cores, rounded up to 11 cores; the Bronze-level channel (VT-BZ-NodeA-01) serves low-priority, cacheable files (low sensitive to format), with a quota = 80 × 6.97% × 0.8 ≈ 80 × 0.0697 × 0.8 ≈ 4.46 cores, rounded up to 4 cores.

[0064] The calculation logic for memory quota is similar to that of CPU, with the formula: Memory Quota (GB) = Total Allocable Memory × Normalized Weight × Cache Demand Coefficient. The cache demand coefficient is positively correlated with the average file size—1.5 for average file size > 1GB, 1.2 for average file size between 100MB and 1GB, and 1.0 for average file size < 100MB. For the Gold-level main channel service, with an average file size of 5GB and a cache demand coefficient of 1.5, the memory quota is approximately 409.6 × 20.57% × 1.5 ≈ 126.1GB, rounded down to 126GB. For the Platinum-level main channel, with an average file size of 50MB, the quota is approximately 409.6 × 13.69% × 1.0 ≈ 56.1GB, rounded down to 56GB. For the Bronze-level channel, with an average file size of 20MB, the quota is approximately 409.6 × 6.97% × 1.0 ≈ 28.5GB, rounded down to 29GB.

[0065] Disk I / O quotas are directly linked to the estimated allocated bandwidth of the channel. Since I / O bandwidth determines the speed at which files are read from the disk and sent to the network, it needs to be matched with the network bandwidth to avoid bottlenecks. The formula is: I / O quota (MB / s) = estimated allocated bandwidth (Mbps) × 1.25 × I / O guarantee coefficient, where 1.25 is the coefficient for converting Mbps to MB / s (1Mbps = 125KB / s = 0.125MB / s), and the I / O guarantee coefficient is 1.1 (reserving 10% I / O redundancy). The estimated bandwidth of the Gold-level main channel is 20.57Mbps, and the I / O quota is approximately 20.57 × 1.25 × 1.1 ≈ 20.57 × 1.375 ≈ 28.28 MB / s, rounded up to 28 MB / s. The estimated bandwidth of the Platinum-level main channel is 13.69Mbps, and the quota is approximately 13.69 × 1.25 × 1.1 ≈ 18.82 MB / s, rounded up to 19 MB / s. The estimated bandwidth of the Bronze-level channel is 6.97Mbps, and the quota is approximately 6.97 × 1.25 × 1.1 ≈ 9.69 MB / s, rounded up to 10 MB / s.

[0066] The computing resource quota allocation table integrates quota information for three types of resources, including fields such as channel ID, service level, normalized weight, CPU quota (cores), memory quota (GB), I / O quota (MB / s), and resource allocation basis (such as complexity coefficient, cache coefficient). For example, the table entry description is: "Channel ID: VT-GD-NodeA-01, Service Level: Gold, Normalized Weight: 20.57%, CPU Quota: 20 cores (basis: 80×20.57%×1.2), Memory Quota: 126GB (basis: 409.6×20.57%×1.5), I / O Quota: 28MB / s (basis: 20.57×1.25×1.1), Resource Status: Reserved." After the allocation table is generated, the resource management system sends resource reservation instructions to each distributed node, binding a specified number of CPU, memory, and I / O resources to the corresponding channel instance. The reserved resources can only be used by that channel and cannot be occupied by other tasks, ensuring resource stability during channel transmission. For the backup channel, the resource reservation adopts the "hot standby" mode - 50% of CPU and memory are reserved, and 80% of I / O is reserved, which reduces resource waste and can quickly switch to full-load operation when the main channel fails.

[0067] S204 monitors the throughput performance and node load of each virtual transmission channel in real time during data transmission, and dynamically adjusts the cross-channel traffic load distribution strategy based on channel health to maintain optimal global transmission efficiency. Specifically, a performance monitoring agent can be deployed on each virtual transport channel instance to collect throughput, latency, packet loss rate, and node CPU and memory usage data in real time, and generate a channel performance metric stream. The performance monitoring agent adopts a collaborative deployment architecture of "end-edge-cloud". Agents must be deployed on both ends (source node and target node) and intermediate routing nodes of each virtual transmission channel instance to ensure data collection covers the entire channel link. The agent is a stateless, lightweight process, with resource consumption controlled to ≤1% CPU utilization and ≤50MB memory per node, avoiding interference with transmission performance. Deployment uses containerized rapid distribution. The agent container image is pushed to the target node through the resource management system. When the image starts, it is automatically associated with its channel ID, achieving a one-to-one binding between "channel and agent". For example, in the virtual channel VT-GD-NodeA-01 (Gold-level main channel), the source node Node-A, target node Node-D, and intermediate node Node-C all deploy agent instances associated with that channel ID.

[0068] The core metrics collected must be strongly correlated with the channel service level to ensure their relevance: Throughput refers to the amount of effective data transmitted by the channel per unit time, measured in Mbps (1Mbps = 125KB / s), reflecting the channel's actual data carrying capacity. Platinum / Gold level channels require a collection accuracy of 0.1Mbps, while Silver / Bronze level requires 0.5Mbps. Latency is measured using Round-Trip Time (RTT), the total time it takes for data to travel from the source node to the target node and receive confirmation, measured in milliseconds (ms). Gold level channels, due to their large file loads, require focused monitoring with a collection accuracy of 1ms, while other levels require 5ms. Packet loss rate is the percentage of lost data packets out of the total transmitted data packets, measured in %. Platinum level channels have stringent packet loss rate requirements, with a collection accuracy of 0.01%, while others require 0.1%. Node CPU and memory usage reflect the hardware resource status of the channel. CPU usage is the percentage of CPU cores currently used by the process, and memory usage is the percentage of physical memory used compared to the total memory of the node. Both have a collection accuracy of 0.1%. Only resource usage data of the associated nodes are collected for each channel to avoid cross-channel interference.

[0069] The sampling frequency adopts a "differentiated dynamic adjustment" strategy, automatically adapting based on the current transmission status of the channel: when the channel is in an active transmission state (files are being transmitted), the sampling frequency for throughput, latency, and packet loss rate is increased to 100ms / time (Gold / Platinum level) or 500ms / time (Silver / Bronze level), while CPU and memory utilization are sampled 1s / time; when the channel is idle (no transmission tasks), the sampling frequency for all indicators is reduced to 5s / time to reduce resource consumption. For example, when the VT-GD-NodeA-01 channel is transmitting 5GB of engineering drawings, throughput is sampled every 100ms, and CPU utilization is sampled every 1s; after the transmission is completed, all indicators are sampled every 5s.

[0070] The collected data needs to be preprocessed to generate a standardized indicator stream. Preprocessing includes data cleaning and format standardization: Data cleaning uses the "3σ criterion" to remove outliers. For example, if the delay sampling value of a certain channel is 1000ms, which is much higher than three times the standard deviation of the mean of 50ms (assuming the standard deviation is 15ms, 3σ=45ms, 1000ms>50+45=95ms), it is judged as an outlier and removed, and replaced with the moving average of the first three samples; Format standardization uses JSON format for encapsulation. Each indicator record contains six major fields: "channel ID, indicator type, value, unit, collection timestamp (accurate to milliseconds), and node identifier", to ensure data traceability.

[0071] The channel performance metric stream is a pre-processed data sequence organized in timestamp order and pushed to the performance analysis node in real time using a streaming transmission method. Transmission uses the UDP protocol (ensuring low latency), and CRC-32 checksums are added to key metrics (such as latency for the Gold-level channel) to ensure data transmission integrity. The receiving end of the metric stream uses a buffer to temporarily store the data. The buffer size is allocated according to the channel level: 10MB for Gold / Platinum-level channels (capable of storing 100,000 metrics at 100ms intervals), and 5MB for Silver / Bronze-level channels, to avoid data overflow.

[0072] Based on the channel performance metrics stream, calculate the health score for each virtual transmission channel. The health score comprehensively considers throughput performance stability and node load balancing to generate a channel health assessment report. The health score is based on a 100-point scale (0-100 points), with a higher score indicating a better channel condition. The calculation dimensions are "throughput performance stability" (weight 0.6) and "node load balancing" (weight 0.4). The weight allocation is based on the core requirement of the channel service—throughput performance directly determines transmission efficiency, hence it is given a higher weight. Each dimension first calculates a sub-score (0-100 points), and then the sub-scores are weighted and summed to obtain the final health score. The formula is: Health Score = Throughput Performance Stability Score × 0.6 + Node Load Balancing Score × 0.4.

[0073] The throughput performance stability score is calculated from two sub-items: "Throughput Target Compliance Rate" and "Indicator Fluctuation Level," with weights of 0.7 and 0.3, respectively. Throughput Target Compliance Rate = (Actual Average Throughput / Target Throughput) × 100. The target throughput is the estimated allocated bandwidth in the initial bandwidth weight allocation table. For example, the target throughput of the VT-GD-NodeA-01 channel is 20.57 Mbps, and the actual average throughput during a certain period is 18.6 Mbps. The compliance rate = (18.6 / 20.57) × 100 ≈ 90.4 points. Indicator fluctuation level is measured using the coefficient of variation (CV), where CV = (Indicator Standard Deviation / Indicator Mean Deviation). The fluctuation score is calculated as follows: (value) × 100, fluctuation score = 100 - coefficient of variation × 2 (coefficient 2 is used to map the fluctuation range to 0-100 points). For example, if the throughput sampling values ​​of this channel are 18.6, 18.8, 18.5, 18.9, and 18.7 Mbps, the mean is 18.7 Mbps, the standard deviation is approximately 0.1414, the coefficient of variation is (0.1414 / 18.7) × 100 ≈ 0.756, and the fluctuation score is 100 - 0.756 × 2 ≈ 98.49 points. Therefore, the throughput performance stability score is 90.4 × 0.7 + 98.49 × 0.3 ≈ 63.28 + 29.55 ≈ 92.83 points.

[0074] The node load balancing score is calculated for all nodes associated with the channel (source, destination, and intermediate routing nodes), with "CPU load balancing score" and "memory load balancing score" each weighted at 0.5. Taking CPU load as an example, the average CPU utilization (μ) of all associated nodes is first calculated, and then the absolute deviation of each node's utilization from the average (|x_i-μ|) is calculated. Load deviation = (Σ|x_i-μ| / n) / μ×100 (where n is the number of nodes), and load balancing score = 100 - load deviation × 1.5 (coefficient 1.5 strengthens the effect of load balancing). For example, VT-GD-NodeA-01 is associated with three nodes: Node-A, Node-C, and Node-D. Their CPU utilization rates are 35%, 32%, and 33%, respectively, with a mean μ = 33.33%. The sum of absolute deviations is 1.67 + 1.33 + 0.33 ≈ 3.33, and the load deviation is (3.33 / 3) / 33.33 × 100 ≈ 3.33%. The CPU load balancing score is 100 - 3.33 × 1.5 ≈ 95.01. Their memory utilization rates are 45%, 42%, and 43%, respectively. Similarly, the memory load balancing score is ≈ 95.2. Therefore, the node load balancing score is 95.01 × 0.5 + 95.2 × 0.5 ≈ 95.11.

[0075] The final health score = 92.83 × 0.6 + 95.11 × 0.4 ≈ 55.7 + 38.04 ≈ 93.74 points. The score is divided into health levels: 90-100 points are "Excellent", 80-89 points are "Good", 70-79 points are "Average", 60-69 points are "Poor", and <60 points are "Faulty". The channel health assessment report is generated per channel and includes the channel ID, service level, sub-scores for each dimension, final health score, health level, and core issue alerts (e.g., indicating whether a low score indicates insufficient throughput or load imbalance). For example, a gold-level channel with a score of 72 points would report: "Health Level: Average, Core Issue: Throughput compliance rate 75%, lower than target value; Node-C node CPU utilization reaches 65%, load is high". The assessment report is updated every 10 seconds and synchronized to the SDN controller and resource management system to support real-time decision-making.

[0076] Based on the channel health assessment report, identify virtual transmission channels whose health is below the preset health threshold and trigger a dynamic load adjustment mechanism to generate a load adjustment trigger signal; The preset health thresholds are strongly correlated with the channel service level, following the principle of "higher thresholds for higher-level channels" to ensure that the transmission quality of core channels is prioritized: Platinum-level channels have a threshold of 80 points (the lower limit of the "good" health level), Gold-level 75 points, Silver-level 70 points, and Bronze-level 60 points. The threshold settings are based on the service commitments of each channel level. For example, Platinum-level channels promise low latency and zero packet loss, and adjustments should be made promptly when the status approaches "average." Bronze-level channels have lower performance requirements and adjustments are only initiated when approaching "fault." Thresholds can be dynamically optimized using historical data. If a channel of a certain level frequently experiences transmission anomalies due to excessively low thresholds, the threshold can be appropriately increased. For example, if a Gold-level channel repeatedly experiences excessive packet loss at a score of 70, the threshold can be raised to 78 points.

[0077] The health assessment process is executed periodically by the performance analysis node (consistent with the update frequency of the assessment report, once every 10 seconds). It iterates through the health assessment reports of all channels, compares the final score in the report with the corresponding threshold, and marks channels with "score < threshold" as "channels to be adjusted". At the same time, it records the difference between the score and the threshold (Δ = threshold - score) for determining the degree of adjustment in the future. For example, after iteration, three channels to be adjusted are found: Gold level VT-GD-NodeA-01 (score 72, threshold 75, Δ=3), Silver level VT-SL-NodeA-02 (score 68, threshold 70, Δ=2), and Bronze level VT-BZ-NodeA-01 (score 58, threshold 60, Δ=2). The channels are sorted by service level, and the Gold level channel is processed first.

[0078] The dynamic load balancing mechanism employs a "tiered response" strategy, classifying adjustment levels into three categories based on Δ values: Slight (Δ≤5), Moderate (5<Δ≤10), and Severe (Δ>10). Different levels correspond to different adjustment intensities and response speeds. Slight adjustment level: response time ≤5 seconds, migrating only 20%-30% of the traffic to be adjusted; Moderate adjustment level: response time ≤3 seconds, migrating 30%-50% of the traffic while checking the status of the backup channel; Severe adjustment level: response time ≤1 second, migrating more than 50% of the traffic, immediately activating the backup channel to handle the load, and triggering node resource expansion if the backup channel is unavailable. For example, VT-GD-NodeA-01 with Δ=3 is considered a slight adjustment; a Bronze-level channel with Δ=2, although small in value, is still treated as a slight adjustment; if a Platinum-level channel scores 65 (threshold 80, Δ=15), it is considered a severe adjustment, and the backup channel is immediately activated.

[0079] Load adjustment trigger signals are structured data containing adjustment instructions in JSON format. Core fields include "trigger timestamp, channel ID to be adjusted, service level, health score, threshold, Δ value, adjustment level, response requirements, and proportion of traffic to be migrated." Trigger signals are reliably transmitted to the load scheduling node via TCP protocol to ensure no instructions are lost, and a digital signature is attached to prevent tampering. For multiple trigger signals, the load scheduling node processes them according to service level priority: Platinum > Gold > Silver > Bronze, ensuring that adjustments to core channels are executed first.

[0080] Based on the load adjustment trigger signal, a traffic redistribution algorithm is used to migrate the traffic of virtual transmission channels below the preset health threshold to virtual transmission channels above the preset health threshold. At the same time, the bandwidth weight of each channel is updated to generate a dynamically adjusted traffic load distribution strategy.

[0081] The traffic redistribution algorithm employs a "source-target channel matching + proportional traffic migration" strategy. The core steps are target channel selection, migration traffic calculation, and traffic migration execution. Target channel selection must meet two conditions: first, its health score must be at least 10 points higher than its own service level threshold (ensuring sufficient redundancy to handle additional traffic); second, its service level must be the same as or higher than the channel to be adjusted (avoiding performance degradation caused by migrating high-priority traffic to low-priority channels). For example, if the channel to be adjusted is the Gold-level VT-GD-NodeA-01, the following backup channels of the same level are selected: Gold-level backup channel VT-GD-NodeA-02 (health score 88, threshold 75, satisfying 88-75≥10) and Platinum-level backup channel VT-PT-NodeA-02 (health score 90, threshold 80, satisfying 90-80≥10) as target channels. Priority is given to channels of the same level to handle traffic, reducing resource waste.

[0082] Migration traffic calculation is divided into two scenarios: single-target and multi-target. For single-target scenarios, migration traffic = current actual traffic of the channel to be adjusted × migration ratio (specified in the trigger signal). For example, if the current actual traffic of VT-GD-NodeA-01 is 18.6Mbps and the migration ratio is 30%, then the migration traffic = 18.6 × 30% ≈ 5.58Mbps, all of which will be migrated to VT-GD-NodeA-02. For multi-target scenarios, migration traffic is allocated according to the redundancy bandwidth ratio of the target channel. Redundancy bandwidth = (target channel health score / 100) × target channel estimated bandwidth - target channel current actual traffic. For example, if VT-GD-NodeA-01's current actual traffic is 18.6Mbps and the migration ratio is 30%, then the migration traffic = 18.6 × 30% ≈ 5.58Mbps, all of which will be migrated to VT-GD-NodeA-02. Redundant bandwidth of 02 = (88 / 100) × 20.57 - 12 ≈ 18.10 - 12 = 6.10 Mbps, redundant bandwidth of VT-PT-NodeA-02 = (90 / 100) × 13.69 - 8 ≈ 12.32 - 8 = 4.32 Mbps, total redundant bandwidth = 6.10 + 4.32 = 10.42 Mbps, traffic migrated to VT-GD-NodeA-02 = 5.58 × (6.10 / 10.42) ≈ 3.29 Mbps, traffic migrated to VT-PT-NodeA-02 = 5.58 × (4.32 / 10.42) ≈ 2.29 Mbps.

[0083] Traffic migration is implemented by dynamically adjusting channel flow tables through the SDN controller. The controller sends a "traffic splitting instruction" to the source node of the channel to be adjusted, forwarding a specified proportion of traffic to the target channel along the new routing path. Simultaneously, it sends a "traffic reception instruction" to the nodes of the target channel to update the flow table to receive the new traffic. The migration process employs a "smooth switchover" mechanism, gradually reducing the traffic share of the channel to be adjusted within 1 second while simultaneously increasing the traffic share of the target channel to avoid packet loss caused by sudden traffic changes. For example, when migrating 5.58Mbps traffic, 1.12Mbps is split from 0-200ms, another 1.12Mbps is split from 200-400ms, until the entire migration is completed within 1 second. The packet loss rate is monitored in real time during the migration process; if the packet loss rate is greater than 0.5%, the migration is paused and the splitting speed is adjusted.

[0084] The bandwidth weight update needs to be synchronized with the traffic migration ratio. The new weight of the channel to be adjusted = the original weight × (1 - migration ratio), and the new weight of the target channel = the original weight + the original weight of the channel to be adjusted × migration ratio × (the proportion of redundant bandwidth of the target channel). For example, VT-GD-NodeA-01 originally had a weight of 0.2057 and a migration rate of 30%, so its new weight is 0.2057 × (1 - 0.3) ≈ 0.144; VT-GD-NodeA-02 originally had a weight of 0.2057 and carried 3.29Mbps of traffic (accounting for 59% of the total migrated traffic), so its new weight is 0.2057 + 0.2057 × 0.3 × 0.59 ≈ 0.2057 + 0.0367 ≈ 0.2424; VT-PT-NodeA-02 originally had a weight of 0.1369 and carried 2.29Mbps of traffic (accounting for 41%), so its new weight is 0.1369 + 0.2057 × 0.3 × 0.41 ≈ 0.1369 + 0.0252 ≈ 0.1621. After the weights are updated, the normalized weights need to be recalculated to ensure that the sum of the weights of all channels is 1. For example, if the total weight is still 4.5455 after the update, the normalized weight of VT-GD-NodeA-01 is 0.144 / 4.5455≈3.17%, and the corresponding new estimated bandwidth is 100×3.17%≈3.17Mbps.

[0085] The dynamically adjusted traffic load balancing strategy is a structured integration of the above adjustment results, comprising six modules: "Adjustment Time, Details of Channels to be Adjusted, Details of Target Channels, Traffic Migration List, New Bandwidth Weight Allocation Table, and Execution Status." The new bandwidth weight allocation table clearly defines the new weights, normalized weights, and estimated bandwidths for each channel. The execution status is marked as "Adjusting," "Completed," or "Failed." For example, the strategy records: "VT-GD-NodeA-01 migrates 5.58Mbps traffic to VT-GD-NodeA-02 (3.29Mbps) and VT-PT-NodeA-02 (2.29Mbps), with new weights of 0.144, 0.2424, and 0.1621 respectively. Execution status: Completed. Packet loss rate after migration is 0.1%, meeting requirements." This strategy is synchronized to all relevant nodes and modules, serving as a new basis for subsequent transmission scheduling.

[0086] S205, based on a dynamically adjusted traffic load distribution strategy, coordinates the multi-path parallel and sequential synchronous transmission of archive data in the virtual transmission channel cluster, realizing the synchronous scheduling of archive files among distributed nodes.

[0087] Specifically, it can parse the dynamically adjusted traffic load allocation strategy, obtain the bandwidth weight and traffic allocation ratio of each virtual transmission channel, and generate a channel traffic scheduling instruction set; Dynamically adjusted traffic load balancing strategies are typically stored in XML or JSON format, containing core modules such as adjustment time, channel list, new bandwidth weight, traffic allocation ratio, and execution status. The parsing process is handled by a scheduling strategy parser, which possesses multi-format compatibility and fault-tolerant verification capabilities—supporting simultaneous parsing of XML and JSON formats. Before parsing, it verifies the integrity of the strategy file using an MD5 checksum; if verification fails, it automatically invokes a historical backup strategy to prevent scheduling interruptions due to parsing errors. The parser calculates the file hash value and compares it with this value; if they match, parsing begins.

[0088] The core objective of the parsing is to extract the "effective bandwidth weight" and "actual traffic allocation ratio" of each virtual transmission channel. These two parameters directly determine the upper limit of the channel's transmission capacity. The effective bandwidth weight is a dynamically adjusted normalized weight that reflects the channel's resource share in the cluster. The actual traffic allocation ratio is the specific bandwidth share corresponding to the weight, which is multiplied by the global available bandwidth to obtain the channel's real-time available bandwidth. During parsing, the adjusted data must be matched with the set of virtual transmission channel instances according to the "channel ID association" principle to ensure that the parameters correspond one-to-one with the channels. For example, after parsing the strategy, the effective bandwidth weight of the Platinum-level main channel VT-PT-NodeA-01 is 0.1621, and the traffic allocation ratio is 16.21%; the effective bandwidth weight of the Gold-level main channel VT-GD-NodeA-01 is 0.144, and the traffic allocation ratio is 14.4%. If the global available bandwidth is 100Mbps, then the real-time available bandwidth of VT-PT-NodeA-01 = 100 × 16.21% ≈ 16.21Mbps, and VT-GD-NodeA-01 ≈ 14.4Mbps.

[0089] To ensure parameter accuracy, a "logical check" is required during the parsing process: First, the sum of the traffic allocation ratios for all channels must be 100% (allowing for a calculation error of ±0.1%). If the sum of the ratios after parsing for a certain strategy is 99.8%, it is padded to 100% according to the weight ratio of each channel. Second, the effective bandwidth weight of a channel must not exceed the maximum weight threshold of its service level (e.g., the maximum weight threshold for a Platinum channel is 30%), to avoid excessive resource consumption by a single channel. For example, if the weight of a Platinum channel after parsing is 32%, exceeding the threshold, the parser automatically allocates the excess portion proportionally to other channels of the same level.

[0090] The channel traffic scheduling instruction set is a structured presentation of the parsed results, adhering to the principle of "single channel, single instruction." Each instruction contains eight fields: instruction ID, channel ID, service level, effective bandwidth weight, traffic allocation ratio, real-time available bandwidth (Mbps), execution priority, and effective time. The execution priority is consistent with the channel service level (Platinum > Gold > Silver > Bronze), and the effective time is 1 second after instruction generation to ensure timely adjustments. The instruction set is encapsulated in JSON format for easy parsing by the executor. After the instruction set is generated, it is pushed to the executor nodes of each channel via an encrypted TCP channel, while a copy is retained in the scheduling log for traceability.

[0091] According to the channel traffic scheduling instruction set, the archive data files to be transmitted are allocated to the corresponding virtual transmission channels according to the transmission priority category, and a transmission task descriptor is generated for each file, thus generating a file transmission task queue. File and channel allocation follows the "priority strong matching + redundancy" rule. The core logic is that "high-priority files are only allocated to channels of the same or higher service level, low-priority files can be allocated to channels of the same or lower level, and high-level channels can temporarily accept low-priority files when there is remaining bandwidth (marked as 'temporarily occupied')." Before allocation, the "remaining capacity" of each channel must be calculated. The formula is: Remaining capacity = Real-time available bandwidth - Estimated total bandwidth of allocated files. For example, the real-time available bandwidth of the VT-PT-NodeA-01 channel is 16.21Mbps, and 3 files have been allocated, occupying a total of 10Mbps, leaving a remaining capacity of 6.21Mbps.

[0092] The allocation process follows a descending file priority order, prioritizing high-priority files to ensure core data transmission is prioritized. For example, the files to be transmitted include: "Approval Document-001.docx" (estimated bandwidth 2Mbps), "Engineering Drawings-001.dwg" (estimated bandwidth 8Mbps), and "Labor Contract-001.docx" (estimated bandwidth 1.5Mbps). The traversal order is: High-priority high-priority → High-priority low-priority → Medium-priority. "Engineering Drawings-001.dwg" is allocated to the Gold-level main channel VT-GD-NodeA-01 (9Mbps remaining capacity, meeting the 8Mbps requirement); "Approval Document-001.docx" is allocated to the Platinum-level main channel VT-PT-NodeA-01 (6.21Mbps remaining, meeting the 2Mbps requirement); and "Labor Contract-001.docx" is allocated to the Silver-level main channel VT-SL-NodeA-01 (5Mbps remaining capacity, meeting the 1.5Mbps requirement). If the channel corresponding to a high-priority file has no remaining capacity (e.g., another high-priority low-power file "Approval Document-002.docx" has an estimated bandwidth of 5Mbps, but VT-PT-NodeA-01 has only 4.21Mbps remaining), it will be allocated to the same level of backup channel VT-PT-NodeA-02 (with a remaining capacity of 7Mbps) to ensure that high-priority files are not downgraded in allocation.

[0093] The transfer task descriptor is a standardized carrier of file transfer requirements. It must contain core fields such as "task ID, file_id, file name, transfer priority category, channel ID, estimated bandwidth (Mbps), source storage path, target storage path, data dependency list, transfer timeout threshold (seconds), and integrity verification method". The data dependency list is directly extracted from the dependency_graph of the structured feature description. The timeout threshold is set according to priority (Platinum level 10 seconds, Gold level 30 seconds, Silver level 60 seconds, Bronze level 120 seconds). Integrity verification uses SHA-256 hash verification.

[0094] The file transfer task queue is split by "channel dimension," with each virtual transfer channel corresponding to an independent queue. Tasks within the queue are ordered according to the rule of "descending transfer priority + ascending estimated transfer time"—files with the same priority are executed first, reducing overall queue time. For example, the VT-PT-NodeA-01 queue contains "Approval Document-001.docx" (transfer time 2 seconds) and "Approval Document-003.docx" (transfer time 3 seconds), with the former ordered first. If temporarily assigned medium-priority files (marked "temporarily occupied") are mixed into the queue, they are placed after high-priority files. The queue uses a "first-in, first-out + priority preemption" mechanism. If a new high-priority task is added, it can be inserted at the head of the queue for priority execution. For example, if a medium-priority task already exists in the queue, and a new platinum-level high-priority task is added, the current medium-priority task is immediately paused, the high-priority task is executed, and the medium-priority task resumes after execution.

[0095] Coordinate the executors of each virtual transmission channel to start multiple parallel transmissions in the order of the file transmission task queue, while monitoring the transmission progress and anomalies, and generating the parallel transmission execution status. Each virtual transmission channel's executor is a dedicated process deployed on a node, responsible for receiving task queues, initiating data transmission, and interacting with the scheduling node. The binding relationship between the executor and the channel is fixed through the channel ID; for example, the executor of VT-PT-NodeA-01 only handles the task queue of that channel. The coordination of the executors is handled by the global scheduling node. The scheduling node dynamically adjusts the task initiation timing of each executor through a "real-time channel bandwidth utilization feedback" mechanism to avoid bandwidth contention during multi-channel parallel transmission. When the total bandwidth utilization of a channel at a certain service level exceeds 90% of the reserved bandwidth for that level, the initiation of new tasks at that level is temporarily suspended, prioritizing the transmission quality of already initiated tasks. For example, if the total reserved bandwidth for a gold-level channel is 40Mbps, and the current total utilization reaches 36Mbps (90%), the initiation of new gold-level tasks is suspended until the utilization drops below 80%.

[0096] The "parallelism control" of multi-channel parallel transmission is positively correlated with the channel bandwidth weight. Parallelism = ceil(effective bandwidth weight × 10), meaning that a channel with a higher weight can execute more tasks simultaneously. For example, VT-PT-NodeA-01 has an effective bandwidth weight of 0.1621, and its parallelism = ceil(0.1621 × 10) = 2, allowing it to execute 2 tasks simultaneously; VT-GD-NodeA-01 has a weight of 0.144, and its parallelism = 2; VT-BZ-NodeA-01 has a weight of 0.0697, and its parallelism = 1. Parallel transmission adopts a "fragmented transmission + pipeline" mode, dividing large files (>1GB) into 128MB fragments. After the previous fragment is transmitted to the target node and received, the transmission of the next fragment begins immediately, while the source node continuously reads the next fragment data, reducing I / O wait time. For example, a 5GB file named "Engineering Drawing-001.dwg" is divided into 40 128MB segments. After the first segment is transmitted to the target node, the transmission of the second segment begins immediately without waiting for all segments to be read.

[0097] The transmission progress monitoring adopts a "fragment-level progress accumulation + real-time feedback" mechanism. The progress of each task = (number of successfully transmitted fragments / total number of fragments) × 100%. Small files (≤1GB) are calculated as a whole (100% upon completion of transmission). Each time the executor completes a fragment transmission or a small file transmission, it immediately reports the progress to the scheduling node at a frequency of once per second. The progress data includes "task ID, current progress (%), amount of data transmitted (MB), and remaining time (seconds)".

[0098] Transmission anomaly monitoring focuses on three core anomalies: "delay timeout, excessive packet loss, and verification failure." Anomaly judgment thresholds are tied to the channel service level: Delay timeout refers to a single fragment transmission's RTT exceeding twice the timeout threshold (e.g., for Platinum-level tasks, the timeout threshold is 10 seconds; if the RTT > 20 seconds, it's considered an anomaly); Excessive packet loss refers to a task's packet loss rate exceeding the channel's maximum allowed packet loss rate for three consecutive samples (e.g., Gold-level allows 0.1%; three consecutive > 0.1% results in an anomaly); Verification failure refers to the target node's calculated SHA-256 value not matching the verification value in the task descriptor after receiving the fragment. When an anomaly occurs, the executor immediately triggers a "tiered processing mechanism": Minor anomalies (e.g., a single packet loss rate of 0.15%) automatically initiate retransmission (retransmission count ≤ 3 times); Moderate anomalies (e.g., two consecutive verification failures) suspend the current task and send an alarm to the scheduling node; Severe anomalies (e.g., delay timeout and retransmission failure) terminate the task, mark it as "transmission failed," and trigger manual intervention. For example, if the verification of the 5th fragment of "Engineering Drawing-001.dwg" fails, the executor will start retransmission. If it succeeds after 2 retransmissions, the subsequent fragment transmission will continue. If it fails to retransmit 3 times, the task will be paused and an alarm will be issued.

[0099] The parallel transmission execution status is a summary of all task progress and exception information, presented in groups by channel, including "channel ID, number of currently running tasks, total number of tasks, overall progress (%), and list of exception tasks". The status data is updated every 2 seconds and synchronized to the visual monitoring interface of the scheduling node, allowing maintenance personnel to monitor the transmission status in real time.

[0100] Based on the parallel transmission execution state, for files with data dependencies, sequential synchronization control points are inserted into the transmission task queue to ensure that dependent files are transmitted before the target file, ultimately achieving complete synchronous scheduling of archive files.

[0101] Data dependency identification is based on the "dependency_graph" field in the structured feature description. This field is parsed to obtain the "dependency-dependent" pairs for each file, forming a dependency list. For example, "Project Report.docx" (file_id: Node-A-78901234) depends on "Progress Chart.xlsx" (file_id: Node-A-78904321), and "Progress Chart.xlsx" depends on "Original Data.csv" (file_id: Node-A-78909876), forming a chain dependency relationship of "Project Report → Progress Chart → Original Data". During the identification process, "circular dependencies" (such as A depending on B, and B depending on A) must be excluded. If a circular dependency is detected, it is immediately marked as a "dependency anomaly," triggering manual review of the dependencies before transmission.

[0102] Sequence synchronization control points are logical instructions used to enforce the transmission order, and are divided into two categories: "waiting control points" and "checking control points." Waiting control points are used to "pause the transmission of the target file when the dependent file has not been completed." They are inserted before the target file task, and the instruction content is "wait for the file with file_id XXX to be transmitted and for verification to pass." Checking control points are used to "verify the integrity of the dependent file after its transmission is complete, and only start the transmission of the target file after confirming that it is correct." They are inserted after the dependent file task and before the target file task, and the instruction content is "verify the SHA-256 value of the file with file_id XXX, and confirm that it matches the descriptor." For example, a waiting control point is inserted before the "Project Report.docx" task (waiting for "Progress Chart.xlsx" to complete), and a checking control point is inserted after the "Progress Chart.xlsx" task (verifying its integrity) to ensure that the "Progress Chart" is complete before starting the transmission of "Project Report."

[0103] The timing of control point insertion is dynamically determined by the scheduling node based on the parallel transmission execution status: when the transmission progress of the dependent file reaches 100% and there are no anomalies, a verification control point is inserted; after the verification control point executes successfully (integrity verification is successful), the waiting control point before the target file is removed, allowing the target file to start transmission; if the transmission of the dependent file is abnormal (such as verification failure), the target file task is marked as "waiting for an anomaly" until the dependent file problem is resolved. For example, after the transmission of "raw data.csv" is completed and verification is passed, the waiting control point before the task of "progress chart.xlsx" is removed, and the transmission of "progress chart" is started; after the verification of "progress chart" is passed, the waiting control point of "project report" is removed, and the transmission of "project report" is started.

[0104] For complex dependency relationships such as "one-to-many" or "many-to-one" (e.g., multiple files depend on the same core file, or one file depends on multiple files), "aggregate control" is adopted: when multiple files depend on the same core file, a unified wait control point is inserted before all dependent file tasks, and all dependent files are released synchronously after the core file is completed; when one file depends on multiple files, multiple wait control points are inserted before the target file task, and the target file is released only after all dependent files are completed and verified.

[0105] The final verification of complete synchronous scheduling employs a "dependency relationship closed-loop verification." After all file transfers are completed, the scheduling node traverses the dependency relationship list, checking whether each file's "dependent files have all been transferred" and "dependent files have been transferred after itself." Simultaneously, it verifies the integrity of each file (SHA-256 value matching) and the correctness of the storage path (consistent with the target path of the task descriptor). Upon successful verification, a "Synchronization Completion Report" is generated, containing information such as "total number of transferred files, number of successful transfers, number of failed transfers, dependency relationship verification results, and total transfer time." Failed verification is marked as a "failed item," such as "Project report.docx is completed, but the dependent progress chart.xlsx failed verification," facilitating targeted problem investigation. For example, if a batch of transfers contains 100 files, with 98 successful and 2 failing due to dependency anomalies, the synchronization completion report explicitly records the failed item as "file_id:Node-A-11112222 has a circular dependency and was not transferred."

[0106] Another embodiment of the present invention provides a synchronous transmission and scheduling system for archival data files, see [link to relevant documentation]. Figure 3 The system may include: The parsing module 301 is used to perform multi-dimensional feature extraction and structural parsing on the archive data files to be synchronized, and generate a structured feature description that includes file business attributes, format sensitivity and data dependency relationships; The determination module 302 is used to dynamically determine the transmission priority category of each file and estimate its network transmission overhead based on the structured feature description, through a pre-trained classification model and a business rule engine. The allocation module 303 is used to self-organize and construct a virtual transmission channel cluster with differentiated service levels according to the transmission priority category and the network transmission overhead, and pre-allocate an initial bandwidth weight and computing resource quota for each channel; The adjustment module 304 is used to monitor the throughput performance and node load of each virtual transmission channel in real time during data transmission, and dynamically adjust the cross-channel traffic load distribution strategy according to the channel health to maintain optimal global transmission efficiency. The scheduling module 305 is used to coordinate the multi-path parallel and sequential synchronous transmission of archive data in the virtual transmission channel cluster based on the dynamically adjusted traffic load distribution strategy, so as to realize the synchronous scheduling of archive files among distributed nodes.

[0107] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0108] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0109] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. An archival data file synchronization transmission scheduling method, characterized by, The method comprises: Multi-dimensional feature extraction and structure analysis are performed on the archive data files to be synchronized to generate a structured feature description containing file business attributes, format sensitivity and data dependency relationship; Based on the structured feature description, the transmission priority category of each file is dynamically determined and its network transmission cost is estimated through a pre-trained classification model and a business rule engine; According to the transmission priority category and the network transmission cost, a virtual transmission channel cluster with differentiated service levels is self-organized, and each channel is pre-allocated an initial bandwidth weight and a computing resource quota; During data transmission, the throughput performance and node load of each virtual transmission channel are monitored in real time, and the traffic load distribution strategy across channels is dynamically adjusted according to the channel health degree to maintain the optimal global transmission efficiency; Based on the dynamically adjusted traffic load distribution strategy, multi-path parallel and sequential synchronization transmission of archive data is coordinated and executed in the virtual transmission channel cluster to realize synchronization scheduling of archive files among distributed nodes.

2. The method of claim 1, wherein, The multi-dimensional feature extraction and structure analysis of the archive data files to be synchronized to generate a structured feature description containing file business attributes, format sensitivity and data dependency relationship comprises: Scan the set of archive data files to be synchronized, extract the basic metadata of each file, including file size, creation time, modification time, file type and storage path, and generate a basic metadata set; Based on the basic metadata set, the file content abstract and keywords are analyzed by natural language processing technology, and the business domain attribution is identified in combination with the file type to generate a business attribute feature set; According to the business attribute feature set, the complexity and compression characteristics of the file format are analyzed, the format conversion and compatibility indicators are calculated, and the format sensitivity evaluation results are generated; In combination with the business attribute feature set and the format sensitivity evaluation results, the reference relationship and version dependency among files are queried through a graph database to construct a data dependency graph, and finally the structured feature description is integrated and generated.

3. The method of claim 2, wherein, The dynamic determination of the transmission priority category of each file and the estimation of its network transmission cost based on the structured feature description through a pre-trained classification model and a business rule engine comprises: Input the structured feature description into the pre-trained classification model, which is trained based on historical transmission logs, to output the preliminary priority score and transmission time prediction of the file, and generate a model prediction result; Input the model prediction result into the business rule engine, apply the preset business rules to correct the priority score, and generate the priority after rule correction; Based on the priority after rule correction, in combination with the file size and network historical performance data, the bandwidth estimation algorithm is used to calculate the transmission cost of each file to generate the network transmission cost estimation value; According to the network transmission cost estimation value and the priority after rule correction, the files are divided into different transmission priority categories using a clustering algorithm to generate a transmission priority category division result.

4. The method of claim 3, wherein, The self-organization of a virtual transmission channel cluster with differentiated service levels according to the transmission priority category and the network transmission cost, and the pre-allocation of an initial bandwidth weight and a computing resource quota to each channel comprise: According to the transmission priority class division result, the number and service levels of the required virtual transmission channels are determined, each service level corresponds to a priority class, and a virtual transmission channel configuration scheme is generated; Based on the virtual transmission channel configuration scheme, virtual channel instances are dynamically created on the distributed network nodes, and each instance is assigned a unique identifier and an initial routing path, generating a virtual transmission channel instance set; In combination with the network transmission overhead estimation value, the initial bandwidth weight of each virtual transmission channel instance is calculated, the weight is proportional to the priority of the served file and the transmission overhead, and an initial bandwidth weight allocation table is generated; According to the initial bandwidth weight allocation table, the calculation resources are reserved for each virtual transmission channel instance from the global resource pool, and a calculation resource quota allocation table is generated.

5. The method of claim 4, wherein, In the data transmission process, the throughput performance and node load of each virtual transmission channel are monitored in real time, and the traffic load distribution strategy across channels is dynamically adjusted according to the channel health degree to maintain the optimal global transmission efficiency, including: Deploy a performance monitoring agent on each virtual transmission channel instance to collect throughput, delay, packet loss rate, and node CPU and memory usage data in real time, generating a channel performance indicator stream; Based on the channel performance indicator stream, the health score of each virtual transmission channel is calculated, which considers the throughput performance stability and node load balancing, and a channel health evaluation report is generated; According to the channel health evaluation report, virtual transmission channels with a health score below a preset health threshold are identified, and a dynamic load adjustment mechanism is triggered, generating a load adjustment trigger signal; According to the load adjustment trigger signal, a traffic redistribution algorithm is used to migrate part of the traffic of the virtual transmission channel below the preset health threshold to the virtual transmission channel above the preset health threshold, and the bandwidth weight of each channel is updated, generating a dynamically adjusted traffic load distribution strategy.

6. The method of claim 5, wherein, Based on the dynamically adjusted traffic load distribution strategy, multi-path parallel and sequential synchronous transmission of archive data is coordinated and executed in the virtual transmission channel cluster, realizing synchronous scheduling of archive files between distributed nodes, including: Parse the dynamically adjusted traffic load distribution strategy to obtain the bandwidth weight and traffic allocation ratio of each virtual transmission channel, and generate a channel traffic scheduling instruction set; According to the channel traffic scheduling instruction set, the archive data files to be transmitted are distributed to the corresponding virtual transmission channels according to the transmission priority class, and a transmission task descriptor is generated for each file, generating a file transmission task queue; Coordinate the executors of each virtual transmission channel to start multi-path parallel transmission according to the order in the file transmission task queue, while monitoring the transmission progress and exceptions, generating a parallel transmission execution state; Based on the parallel transmission execution state, for files with data dependency, a sequential synchronization control point is inserted in the transmission task queue to ensure that dependent files are transmitted before the target file, and finally the complete synchronous scheduling of archive files is realized.

7. An archival data file synchronization transmission scheduling system, characterized by, The system comprises: An analysis module for multi-dimensional feature extraction and structure analysis of the archive data files to be synchronized, generating a structured feature description containing file business attributes, format sensitivity, and data dependency; A determining module is configured to determine a transmission priority category of each file and estimate a network transmission cost of the file based on the structured feature description, a pre-trained classification model, and a business rule engine; An allocating module is configured to self-organize a cluster of virtual transmission channels with differentiated service levels according to the transmission priority category and the network transmission cost, and pre-allocate an initial bandwidth weight and a computing resource quota to each channel; An adjusting module is configured to monitor a throughput performance and a node load of each virtual transmission channel in real time during data transmission, and dynamically adjust a traffic load distribution strategy across channels according to a channel health degree, so as to maintain an optimal global transmission efficiency; A scheduling module is configured to perform multi-path parallel and sequential synchronous transmission of archive data in the cluster of virtual transmission channels based on the dynamically adjusted traffic load distribution strategy, so as to realize synchronous scheduling of archive files among distributed nodes.

8. The system of claim 7, wherein, The analyzing module is specifically configured to: scan a set of to-be-synchronized archive data files, extract basic metadata of each file, including file size, creation time, modification time, file type, and storage path, and generate a basic metadata set; analyze the file content abstract and keywords based on the basic metadata set through a natural language processing technology, and identify a business domain attribute of the file according to the file type, to generate a business attribute feature set; analyze the complexity and compression characteristics of the file format according to the business attribute feature set, calculate a format conversion and compatibility index, and generate a format sensitivity evaluation result; combine the business attribute feature set and the format sensitivity evaluation result, query a reference relationship and a version dependency among files through a graph database, construct a data dependency graph, and finally integrate and generate a structured feature description.

9. A storage medium, characterized by The storage medium stores a computer program, and the computer program is configured to execute the method in any one of claims 1-6 when running.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method in any one of claims 1-6 by running the computer program.