Intelligent transmission system and method for large files based on supercomputing platform of SFTP nodes
Patent Information
- Application Number
- CN202610971192.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]因此,本发明提供了基于SFTP节点的超算平台大文件智能传输方法解决对不同数据块在历史传输中的稳定性差异缺乏刻画,难以避免高风险数据占用关键传输通道,进而影响整体传输性能,传统方法通常在传输开始前确定调度策略,在传输过程中仅进行有限调整,未能充分利用实时反馈数据对策略进行动态优化,导致系统对网络波动和节点状态变化的适应能力不足问题
[0016]本发明有益效果为:通过对文件内容进行结构解析并结合历史访问行为进行关联分析,将大文件划分为具有结构特征和语义特征的数据分块,并基于数据指纹实现历史行为的继承,使数据分块由单纯的物理划分单元转变为具备语义重要性的传输单元,从而提升分块合理性并降低不合理分块带来的传输效率损失及重传成本,通过将语义重要性与基于历史传输行为构建的风险参数以及SFTP节点的实时资源状态进行联合建模,实现面向数据特性与系统资源的多通道调度,使高重要性且低风险的数据优先占用优质通道资源,而高风险数据被分散或延后处理,从而提高资源利用效率并增强整体传输稳定性。
Smart Images

Figure CN122824731A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer network technology, and in particular to a large file intelligent transfer system and method for supercomputing platforms based on SFTP nodes. Background Technology
[0002] With the rapid development of high-performance computing and supercomputing platforms, applications such as large-scale scientific computing, meteorological simulation, bioinformatics analysis, and artificial intelligence training are increasingly reliant on massive data processing capabilities. Against this backdrop, efficient transmission of large files between different computing nodes has become a crucial foundation for supporting the operation of supercomputing tasks. Due to its excellent security and cross-platform compatibility, it is widely used for data transmission in supercomputing environments.
[0003] Existing scheduling mechanisms are mostly based on system resource utilization, and lack characterization of the stability differences of different data blocks in historical transmission. This makes it difficult to avoid high-risk data occupying critical transmission channels, thereby affecting overall transmission performance. Traditional methods usually determine the scheduling strategy before transmission begins and only make limited adjustments during transmission. They fail to make full use of real-time feedback data to dynamically optimize the strategy, resulting in insufficient adaptability of the system to network fluctuations and node state changes. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a method for intelligent transmission of large files on a supercomputing platform based on SFTP nodes. This method addresses the lack of characterization of the stability differences of different data blocks in historical transmission, the difficulty in avoiding high-risk data occupying critical transmission channels, and the resulting impact on overall transmission performance. Traditional methods typically determine the scheduling strategy before transmission begins and make only limited adjustments during transmission, failing to fully utilize real-time feedback data to dynamically optimize the strategy, resulting in insufficient adaptability of the system to network fluctuations and node status changes.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for intelligent transmission of large files on a supercomputing platform based on SFTP nodes, which includes acquiring file feature data, historical transmission data and SFTP node status data of the large file to be transmitted. Based on the file feature data, historical transmission data, and SFTP node status data, correlation analysis is performed to obtain transmission feature parameters, and an adaptive chunking strategy and chunk transmission priority are generated according to the transmission feature parameters. Based on the block transmission priority and SFTP node status, a multi-channel SFTP transmission strategy is established and dynamically adjusted to perform concurrent transmission of each data block. During the transmission process, transmission feedback data is acquired, and the block transmission priority and multi-channel SFTP transmission strategy are dynamically updated based on the transmission feedback data. If the transmission is interrupted or completed, the transmission is resumed and the data is reassembled based on the status of the transmitted segments and the verification results to obtain the complete file.
[0007] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the method for generating adaptive block partitioning strategies and block transfer priorities includes: The large file to be transmitted is scanned by a sliding window of a preset length. The data in each window is read, the byte distribution features and continuous repetition features are extracted, and the information entropy and repetition index are calculated respectively, thereby forming a data feature sequence that reflects the local structural changes of the data. Traverse the data feature sequence, calculate the difference in information entropy and the difference in repetition index between adjacent sliding windows in turn, compare the calculated difference with the preset mutation threshold, mark the position where the difference exceeds the mutation threshold as the mutation point of the data structure, use the mutation point as the segmentation boundary, and physically divide the large file to be transmitted to obtain multiple data structure partitions with independent data structure features. The starting position, length, and data distribution characteristics of each data structure partition are encoded, and a content summary identifier is generated for each partition as a unique data fingerprint identifier for the data structure partition. Data fingerprints are extracted from existing data blocks in the historical transmission data in the same way, and the data fingerprints are correlated with the corresponding historical access records. The access order, access frequency and access concentration in the access records are statistically analyzed to obtain the access behavior characteristics of the historical data blocks. The data fingerprint of the current data structure partition is matched with the data fingerprint of the historical data block. The data structure partition that matches successfully inherits the corresponding access behavior features, and the data structure partition that does not match is assigned default behavior parameters, thereby establishing semantic importance parameters for each data structure partition. The semantic importance parameter is bound to the data structure partition, and the data structure partition is used as the basic block unit. Each block is sorted according to the semantic importance parameter, thereby generating an adaptive block strategy and block transmission priority.
[0008] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the multi-channel SFTP transfer strategy includes: The data fingerprint identifier is used to search the historical transmission data, extract the transmission log corresponding to the data block that is the same as or similar to the current data fingerprint, and perform statistical analysis on the transmission duration, retransmission count and transmission interruption status in the transmission log to obtain the historical transmission behavior characteristics corresponding to the data fingerprint. The historical transmission behavior characteristics are normalized, and the transmission stability, retransmission frequency and interruption probability are comprehensively evaluated to establish structural risk parameters for each data fingerprint, which are used to characterize the instability of data during transmission. The structural risk parameters and semantic importance parameters are jointly analyzed, and each data block is comprehensively scheduled and evaluated. Data blocks with high semantic importance are given priority to enter the transmission queue first, while data blocks with low transmission priority are delayed or dispersed according to stability, thereby generating a block scheduling sequence. The SFTP transmission channel is allocated according to the block scheduling sequence. Data blocks with high transmission priority are allocated transmission resources first, and an initial multi-channel transmission strategy is established based on the node's current bandwidth and load. During the transmission process, the actual transmission behavior of each data block is continuously monitored, the actual transmission duration, retransmission count and abnormal situations of the corresponding data fingerprint are recorded, and the data are compared and analyzed with structural risk parameters to obtain risk deviation information. The structural risk parameters of the corresponding data fingerprint are dynamically corrected based on the risk deviation information, and the correction results are fed back to the subsequent data block scheduling and evaluation process, thereby realizing an adaptive transmission scheduling update mechanism based on data fingerprint.
[0009] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the acquisition of transfer feedback data includes: The transmission rate, transmission delay, and retransmission status of each data block are collected to obtain real-time transmission performance data. The connection status, abnormal interruptions, and bandwidth fluctuations during the transmission process are recorded to obtain transmission status feedback data.
[0010] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the dynamic updating of block transfer priority and multi-channel SFTP transfer strategy includes: Statistical analysis of the transmission feedback data is performed, and the actual transmission cost of each block is recalculated to obtain an updated priority sequence. The transmission channel parameters are adjusted based on changes in node resources to obtain an updated multi-channel transmission strategy.
[0011] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the method includes: resuming interrupted transfers and data reassembly. Record the status of the transmitted segments and mark the incomplete segments to obtain a set of resume tasks; The complete file is obtained by verifying and concatenating all blocks according to the block index order.
[0012] As a preferred embodiment of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes described in this invention, the updating of the transfer characteristic parameters includes: By fusing historical transmission data with current transmission feedback data, updated training data is obtained; The transmission feature parameters are recalculated based on the updated training data to achieve continuous optimization of the transmission strategy.
[0013] Secondly, the present invention provides a large file intelligent transmission system for a supercomputing platform based on an SFTP node, including a data acquisition module, a feature analysis module, a strategy generation module, a dynamic transmission module, and a reconstruction and recovery module; The data acquisition module is used to acquire file characteristic data, historical transmission data, and SFTP node status data of the large file to be transferred. The feature analysis module is used to perform correlation analysis on the file feature data, historical transmission data and SFTP node status data to obtain transmission feature parameters, and generate an adaptive block splitting strategy and block transmission priority based on the transmission feature parameters. The strategy generation module is used to establish and dynamically adjust the multi-channel SFTP transmission strategy based on the block transmission priority and SFTP node status, and to transmit each data block concurrently. The dynamic transmission module is used to acquire transmission feedback data during the transmission process and dynamically update the block transmission priority and multi-channel SFTP transmission strategy based on the transmission feedback data. The reconstruction and recovery module is used to resume transmission and reconstruct data based on the status of the transmitted blocks and the verification results after transmission is interrupted or completed, so as to obtain a complete file.
[0014] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in the first aspect of the present invention.
[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the intelligent large file transfer method for a supercomputing platform based on an SFTP node as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: By performing structural analysis on file content and combining it with historical access behavior for correlation analysis, large files are divided into data blocks with structural and semantic features. Based on data fingerprints, historical behavior is inherited, transforming data blocks from simple physical partitioning units into transmission units with semantic importance. This improves the rationality of block division and reduces the transmission efficiency loss and retransmission cost caused by unreasonable block division. By jointly modeling semantic importance with risk parameters constructed based on historical transmission behavior and the real-time resource status of SFTP nodes, multi-channel scheduling oriented towards data characteristics and system resources is achieved. High-importance and low-risk data are given priority to occupy high-quality channel resources, while high-risk data is distributed or processed later, thereby improving resource utilization efficiency and enhancing overall transmission stability. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a method for intelligent large file transfer on a supercomputing platform based on SFTP nodes. Figure 2 This is a schematic diagram of a large file intelligent transfer system for a supercomputing platform based on SFTP nodes. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Reference Figures 1-2 This is one embodiment of the present invention, which provides a method for intelligent large file transfer on a supercomputing platform based on an SFTP node, including the following steps: S1. Obtain file characteristic data, historical transfer data, and SFTP node status data of the large file to be transferred.
[0023] Furthermore, the system performs attribute parsing on large files to be transferred, extracting basic information such as file size, type, path, creation time, and modification time to describe the file's basic attributes.
[0024] The file content is scanned using a sliding window method. Data sequences are read in each window, and the byte distribution and continuous repetition are read to form the basic data sequence for subsequent structural calculations, which is used to characterize the data structure features inside the file.
[0025] By combining the scanning results, the data change trends in different regions within the document are characterized, making the document exhibit a continuously changing structural distribution.
[0026] Analyze historical transmission logs to identify transmission duration, transmission rate, failures and retransmissions in past transmission tasks, and record the location of interruptions and bandwidth usage during transmission.
[0027] For data blocks in historical transmissions, the corresponding data identification information is extracted and associated with the actual transmission results and access records to reflect the behavioral characteristics of the data during use.
[0028] The operating status of the SFTP node is collected in real time, including CPU usage, memory usage, disk read / write capability, and network bandwidth utilization. At the same time, the current number of connections and task queuing status are recorded to reflect the real-time load level of the node.
[0029] File attribute information, content structure characteristics, historical transmission information, and node operation status are uniformly organized to ensure that all types of data are consistent in format and scale.
[0030] It should be noted that by introducing a joint collection method of file attributes, content structure, historical behavior and node status before large file transmission, the data foundation before transmission is no longer limited to file size or file type, but can simultaneously reflect the internal structural characteristics of the file, historical transmission performance and current node load status, thus providing multi-dimensional judgment basis for subsequent adaptive block division and transmission scheduling.
[0031] It not only obtains basic file attribute information, but also scans the file content through a sliding window to analyze byte distribution, repetition features, and information entropy changes, thus characterizing the file's composition from the perspective of its internal data structure.
[0032] At the same time, historical transmission data and node operating status are introduced, so that the data acquisition process not only reflects the "file itself", but also the "historical behavior" and the "current environment".
[0033] By jointly organizing multi-source data, the input data simultaneously possesses structural, behavioral, and environmental characteristics, providing a unified foundation for subsequent data feature-based segmentation strategies and scheduling decisions, thereby avoiding decision-making biases caused by relying solely on single-dimensional information.
[0034] S2. Based on file feature data, historical transmission data and SFTP node status data, perform correlation analysis to obtain transmission feature parameters, and generate an adaptive chunking strategy and chunk transmission priority based on the transmission feature parameters.
[0035] Furthermore, a sliding window scan is performed on the large file to be transmitted according to a preset length. Data within each window is read, byte distribution features and continuous repetition features are extracted, and information entropy and repetition index are calculated respectively, thereby forming a data feature sequence that reflects the local structural changes of the data.
[0036] Traverse the data feature sequence, calculate the difference in information entropy and the difference in repetition index between adjacent sliding windows, compare the calculated difference with the preset mutation threshold, mark the position where the difference exceeds the mutation threshold as the mutation point of the data structure, use the mutation point as the segmentation boundary, physically split the large file to be transmitted, and obtain multiple data structure partitions with independent data structure characteristics.
[0037] The starting position, length, and data distribution characteristics of each data structure partition are encoded, and a content summary identifier is generated for each partition as a unique data fingerprint identifier for the data structure partition.
[0038] Data fingerprints are extracted from existing data blocks in the historical transmission data in the same way, and the data fingerprints are correlated with the corresponding historical access records. The access order, access frequency and access concentration in the access records are statistically analyzed to obtain the access behavior characteristics of the historical data blocks.
[0039] The data fingerprints of the current data structure partition are matched with the data fingerprints of historical data blocks. The corresponding access behavior features are inherited for the data structure partitions that match successfully, and default behavior parameters are assigned to the data structure partitions that do not match, thereby establishing semantic importance parameters for each data structure partition.
[0040] The semantic importance parameter is bound to the data structure partition, and the data structure partition is used as the basic block unit. The blocks are sorted according to the semantic importance parameter to generate an adaptive block strategy and block transmission priority.
[0041] It should be noted that by performing structural parsing on the file content before segmentation and determining the segmentation boundaries based on the changes in data distribution within the sliding window, the segmentation process can be adaptively executed according to changes in the internal structure of the file. This ensures that each segment maintains the consistency of its internal data structure as much as possible, reducing processing costs in subsequent transmission and retransmission processes.
[0042] To transform data distribution characteristics into quantifiable structural indicators, we perform probabilistic modeling of the data distribution within the sliding window and proportional modeling of continuous repetition features, thereby characterizing the randomness and regularity of the data respectively.
[0043] To ensure a unified quantification of data distribution characteristics, and considering that the probability of byte occurrence reflects the data distribution within a window, the probability of occurrence of each type of byte is used as the basic statistical object. Furthermore, considering that highly random data regions typically exhibit a uniform distribution of multiple byte types, a logarithmic weighting method is employed to amplify the differences between different probability distributions, enabling the effective differentiation of highly random data regions. The information entropy expression is constructed as follows: ; in, Indicates the first Information entropy of a sliding window Indicates the total number of byte categories. Indicates the first The first sliding window The probability of a byte-like object appearing.
[0044] The more uniform the byte distribution, the larger the overall value, thus enabling the identification of highly random data regions and allowing different data structures to be compared on a uniform scale.
[0045] To identify continuous repetitive structures that are difficult to fully represent with information entropy, considering that regular data usually exhibits the concentrated occurrence of consecutive identical bytes or repetitive segments, the length of continuous repetitive data is used as an amplification factor for the regularity feature. By scaling the length of repetitive data to the total window length, the degree of repetition between different windows can be compared on a uniform scale, thus constructing a repetition expression: : in, Indicates the first The repeatability index of a sliding window. Indicates the first The cumulative length of consecutive repeating data within a sliding window Indicates the length of the sliding window.
[0046] The higher the proportion of duplicate data, the larger the indicator value, thereby strengthening the ability to identify regular data and enabling the system to distinguish between areas with high duplication data and areas with high random data.
[0047] To uniformly characterize the degree of data structure change between adjacent sliding windows, the information entropy difference and the repetition difference are jointly modeled. The information entropy difference reflects the randomness of data changes, while the repetition difference reflects the regularity of data changes. This ensures that any abrupt change in any dimension can be reflected in the structural change. To avoid the problem that simple linear superposition cannot reflect the coordinated changes of multi-dimensional features, the information entropy change and the repetition change are coupled and modeled. The response capability when both types of features change simultaneously is enhanced through a product approach. The expression for the structural change is constructed as follows: ; in, Indicates the first The sliding window and the first The amount of structural change between sliding windows and These represent the information entropy of adjacent sliding windows, and These represent the repetition indices of adjacent sliding windows. Represents the weight of the information entropy difference. This represents the weight of the repetition difference.
[0048] This allows for a unified measurement of random and regular variations. By introducing weighting coefficients, the influence of the two types of features under different data types can be adjusted, thereby improving the adaptability of boundary recognition.
[0049] To enable data structure partitions to be identified across different files or transmission tasks, partition location, partition length, and structural features are all incorporated into the fingerprint generation process. By introducing the partition start position, the physical relationship of data blocks within the original file can be recorded. By introducing the partition length, data blocks of different scales can be distinguished. By introducing information entropy and repetition features, the content structure of the data blocks can participate in the identifier generation, thus constructing the following data fingerprint expression: ; in, Indicates the first Data fingerprint identifier for each data structure partition. This represents the function for calculating the summary. Indicates the first The starting position of each data structure partition. Indicates the length of the data structure partition. This represents the information entropy feature corresponding to the partition of the data structure. This indicates the repetition characteristic corresponding to the partition of the data structure.
[0050] By simultaneously introducing information entropy and repetition, the data structure can be jointly described from the two dimensions of "randomness" and "regularity," thereby providing a computational basis for subsequent structural mutation identification.
[0051] To transform historical access behavior into semantic importance parameters that can be used for transmission ordering, access frequency reflects the intensity of data usage. However, excessively high access frequency, if directly involved in the calculation, can lead to the result being dominated by a single factor. A logarithmic function is used to compress the impact of increased access frequency. Data accessed earlier typically has a greater impact on task initiation or critical processes; therefore, access order is used as a reverse adjustment factor. Access concentration reflects the characteristic of data being concentratedly accessed at a specific stage; therefore, access concentration is used as a semantic enhancement factor to construct a semantic importance expression. ; in, Indicates the first The semantic importance parameter of each data structure partition. Indicates access frequency characteristics. Indicates access order characteristics. This indicates the degree of concentration of access.
[0052] This approach gradually flattens the impact of increased access frequency on results, strengthens the priority of key data by using a reverse access order, and enhances the impact of phased data through concentration, thereby achieving a coupled expression of multi-dimensional behaviors.
[0053] By introducing data fingerprinting and behavior inheritance mechanisms, the blocks not only have structural consistency but also semantic importance identifiers, thus providing a basis for subsequent scheduling and realizing the transformation from "structure-driven" to "structure + semantic joint-driven".
[0054] S3. Establish and dynamically adjust the multi-channel SFTP transmission strategy based on the block transmission priority and SFTP node status to transmit each data block concurrently.
[0055] Furthermore, a sliding window scan is performed on the large file to be transmitted according to a preset length. Data within each window is read, byte distribution features and continuous repetition features are extracted, and information entropy and repetition index are calculated respectively, thereby forming a data feature sequence that reflects the local structural changes of the data.
[0056] Traverse the data feature sequence, calculate the difference in information entropy and the difference in repetition index between adjacent sliding windows, compare the calculated difference with the preset mutation threshold, mark the position where the difference exceeds the mutation threshold as the mutation point of the data structure, use the mutation point as the segmentation boundary, physically split the large file to be transmitted, and obtain multiple data structure partitions with independent data structure characteristics.
[0057] The starting position, length, and data distribution characteristics of each data structure partition are encoded, and a content summary identifier is generated for each partition as a unique data fingerprint identifier for the data structure partition.
[0058] Data fingerprints are extracted from existing data blocks in the historical transmission data in the same way, and the data fingerprints are correlated with the corresponding historical access records. The access order, access frequency and access concentration in the access records are statistically analyzed to obtain the access behavior characteristics of the historical data blocks.
[0059] The data fingerprints of the current data structure partition are matched with the data fingerprints of historical data blocks. The corresponding access behavior features are inherited for the data structure partitions that match successfully, and default behavior parameters are assigned to the data structure partitions that do not match, thereby establishing semantic importance parameters for each data structure partition.
[0060] The semantic importance parameter is bound to the data structure partition, and the data structure partition is used as the basic block unit. The blocks are sorted according to the semantic importance parameter to generate an adaptive block strategy and block transmission priority.
[0061] It should be noted that by introducing the semantic importance of data blocks and historical transmission risks into the multi-channel SFTP transmission process, channel allocation is no longer based solely on the current resource availability, but is scheduled in conjunction with the transmission value and stability of the data blocks themselves, thereby improving the transmission priority of critical data and the stability of the overall transmission process.
[0062] Based on block prioritization, historical transmission behavior associated with data fingerprints is introduced to analyze the stability of different data types in historical transmission, thereby identifying data blocks prone to retransmission or interruption. The same type of data often exhibits similar transmission characteristics in different transmission tasks; therefore, risk experience can be reused through data fingerprinting.
[0063] To enable multi-channel transmission strategies to identify the transmission stability of different data blocks, a transmission risk coefficient is constructed based on historical transmission behavior associated with data fingerprints. The retransmission rate reflects the frequency of repeated transmission of data blocks in historical transmissions, the average packet loss rate reflects the instability of data blocks in the transmission link, and the interruption probability reflects the correlation between data blocks and transmission interruptions. This transforms historical transmission behavior into calculable risk parameters. To avoid underestimating the overall risk due to independent weighting of various risk factors, the retransmission rate, packet loss rate, and interruption probability are treated as joint events, and a risk model is constructed using probability complements, expressed as: ; in, Indicates the first The transmission risk coefficient of each data block. This indicates the retransmission rate of the historical data fingerprint corresponding to the data block. This indicates the average packet loss rate corresponding to the historical data fingerprints of the data blocks. This indicates the probability of interruption of the historical data fingerprint corresponding to the data block.
[0064] To ensure that channel scheduling is constrained by both semantic importance and transmission risk, considering that semantic importance should positively increase the transmission priority of data blocks, while transmission risk should negatively suppress the occupation of critical channels by unstable data blocks, semantic importance is used as the base weight, and the risk coefficient is used as an exponential decay factor. This allows low-risk data blocks to maintain their original semantic priority, while gradually reducing the channel allocation weight as the risk increases, avoiding abrupt weight changes caused by simple division. The channel weight expression is constructed as follows: ; in, Indicates the first Channel weights are assigned to each data block. Indicates the first The semantic importance parameter of each data block. Indicates the first The transmission risk coefficient of each data block ensures that low-risk data maintains its original priority, while the weight of high-risk data is gradually reduced, thereby preventing unstable data from occupying critical channel resources.
[0065] To ensure that channel weights can be translated into specific SFTP channel selections, considering that the scheduling weight of a data block only reflects transmission demand, while the available channel resources and current load jointly reflect the channel's carrying capacity, available channel resources are used as a positive matching factor, and real-time channel load is used as a suppressive factor. This is achieved by introducing... This ensures that the matching value is lower for channels with higher load, thus preventing multiple high-priority data blocks from entering the same high-load channel simultaneously. The channel allocation expression is constructed as follows: ; in, Indicates the first Each data block is allocated an SFTP transfer channel. This indicates the weight assigned to the transmission channel for data chunking. Indicates the first The available resource status of each candidate transmission channel. This indicates that the channel with the largest calculated result is selected from all candidate channels. Indicates the first The current real-time load level of each candidate channel.
[0066] This prioritizes allocating data blocks to channels with sufficient resources and low load, thereby improving overall transmission efficiency.
[0067] This approach allows high-importance, low-risk data blocks to receive greater transmission bandwidth weight, while high-risk data is distributed and transmitted in a distributed manner. Through this risk-aware weighted scheduling, critical resources are prevented from being occupied by unstable data blocks, thereby improving the overall stability and efficiency of transmission.
[0068] During the scheduling process, semantic importance and transmission risk are considered together, so that important and stable data are given priority to occupy the transmission channel, while high-risk data is transmitted in a delayed or distributed manner, thereby avoiding resource conflicts.
[0069] This allows multi-channel transmission to no longer rely solely on resource status, but instead integrates data characteristics for scheduling, thereby improving the overall stability and efficiency of transmission.
[0070] S4. Acquire transmission feedback data during the transmission process, and dynamically update the block transmission priority and multi-channel SFTP transmission strategy based on the transmission feedback data.
[0071] Furthermore, the transmission rate, transmission delay, and retransmission status of each data block are collected to obtain real-time transmission performance data.
[0072] The connection status, abnormal interruptions, and bandwidth fluctuations during the transmission process are recorded to obtain transmission status feedback data.
[0073] Statistical analysis is performed on the transmission feedback data, and the actual transmission cost of each block is recalculated to obtain an updated priority sequence.
[0074] The transmission channel parameters are adjusted based on changes in node resources to obtain an updated multi-channel transmission strategy.
[0075] It should be noted that by continuously collecting feedback data such as real-time transmission rate, latency, retransmission, and abnormal interruption during the transmission process, and then applying the feedback data in reverse to the block priority and risk model, the transmission strategy can be dynamically corrected based on the actual operating results, thereby improving the adaptability in complex network environments.
[0076] By collecting data on transmission rate, latency, and retransmission status in real time, and combining this with connection status and anomaly information, transmission feedback data is generated. The actual transmission performance is then compared with the expected scheduling results, thereby identifying the deviation between the scheduling strategy and the actual situation.
[0077] By introducing a feedback mechanism, the transmission system can continuously adjust its priority and resource allocation methods based on actual operating results, thereby gradually approaching the optimal scheduling state, realizing dynamic adaptive adjustment of the transmission strategy, improving the robustness of the system in complex network environments, and constructing a real-time transmission cost to quantify real-time performance deviations during the transmission process.
[0078] To quantify real-time performance deviations during transmission, considering that the deviation between actual and expected transmission delays reflects transmission efficiency anomalies, while the number of retransmissions reflects transmission stability anomalies, both delay deviation and retransmission count are considered as components of real-time transmission cost. When delay anomalies and retransmissions occur simultaneously, it indicates that the transmission anomaly of the data block is more severe than a single anomaly. Therefore, the number of retransmissions is used as an amplification factor for delay deviation, amplifying the composite anomaly through a product coupling method, and constructing the expression for real-time transmission cost: ; in, Indicates the first The real-time transmission cost of data blocks Indicates the first The actual transmission delay of each data block Indicates the first The expected transmission latency of each data block Indicates the first The number of retransmissions for each data block.
[0079] This amplifies the transmission cost when delay anomalies and retransmissions occur simultaneously, thereby enabling more accurate identification of abnormal data blocks.
[0080] To ensure that the real-time transmission cost directly affects the subsequent transmission order, the block priority is dynamically adjusted. By retaining the initial priority, the semantic importance analysis results mentioned above continue to participate in scheduling. An exponential decay term is introduced to introduce the real-time transmission cost, so that the priority of data blocks with poor transmission performance is smoothly reduced in subsequent scheduling. The priority update expression is constructed as follows: ; in, Indicates the first The priority of each data block update Indicates the first The initial priority of each data block, This represents the priority attenuation coefficient. Indicates the first The real-time transmission cost of data blocks.
[0081] This allows data blocks that perform poorly in actual transmission to be dynamically downgraded or have their channels adjusted, thereby enabling the transmission strategy to gradually approach the optimal state during operation and greatly improving the system's robustness in complex network environments.
[0082] To enable real-time feedback to correct the risk model corresponding to the data fingerprint, and considering that not all real-time feedback should update historical risks by the same magnitude, a gating function is introduced to control the update intensity. When the real-time transmission cost is higher than the historical risk, it indicates that the historical risk underestimates the instability of the data block. In this case, the gating factor is increased to allow the model to quickly absorb abnormal experience. When the real-time transmission cost is close to the historical risk, it indicates that the current feedback is within the normal fluctuation range. In this case, the gating factor is small to avoid frequent model oscillations. The risk update expression is constructed as follows: ; in: ; in, Indicates the first The transmission risk coefficient after updating each data block Indicates the first The transmission risk coefficient before updating each data block. Indicates the first The real-time transmission cost of data blocks This indicates the risk update weight.
[0083] This increases the update magnitude when real-time transmission performance deviates from historical risks, while reducing updates within normal fluctuation ranges, thereby ensuring model stability.
[0084] This allows newly generated transmission experience to continuously adjust the weights of historical samples, thus enabling the system to have long-term self-learning capabilities. As transmission tasks are continuously executed, the system accumulates more transmission experience, and the generated transmission characteristic parameters become closer to the actual network and data environment, achieving continuous optimization of transmission performance.
[0085] S5. If the transmission is interrupted or completed, resume the transmission and reassemble the data according to the status of the transmitted blocks and the verification results to obtain the complete file.
[0086] Furthermore, the status of the transmitted segments is recorded and incomplete segments are marked to obtain a set of resume tasks.
[0087] The complete file is obtained by verifying and concatenating all blocks according to the block index order.
[0088] The historical transmission data and the current transmission feedback data are fused to obtain updated training data.
[0089] The transmission feature parameters are recalculated based on the updated training data to achieve continuous optimization of the transmission strategy.
[0090] It should be noted that by continuing to retain the feedback data from the current transmission after breakpoint resumption and file reconstruction, and integrating it with historical transmission data, the experience generated in a single transmission process can be precipitated as reusable training data for subsequent transmission tasks, thereby forming a continuously optimized transmission control mechanism.
[0091] When resuming interrupted transmissions and reassembling data, the feedback data generated during the current transmission process is merged with historical transmission data, so that the newly generated data can supplement historical samples, thereby continuously enriching the behavioral data foundation of the system.
[0092] As transmission tasks are continuously executed, the system can gradually accumulate transmission experience for different types of data, and by updating transmission characteristic parameters, make subsequent transmission strategies more closely aligned with actual conditions. Through a continuous optimization mechanism, the system acquires long-term learning capabilities, thereby improving overall transmission performance. To achieve continuous system updates, strategy optimization factors are constructed based on real-time transmission costs and data fingerprint matching degrees.
[0093] To facilitate the transfer of transmission experience between similar data blocks, considering that real-time transmission cost reflects the importance of the current transmission experience, and data fingerprint similarity reflects the transferability of that experience to other data blocks, data fingerprint similarity is used as the weight for experience propagation. Higher similarity indicates that the two data blocks are closer in structural features and historical behavior, and the current transmission experience is more valuable for subsequent data blocks. Lower similarity reduces the impact of the experience on subsequent data blocks. The following strategy optimization factor expression is constructed: ; in, This represents the strategy optimization factor. Indicates the first The real-time transmission cost of a data block Indicates the first Data fingerprint matching degree of each data block This indicates the number of data blocks involved in this update. Indicates the current number Data fingerprints of data blocks With history Data fingerprints of data blocks Similarity between them This represents the total number of historical data blocks involved in the calculation. This represents a smoothing coefficient to prevent the denominator from being zero, enabling similar data blocks to share transmission experience, thereby achieving experience transfer rather than simple average updates.
[0094] This allows newly generated transmission experience to continuously adjust the weights of historical samples, thus enabling the system to have long-term self-learning capabilities. As transmission tasks are continuously executed, the system accumulates more transmission experience, and the generated transmission characteristic parameters become closer to the actual network and data environment, achieving continuous optimization of transmission performance.
[0095] Information entropy, repetition, structural change, semantic importance parameter, transmission risk coefficient, channel allocation weight, real-time transmission cost, and strategy optimization factor are not calculated in isolation. Instead, they are passed sequentially through the link of "structural identification - semantic transfer - risk scheduling - feedback correction". The calculation results constrain the transmission decision in the next stage, transforming data blocks from ordinary transmission units into identifiable transmission units with structural characteristics, historical behavior characteristics, and risk characteristics, thereby forming a data-driven closed-loop transmission control system.
[0096] This embodiment also provides a supercomputing platform large file intelligent transfer system based on SFTP nodes, including: a data acquisition module, a feature analysis module, a strategy generation module, a dynamic transfer module, and a reconstruction and recovery module.
[0097] The data acquisition module is used to acquire file characteristic data, historical transmission data, and SFTP node status data of the large file to be transferred.
[0098] The feature analysis module is used to perform correlation analysis on file feature data, historical transmission data and SFTP node status data to obtain transmission feature parameters, and generate adaptive chunking strategies and chunk transmission priorities based on the transmission feature parameters.
[0099] The strategy generation module is used to establish and dynamically adjust the multi-channel SFTP transmission strategy based on the block transmission priority and SFTP node status, and to transmit each data block concurrently.
[0100] The dynamic transmission module is used to acquire transmission feedback data during the transmission process and dynamically update the block transmission priority and multi-channel SFTP transmission strategy based on the transmission feedback data.
[0101] The reconstruction and recovery module is used to resume transmission and reconstruct data based on the status of the transmitted blocks and the verification results after transmission is interrupted or completed, so as to obtain a complete file.
[0102] This embodiment also provides a computer device applicable to the intelligent large file transfer method for supercomputing platforms based on SFTP nodes, comprising: a memory and a processor. The memory stores computer-executable instructions, and the processor executes the computer-executable instructions to implement the intelligent large file transfer method for supercomputing platforms based on SFTP nodes as proposed in the above embodiment.
[0103] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0104] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the intelligent large file transfer method for a supercomputing platform based on an SFTP node, as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0105] In summary, this invention achieves the following: by performing structural analysis on file content and combining it with historical access behavior for correlation analysis, large files are divided into data blocks with structural and semantic features. Based on data fingerprints, historical behavior is inherited, transforming data blocks from simple physical partitioning units into semantically important transmission units. This improves the rationality of block division and reduces transmission efficiency losses and retransmission costs caused by unreasonable block division. Furthermore, by jointly modeling semantic importance with risk parameters constructed based on historical transmission behavior and the real-time resource status of SFTP nodes, multi-channel scheduling oriented towards data characteristics and system resources is achieved. This prioritizes high-importance, low-risk data for accessing high-quality channel resources, while high-risk data is distributed or processed later, thereby improving resource utilization efficiency and enhancing overall transmission stability.
[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for intelligent large file transfer on a supercomputing platform based on SFTP nodes, characterized by: include: Obtain file characteristic data, historical transfer data, and SFTP node status data of the large file to be transferred; Based on the file feature data, historical transmission data, and SFTP node status data, correlation analysis is performed to obtain transmission feature parameters, and an adaptive chunking strategy and chunk transmission priority are generated according to the transmission feature parameters. Based on the block transmission priority and SFTP node status, a multi-channel SFTP transmission strategy is established and dynamically adjusted to perform concurrent transmission of each data block. During the transmission process, transmission feedback data is acquired, and the block transmission priority and multi-channel SFTP transmission strategy are dynamically updated based on the transmission feedback data. If the transmission is interrupted or completed, the transmission is resumed and the data is reassembled based on the status of the transmitted segments and the verification results to obtain the complete file.
2. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 1, characterized in that: The adaptive chunking strategy and chunk transmission priority include: The large file to be transmitted is scanned by a sliding window of a preset length. The data in each window is read, the byte distribution features and continuous repetition features are extracted, and the information entropy and repetition index are calculated respectively, thereby forming a data feature sequence that reflects the local structural changes of the data. Traverse the data feature sequence, calculate the difference in information entropy and the difference in repetition index between adjacent sliding windows in turn, compare the calculated difference with the preset mutation threshold, mark the position where the difference exceeds the mutation threshold as the mutation point of the data structure, use the mutation point as the segmentation boundary, and physically divide the large file to be transmitted to obtain multiple data structure partitions with independent data structure features. The starting position, length, and data distribution characteristics of each data structure partition are encoded, and a content summary identifier is generated for each partition as a unique data fingerprint identifier for the data structure partition. Data fingerprints are extracted from existing data blocks in the historical transmission data in the same way, and the data fingerprints are correlated with the corresponding historical access records. The access order, access frequency and access concentration in the access records are statistically analyzed to obtain the access behavior characteristics of the historical data blocks. The data fingerprint of the current data structure partition is matched with the data fingerprint of the historical data block. The data structure partition that matches successfully inherits the corresponding access behavior features, and the data structure partition that does not match is assigned default behavior parameters, thereby establishing semantic importance parameters for each data structure partition. The semantic importance parameter is bound to the data structure partition, and the data structure partition is used as the basic block unit. Each block is sorted according to the semantic importance parameter, thereby generating an adaptive block strategy and block transmission priority.
3. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 2, characterized in that: The multi-channel SFTP transfer strategy includes: The data fingerprint identifier is used to search the historical transmission data, extract the transmission log corresponding to the data block that is the same as or similar to the current data fingerprint, and perform statistical analysis on the transmission duration, retransmission count and transmission interruption status in the transmission log to obtain the historical transmission behavior characteristics corresponding to the data fingerprint. The historical transmission behavior characteristics are normalized, and the transmission stability, retransmission frequency and interruption probability are comprehensively evaluated to establish structural risk parameters for each data fingerprint, which are used to characterize the instability of data during transmission. The structural risk parameters and semantic importance parameters are jointly analyzed, and each data block is comprehensively scheduled and evaluated. Data blocks with high semantic importance are given priority to enter the transmission queue first, while data blocks with low transmission priority are delayed or dispersed according to stability, thereby generating a block scheduling sequence. The SFTP transmission channel is allocated according to the block scheduling sequence. Data blocks with high transmission priority are allocated transmission resources first, and an initial multi-channel transmission strategy is established based on the node's current bandwidth and load. During the transmission process, the actual transmission behavior of each data block is continuously monitored, the actual transmission duration, retransmission count and abnormal situations of the corresponding data fingerprint are recorded, and the data are compared and analyzed with structural risk parameters to obtain risk deviation information. The structural risk parameters of the corresponding data fingerprint are dynamically corrected based on the risk deviation information, and the correction results are fed back to the subsequent data block scheduling and evaluation process, thereby realizing an adaptive transmission scheduling update mechanism based on data fingerprint.
4. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 3, characterized in that: The acquisition of transmission feedback data includes: The transmission rate, transmission delay, and retransmission status of each data block are collected to obtain real-time transmission performance data. The connection status, abnormal interruptions, and bandwidth fluctuations during the transmission process are recorded to obtain transmission status feedback data.
5. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 4, characterized in that: The dynamic update strategy for chunked transmission priority and multi-channel SFTP transmission includes: Statistical analysis of the transmission feedback data is performed, and the actual transmission cost of each block is recalculated to obtain an updated priority sequence. The transmission channel parameters are adjusted based on changes in node resources to obtain an updated multi-channel transmission strategy.
6. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 5, characterized in that: The breakpoint resumption and data reconstruction include: Record the status of the transmitted segments and mark the incomplete segments to obtain a set of resume tasks; The complete file is obtained by verifying and concatenating all blocks according to the block index order.
7. The intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in claim 6, characterized in that: The update of the transmission characteristic parameters includes: By fusing historical transmission data with current transmission feedback data, updated training data is obtained; The transmission feature parameters are recalculated based on the updated training data to achieve continuous optimization of the transmission strategy.
8. A supercomputing platform large file intelligent transfer system based on SFTP nodes, based on the supercomputing platform large file intelligent transfer method based on SFTP nodes as described in any one of claims 1 to 7, characterized in that: It includes a data acquisition module, a feature analysis module, a strategy generation module, a dynamic transmission module, and a reconstruction and recovery module; The data acquisition module is used to acquire file characteristic data, historical transmission data, and SFTP node status data of the large file to be transferred. The feature analysis module is used to perform correlation analysis on the file feature data, historical transmission data and SFTP node status data to obtain transmission feature parameters, and generate an adaptive block splitting strategy and block transmission priority based on the transmission feature parameters. The strategy generation module is used to establish and dynamically adjust the multi-channel SFTP transmission strategy based on the block transmission priority and SFTP node status, and to transmit each data block concurrently. The dynamic transmission module is used to acquire transmission feedback data during the transmission process and dynamically update the block transmission priority and multi-channel SFTP transmission strategy based on the transmission feedback data. The reconstruction and recovery module is used to resume transmission and reconstruct data based on the status of the transmitted blocks and the verification results after transmission is interrupted or completed, so as to obtain a complete file.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent large file transfer method for supercomputing platforms based on SFTP nodes as described in any one of claims 1 to 7.