Data transmission method and device, storage medium and electronic equipment
By dynamically adjusting the data synchronization transmission rate, based on the network bandwidth and load state of the storage device cluster, data security problems between data centers are solved, efficient and stable data synchronization is achieved, and the reliability and business continuity of the storage system are ensured.
Patent Information
- Application Number
- CN202510571481.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-04
- Publication Date
- 2025-07-18
AI Technical Summary
When existing data transmission methods synchronize between data centers, there is a problem of low data security, especially during peak data access periods, it is easy to lead to synchronization delay and data inconsistency, affecting the security of the storage system.
By obtaining network bandwidth information and load status in the storage device cluster, dynamically predict the data synchronization transmission rate, and adjust the synchronization rate according to the device load and network bandwidth, ensuring efficient and timely synchronization of data between storage devices.
The efficient and stable transmission of data synchronization is realized in the storage device cluster, reducing the risk of equipment overload, ensuring that the second storage device can quickly take over services in the event of a failure, reducing the risk of data loss and service interruption, and improving data security.
Smart Images

Figure CN120343020A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage, and particularly relates to a data transmission method, a device, a storage medium, and an electronic device. Background Art
[0002] In the case of the increasing demand of enterprises for data storage, using a dual-active storage system can ensure data security and guarantee the availability of services. A dual-active storage system usually consists of two or more geographically dispersed data centers, and these two or more data centers both have complete service processing capabilities and data storage capabilities, and can provide services externally simultaneously without the distinction between the primary and the standby.
[0003] Currently, in the data transmission methods provided by the related technologies among the above-mentioned data centers, a fixed synchronization rate is usually used to synchronize data between two data centers. However, during the peak data access period, significant synchronization delays will occur when using the above synchronization method, making it impossible to synchronize the latest data status between the two data centers in a timely manner, thereby affecting data consistency. Further, due to data inconsistency between the data centers, the data security of the storage system will also become lower. Summary of the Invention
[0004] The present application provides a data transmission method, a device, a storage medium, and an electronic device to at least solve the problem of relatively low data security in the data transmission method in the related technologies.
[0005] The present application provides a data transmission method, including: performing an operation on the data in the first storage device in a storage device cluster to obtain the target data after the operation, where the storage device cluster includes the first storage device and at least one second storage device;
[0006] Determining a target storage device for receiving the target data from at least one second storage device, and obtaining the reference network bandwidth information between the first storage device and the target storage device, as well as the first load status of the first storage device and the second load status of the second storage device;
[0007] Predicting the data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load status, and the second load status;
[0008] Transmitting the target data to the target storage device according to the data synchronization transmission rate.
[0009] The present application also provides a data transmission device, including a data processing module, configured to perform an operation on the data in the first storage device in a storage device cluster to obtain target data after the operation, where the storage device cluster includes the first storage device and at least one second storage device;
[0010] A storage device status determination module, configured to determine a target storage device for receiving the target data from at least one second storage device, and obtain reference network bandwidth information between the first storage device and the target storage device, as well as the first load status of the first storage device and the second load status of the second storage device;
[0011] A transmission rate prediction module, configured to predict the data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load status, and the second load status;
[0012] A data transmission module, configured to transmit the target data to the target storage device according to the data synchronization transmission rate.
[0013] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data transmission methods when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any one of the above data transmission methods when executed by a processor.
[0015] The present application also provides a computer program product, including a computer program, where the computer program implements the steps of any one of the above data transmission methods when executed by a processor.
[0016] In this application, an operation is performed on the data in the first storage device in the storage device cluster to obtain the target data after the operation. The storage device cluster includes the first storage device and at least one second storage device. The target storage device for receiving the target data is determined from at least one second storage device, and the reference network bandwidth information between the first storage device and the target storage device, the first load status of the first storage device, and the second load status of the second storage device are obtained. Based on the reference network bandwidth information, the first load status, and the second load status, the data synchronization transmission rate between the first storage device and the target storage device is predicted. The target data is transmitted to the target storage device according to the data synchronization transmission rate. In the above manner, in the storage device cluster, the target data in the first storage device that may be accessed can be synchronized to the target storage device. During the data synchronization transmission process, the data synchronization transmission rate can be adaptively updated according to the network bandwidth between the first storage device and the target storage device and the load status of the first storage device and the target storage device, avoiding device overload, so as to ensure that in the case of a storage device failure, the second storage device can quickly take over the service, reducing the risk of data loss and service interruption. Therefore, the technical problem of low data security in the data transmission method in the related art can be solved, and the technical effect of improving data security can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic diagram of the hardware environment of an optional data transmission method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an optional data transmission method according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of an optional data transmission method according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of another optional data transmission method according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of yet another optional data transmission method according to an embodiment of the present application;
[0023] Figure 6It is a structural block diagram of an optional data transmission device according to an embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0025] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0026] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0027] According to one aspect of the embodiments of the present application, a data transmission method is provided. As an optional implementation manner, the above data transmission method can be but is not limited to being applied to a data transmission system in a hardware environment as Figure 1 shown. As Figure 1 shown, the data transmission system may include: a first storage device 102, a second storage device 104, and an arbitration server 114. Among them, the first storage device 102 may include a host group one 106 and a storage system 110, and the second storage device 104 may include a host group two 108 and a storage system 112.
[0028] It should be noted that the data transmission system may be a dual-active storage system. The first storage device 102 may be one of the data centers, and the second storage device 104 may be the other data center. It is possible to achieve real-time synchronization of data and seamless switching of services between two geographically dispersed data centers to ensure business continuity and high data availability. Each data center is in an active running state and can simultaneously process data read and write operations, ensuring that even when one of the data centers fails, the business can quickly and imperceptibly switch to the other data center to continue running, thus greatly enhancing the disaster tolerance ability of the system and the continuity of services.
[0029] The first storage device 102 and the second storage device 104 can be physical locations for centralized storage, processing, and distribution of data, and may include a large number of servers, storage devices, network devices, and infrastructure (such as power, cooling systems) ( Figure 1 only the host and storage devices are shown in the figure), which are used to support the operation of the IT systems of enterprises or organizations. A high-speed network connection can be established between the two data centers to achieve real-time or near-real-time data synchronization between the data centers, ensuring that the data remains consistent among all live points. When a failure of a certain data center is detected, the system can automatically and quickly switch the service traffic to other data centers, and users will not perceive service interruption.
[0030] It should be noted that the hosts in host group one 106 and host group two 108 refer to the computing devices responsible for running various application programs and services in the data center, usually referring to servers. In the data center environment, the hosts undertake core functions such as data processing, business logic execution, and user request response. The hosts are usually equipped with high-performance processors, large-capacity memories, highly reliable storage devices, and high-speed network interfaces to ensure efficient processing of large-scale data and high-concurrency user requests. The types and configurations of the hosts depend on the specific application requirements of the data center. For example, Web servers, database servers, mail servers, etc. Each type of server host has its specific optimized configuration to meet specific service requirements.
[0031] Storage system one 110 and storage system two 112 are one of the core components of the data center, responsible for the persistent storage and management of data. It includes storage hardware (such as hard disks, SSDs, storage arrays, tape libraries, etc.) and storage software (such as file systems, storage virtualization, data management software, etc.). It can provide persistent storage space for various application programs and user data. Protect data from loss and damage through mechanisms such as redundant storage, backup, and snapshots. Improve data access speed through technologies such as caching, read / write optimization, and intelligent tiering. Provide functions such as data life cycle management, data distribution, and data encryption to ensure data security and compliance. The design and scale of the storage system can be adjusted according to the business requirements of the data center, ranging from simple direct-attached storage (DAS), storage area network (SAN) to complex network-attached storage (NAS), object storage, and cloud-based storage services, covering a wide range from basic storage to advanced data services.
[0032] In a data transmission system, the arbitration server 114 ensures data and service status consistency among storage devices. When two data centers are running the same service or application simultaneously, to prevent the occurrence of "split-brain" (i.e., both data centers consider themselves the primary data center simultaneously, resulting in data inconsistency), the arbitration server is responsible for monitoring their status and deciding which data center should continue to operate as the primary active node in case of a failure, while the other automatically switches to a read-only or standby state.
[0033] It should be noted that Figure 1 The examples given are only for illustration. The data transmission system can include not only two storage devices but also multiple storage devices, that is, a data transmission system composed of multiple data centers can span a broader geographical area through a first storage device and multiple second storage devices. Among multiple storage devices, the data stored in each storage device can be kept completely consistent, or different service data can be assigned to each storage device so that it stores fixed service data. For example, all data can be stored in the first storage device and one second storage device, and the data of business department A can be stored in another second storage device.
[0034] The above application scenarios and data transmission systems are only examples. In the face of scenarios that require data transmission, the data transmission method of this application can be used.
[0035] An embodiment of this application provides a data transmission method. Figure 2 It is a flowchart of an optional data transmission method according to an embodiment of this application; as Figure 2 shown, this data transmission method includes:
[0036] Step S202: Perform an operation on the data in the first storage device in the storage device cluster to obtain the target data after the operation, where the storage device cluster includes a first storage device and at least one second storage device;
[0037] It should be noted that, as mentioned above, the storage device cluster can refer to a set composed of multiple storage devices (such as disk arrays, solid-state drives SSDs, network storage devices, etc.), which can provide highly available, high-performance, and scalable data storage services. The storage devices in the cluster can be located in the same geographical area or in data centers at different geographical locations to achieve data redundancy, failover, and load balancing.
[0038] The first storage device may refer to a storage device that performs data operations such as reading, writing, and updating. The first storage device is responsible for receiving read and write requests and performing corresponding operations on the data according to the requests. The second storage device may be a storage device that forms a complementary relationship with the first storage device in a cluster or a dual-active system. The second storage device may exist as a redundant backup of the first storage device to achieve real-time data synchronization or disaster recovery. The target data may refer to the data after the operation of the first storage device, such as the updated new version data or the data processed according to business requirements.
[0039] In an alternative embodiment, when a user or a system operates on data, the first storage device receives the operation request and executes the corresponding data processing task. For example, when an application needs to update data, it sends a write request to the first storage device, and the device will perform a write operation to write the new data into storage. The data after being operated on, that is, the target data, will become the main content of the subsequent synchronization process. The target data will be sent to the second storage device for real-time or priority-based synchronization. The synchronization rate adjustment and data volume difference optimization module will work based on this to ensure efficient and timely data synchronization while optimizing resource usage and avoiding system overload.
[0040] In an alternative embodiment, when the data on any storage system in a certain storage device in a storage device cluster is modified or newly added, in order to maintain data consistency in a dual-active or multi-active architecture, data synchronization is usually triggered immediately or at a set time interval to synchronize the updated data to other storage devices. The storage device where the data is operated on can be determined as the first storage device. The operation on the target data can be a data operation instruction sent by the received user or the modified target data sent by other storage devices to cause the first storage device to synchronize the data.
[0041] In an alternative embodiment, the operation on the data can also be a planned data synchronization by the storage device cluster according to a preset schedule. For example, data can be synchronously timed according to a preset period during periods with low system load such as at night or during non-working hours.
[0042] Step S204, determine a target storage device for receiving the target data from at least one second storage device, and obtain the reference network bandwidth information between the first storage device and the target storage device, as well as the first load status of the first storage device and the second load status of the second storage device;
[0043] It should be noted that the reference network bandwidth information may refer to the available network bandwidth information between the first storage device and the second storage device at the current moment or within a recently measured time period. This information is used to evaluate the possibility and limitations of the data synchronization rate to determine the optimal synchronization strategy. The first load state and the second load state respectively refer to the workload conditions of the first storage device and the second storage device, including CPU usage rate, memory occupancy, disk I / O, and the number of ongoing data processing tasks, etc. By analyzing these states, the system can understand the health status and processing capabilities of the devices, and thus make intelligent resource allocation and rate adjustment decisions.
[0044] In an alternative embodiment, it is necessary to identify the second storage device that receives the synchronization data. This can be selected according to whether the second storage device is to store the target data. For example, the first storage device stores data of all business departments of the entire enterprise, and among multiple second storage devices, there may be a second storage device that stores data of all business departments of the entire enterprise and a second storage device that stores data of some business departments. The second storage device that stores the target data or plans to store the target data can be determined as the target storage device.
[0045] Next, the first storage device can query the latest network bandwidth situation between the first storage device and the selected target storage device, which can be obtained through a dedicated network monitoring tool, historical data statistics, or real-time network performance testing. The reference network bandwidth information directly determines the maximum possibility of the synchronization rate because the data transmission speed is hard-limited by the network bandwidth. At the same time, the current load states of the two devices need to be collected, including but not limited to indicators such as CPU usage rate, memory occupancy, and disk I / O. Through this information, it can be judged whether there are sufficient computing and storage resources to support the synchronization operation, and whether the synchronization rate needs to be adjusted to avoid further burdening the devices.
[0046] Step S206, based on the reference network bandwidth information, the first load state, and the second load state, predict the data synchronization transmission rate between the first storage device and the target storage device;
[0047] It should be noted that the data synchronization transmission rate may refer to the speed at which data is transmitted from the first storage device to the target storage device, usually measured in bytes per second (such as Mbps, Gbps) or the number of data blocks transmitted per unit time. It is a key indicator of data synchronization efficiency and is directly affected by the network bandwidth and the load state of the storage device.
[0048] In an alternative embodiment, after collecting the current load status and reference network bandwidth information of the first storage device and the target storage device, based on the collected information, a prediction algorithm is used to estimate the optimal synchronization rate. Various prediction models can be applied, such as simple moving average, exponential smoothing, ARIMA model based on historical trends, or machine learning prediction models such as neural networks or decision trees. The goal of the prediction is to make full use of network resources without overloading either device, so as to achieve efficient data synchronization. The specific process will be described in detail below and will not be elaborated here. After the prediction is completed, the system will adjust the data synchronization rate to the predicted value in real time. This adjustment can be achieved at the software level (such as adjusting the parameters of the data transmission protocol, such as the TCP window size), the hardware level (such as adjusting the traffic control of the switch), or through middleware services. The adjusted rate needs to be continuously monitored to ensure the smooth progress of the synchronization process and data integrity.
[0049] In an alternative embodiment, the analysis of the load status can be to calculate the load status of the first storage device and the second storage device respectively. This can be done by analyzing the historical data and current status of indicators such as CPU utilization, memory usage, and disk I / O rate, and finding the balance point between the two to ensure that the synchronization task does not overload either side and can be completed as quickly as possible within the allowable bandwidth. Combining the predicted network bandwidth and the load analysis results, a rate adjustment algorithm can be designed. For example, if the predicted bandwidth is sufficient and the loads of both devices are within an acceptable range, the algorithm may recommend a higher synchronization rate. Conversely, if the load of any one side is detected to be close to saturation or the predicted bandwidth is below the threshold, the algorithm will automatically reduce the synchronization rate to prevent the synchronization process from affecting normal business operations.
[0050] Step S208, transfer the target data to the target storage device according to the data synchronization transfer rate.
[0051] In an alternative embodiment, once the data synchronization transfer rate is calculated, the first storage device will send the target data to the target storage device through a high-speed network connection (such as fiber optic, high-speed Ethernet, etc.). The high-speed network connection is the physical basis for data synchronization, which provides high-bandwidth and low-latency data transmission capabilities. During the sending process, the data will be transmitted at the calculated dynamic rate, and the system will continuously monitor the network status and the load of the target storage device during the transmission process to ensure the stability and efficiency of the transmission process.
[0052] After receiving the target data, the target storage device will perform checks on data integrity and consistency to ensure that the data has not been damaged or lost during transmission. After passing the checks, the data will be written to the disk or memory of the target storage device to achieve redundant storage of the data. Through this synchronization process, even if the first storage device fails, the target storage device can immediately take over data access and business processing to ensure business continuity and high data availability. After receiving the target data, the target storage device can update its data directory and index so that subsequent read and write operations can correctly identify and locate the data. At the same time, the system will use a heartbeat mechanism and status synchronization to ensure real-time communication between the two storage devices and the consistency of the synchronized status.
[0053] Figure 3 is a schematic diagram of an optional data transmission method according to an embodiment of the present application; as Figure 3 shown, the first storage device 302 is ready to send target data to the target storage device 308. The first storage device 302 may include a plurality of computing nodes and a plurality of storage nodes ( Figure 3 listing computing node one 304 and storage node one 306 therein), and the target storage device 308 may also include a plurality of computing nodes and a plurality of storage nodes ( Figure 3 listing computing node two 310 and storage node two 312 therein).
[0054] In the case of sending target data to the target storage device 308, computing node one 304 may execute step S302 to send a load status request instruction to computing node 310 of the target storage device 308 to obtain the second load status of the target storage device 308. Computing node two 310 may execute step S304 to send the second load status to computing node one 304. In computing node one 304 of the first storage device 302, reference network bandwidth information may also be obtained. The reference network bandwidth information may be in an external server or in an internal server of the first storage device 302, and there is no limitation thereto. After obtaining the reference network bandwidth information, the first load status, and the second load status, computing node one 304 may calculate the data synchronization transmission rate. The first storage device 302 may execute step S306 to send the target data in storage node one 306 to storage node two 312 according to the data synchronization transmission rate.
[0055] Through this application, operations are performed on the data in the first storage device in the storage device cluster to obtain the target data after the operations. The storage device cluster includes the first storage device and at least one second storage device. The target storage device for receiving the target data is determined from at least one second storage device, and the reference network bandwidth information between the first storage device and the target storage device, the first load status of the first storage device, and the second load status of the second storage device are obtained. Based on the reference network bandwidth information, the first load status, and the second load status, the data synchronization transmission rate between the first storage device and the target storage device is predicted. The target data is transmitted to the target storage device according to the data synchronization transmission rate. In the above manner, the target data in the first storage device that may be accessed in the storage device cluster can be synchronized to the target storage device. During the data synchronization transmission process, the data synchronization transmission rate can be adaptively updated according to the network bandwidth between the first storage device and the target storage device and the load status of the first storage device and the target storage device, avoiding device overload, so as to ensure that in the case of a storage device failure, the second storage device can quickly take over the service, reducing the risk of data loss and service interruption. Therefore, the technical problem of low data security in the data transmission method in the related art can be solved, and the technical effect of improving data security can be achieved.
[0056] In an alternative embodiment, predicting the data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load status, and the second load status includes: predicting the target network bandwidth information between the first storage device and the target storage device based on the reference network bandwidth information; determining the target device resources corresponding to the target data according to the data access frequency of the target data, the first load status, and the second load status, where the target device resources are used to indicate the computing resources and storage resources for the target data transmission; and determining the data synchronization transmission rate by using the target network bandwidth information and the target device resources.
[0057] It should be noted that the target network bandwidth information can indicate the amount of network bandwidth that can stably transmit data between the expected first storage device and the target storage device. Predicting the target network bandwidth information provides a basis for determining the actual transmission rate. The target data access frequency refers to the frequency at which the target data is accessed or operated in a recent period of time, and is usually used as an important indicator for evaluating data popularity and synchronization priority. The target device resources may refer to the computing resources and storage resources of the target storage device for receiving, processing, and storing the target data. The computing resources involve the CPU, memory, etc., and the storage resources are related to disk space and I / O rate. Determining the target device resources helps the system evaluate the current bearing capacity of the target storage device and ensure the efficient execution of data synchronization.
[0058] In an alternative embodiment, before data synchronization, a most suitable transmission rate is predicted based on the current network conditions and device load conditions through a mathematical model and algorithm. First, the first storage device uses historical data and real-time monitoring results to predict the target network bandwidth information; secondly, according to the access frequency of the target data, combined with the load status of the first storage device and the target storage device, the demand for target device resources is determined; finally, the target network bandwidth information and the target device resource demand are comprehensively considered to calculate a data synchronization transmission rate that can fully utilize network resources without overloading the device.
[0059] Example 1: A dual-active data center (storage device cluster) of a large institution needs to synchronize a large amount of transaction data during peak business hours. This data not only has a large volume but also a very high access frequency, and has extremely high requirements for the synchronization rate. However, due to network bandwidth fluctuations and load differences of storage devices, the traditional fixed synchronization rate strategy cannot meet the requirements.
[0060] The system monitors network performance, obtains the current bandwidth level and historical values, and predicts the target network bandwidth information. Analyzing the access frequency of the target data, it is found that it is hot data, with the frequent access times reaching more than 1000 times per second. Combining the load status of the first storage device and the target storage device (the CPU usage rate of the first storage device is 80%, and the usage rate of the target storage device is 60%), the system evaluates that the target storage device needs additional CPU resources (such as an increase of 20%) and storage I / O resources to process the synchronization task. Based on the target network bandwidth information (such as 9.2 Gbps) and target device resources (additional CPU and I / O resources), the system calculates the data synchronization transmission rate. Considering the load of the target storage device and resource reservation, the system decides to set the synchronization rate to 8 Gbps to ensure that the data can be synchronized quickly while avoiding overloading the target storage device.
[0061] Through the above embodiments of the present application, predicting the target network bandwidth information and evaluating the target device resource requirements can intelligently determine the data synchronization transmission rate, ensuring that data can be efficiently and stably synchronized under any network and device conditions, while ensuring the normal operation of the device and the high availability of the system.
[0062] In an alternative embodiment, determining the target device resources corresponding to the target data according to the data access frequency of the target data and the first load state and the second load state includes: calculating a target resource weight matching the target data according to the data access frequency, service correlation degree of the target data, and the predicted transmission time of the target synchronous transmission task corresponding to the target data, where the predicted transmission time is used to indicate the expected time to complete the synchronous transmission task, and the service correlation degree is used to indicate the degree of association between the data and the business process; respectively calculating reference resource weights matching each reference synchronous transmission task according to the data access frequency, predicted transmission time, and service correlation degree of the reference data transmitted by each reference synchronous transmission task; and allocating the total device resources of the first storage device and the target storage device according to the target resource weight, at least one reference resource weight, and the first load state and the second load state to obtain the target device resources.
[0063] It should be noted that the target resource weight matching the target data is a quantitative index used to measure the priority of the target data synchronization task when allocating system resources. The weight calculation takes into account the data access frequency, service correlation degree, and predicted transmission time. Among them, the data access frequency represents the data popularity, the service correlation degree reflects the importance of the data to business continuity, and the predicted transmission time estimates the time length required for synchronization.
[0064] The reference synchronous transmission tasks can be data synchronous tasks that are being carried out in parallel. Their characteristics (such as data access frequency, predicted transmission time, service correlation degree) can be used as a reference benchmark for formulating the current data synchronization strategy. The reference resource weight is similar to the target resource weight, but it corresponds to the reference synchronous transmission task and is used to compare and adjust the resource allocation strategy of the current synchronization task to ensure the rationality of resource allocation and optimize the synchronization efficiency.
[0065] The total device resources include all the resources available for the first storage device and the target storage device, such as CPU, memory, disk I / O capabilities, and network bandwidth, etc. These resources play a decisive role in the data synchronization task, and reasonable allocation can significantly improve the synchronization efficiency and system performance.
[0066] In an alternative embodiment, when allocating resources, the access frequency, business relevance, and estimated transmission time of the target data are comprehensively considered to calculate its resource weight. This aims to prioritize data synchronization tasks that are frequently accessed, have a greater impact on the business, and have an urgent transmission time, ensuring the timely synchronization of critical data and improving business continuity. Before determining the resource weight of the target data, past reference synchronization transmission tasks are first analyzed, and their respective reference resource weights are calculated based on their data access frequency, predicted transmission time, and business relevance. By analyzing the weights of historical tasks, the system can better understand the resource demand patterns of various tasks and provide a basis for resource allocation for new tasks. Based on the resource weight of the target data, the resource weights of historical reference tasks, and the load status of the current first storage device and the second storage device (target storage device), intelligent resource allocation is performed. The goal is to maximize the efficiency of the current data synchronization task without affecting the performance of other tasks, ensure that the data synchronization process is fast and stable, and at the same time take into account the load balance of the entire system.
[0067] In an alternative embodiment, the priority can be determined based on the weight of the task (the target synchronization transmission task for transmitting the target data and the reference synchronization transmission task for transmitting the reference data) and the expected completion time of the task. In the synchronization task management of the first storage device, each data synchronization task is assigned a weight, which can reflect the urgency and importance of the task. For example, a hot data synchronization task may be assigned a higher weight due to its high access frequency and business criticality. In addition, the algorithm also estimates the completion time of each task, which can be comprehensively considered based on factors such as task size, network bandwidth, and response time of the storage system.
[0068] For example, the weight can be obtained by performing a weighted sum of the data size, access frequency, estimated completion time, business relevance, etc. And based on the calculated resource weight of each synchronization task, the network bandwidth and storage resources are dynamically adjusted to prioritize tasks with high resource weights, ensuring that important data can be synchronized quickly, and at the same time avoiding overloading a certain data center or node due to processing a large number of low-priority tasks.
[0069] Example 2: Suppose in a storage device cluster of a company, there are two main data centers - Data Center A (the first storage device) and Data Center B (the target storage device). Data Center A needs to synchronize the processed transaction data (target data) to Data Center B.
[0070] The first storage device can monitor that the access frequency of the target data in the past day is very high (e.g., exceeding 90% of the access requests), and it is closely related to the company's high-frequency trading business. At the same time, according to the current network condition and the prediction model, the estimated time to complete the synchronization is about 30 seconds. Based on this information, the system calculates a high resource weight (e.g., the weight value is 85) for this synchronization task. It can calculate parallel transaction data synchronization tasks and find that under similar conditions (i.e., high data access frequency, strong business association, and similar predicted transmission times), the average resource weight of these tasks is about 80.
[0071] Real-time monitoring shows that the CPU usage rate of the first storage device is 60% and the memory usage rate is 70%, while the corresponding metrics of the target storage device are 50% and 65% respectively. Considering the current load status, as well as the resource weights of the target task and the reference task, it can be decided to allocate more CPU and memory resources to this synchronization task, with the expectation of successfully completing the data synchronization within 30 seconds, while ensuring there are sufficient resources to receive and process the large amount of incoming data to avoid causing resource bottlenecks.
[0072] After the resource allocation is completed, immediately start the data synchronization task at the calculated synchronization rate (such as 1 GBps). During the synchronization process, continuously monitor the load status of the two storage devices and the network bandwidth usage to ensure the smooth progress of data synchronization while avoiding negative impacts on the existing business operations.
[0073] Through the above implementation manners of the present application, by using the fine resource weight calculation and dynamic resource allocation strategy, the efficiency and stability of the storage device cluster in processing large-scale and hot data synchronization tasks are effectively improved, providing a solid technical support for ensuring the continuous operation of the business.
[0074] In an optional implementation manner, before determining the target device resources corresponding to the target data according to the data access frequency of the target data and the first load status and the second load status, it includes: calculating the weighted sum of the data access frequency and the business association degree of each of at least one candidate data to obtain the priority attribute corresponding to each candidate data, where at least one candidate data includes the target data; sorting the candidate data according to the size of their corresponding priority attributes to obtain a candidate data sequence, where the candidate data with a larger priority attribute is located in a more forward position in the candidate data sequence; determining the first N candidate data in the candidate data sequence except the target data as reference data, where N is an integer greater than 0.
[0075] It should be noted that candidate data refers to the set of all data that may need to be synchronized or has already been synchronized in the storage device cluster. This data includes not only the target data but also other data blocks that need to be synchronized simultaneously with or may need to be synchronized recently with the target data. The data access frequency measures the number of times data is read or written within a specific time period and is an important indicator for judging data popularity. For hot data, its access frequency is usually much higher than that of other data, and synchronization needs to be prioritized to ensure data consistency.
[0076] The priority attribute is a measure obtained by weighted calculation of the data access frequency and business relevance of candidate data, and is used to determine the priority order of data synchronization. The higher the priority attribute of a data block, the more forward its position in the synchronization queue, and it will be synchronized first. The candidate data sequence is the result of sorting all candidate data according to the size of the priority attribute, where data with a high priority attribute is located in the front position of the sequence. This sequence is used to guide the order of data synchronization to ensure the timely synchronization of hot data and critical business data.
[0077] In an alternative embodiment, in the storage device cluster, data synchronization is a resource-intensive operation, especially in an environment with a large amount of data and limited network bandwidth. Before data synchronization, the first storage device calculates the priority attributes of all candidate data. By weighted calculation of the data access frequency and business relevance of candidate data, it is ensured that data with high business value and frequent access requirements can obtain a higher priority.
[0078] In an alternative embodiment, the priority attribute can be calculated by the following formula:
[0079] W i =α·A i +β·B i
[0080] Wherein, W i can be the priority attribute of the i-th candidate data, A i is the data access frequency corresponding to the i-th candidate data, and B i is the business relevance corresponding to the i-th candidate data. α and β are the calculation weights of the data access frequency and business relevance, which can be set in advance.
[0081] Figure 4 is a schematic diagram of another alternative data transmission method according to an embodiment of the present application; as Figure 4 (a) shows, in the first storage device, the candidate data sequence 404( Figure 4 in which B1 to B N)Sort according to the priority attribute and prepare for synchronous transmission to the target storage device. At this time, the target data 402 is also to be sent. After calculating the priority attribute of the target data 402, it is calculated that its priority attribute is greater than B2. Therefore, the candidate data sequence needs to be updated. As Figure 4 shown in (b), the target data 402 is in the second position of the sequence, and the updated candidate data sequence 406 is obtained.
[0082] Embodiment 3: Suppose there are the following three data blocks to be synchronized in the first storage device: data block A (financial transaction data), data block B (user login log), and data block C (system maintenance record). The access frequency of data block A is 1000 times per second, and the business relevance is 0.9 (financial transaction data is crucial for the operation of the business); the access frequency of data block B is 500 times per second, and the business relevance is 0.6; the access frequency of data block C is 100 times per second, and the business relevance is 0.4.
[0083] According to the weighted calculation formula (priority attribute = α * data access frequency + β * business relevance), assuming that α and β are the weight coefficients of the data access frequency and business relevance respectively, and are set to 0.7 and 0.3 respectively.
[0084] Then:
[0085] The priority attribute of data block A = 0.7 * 1000 + 0.3 * 0.9 = 700 + 0.27 = 700.27;
[0086] The priority attribute of data block B = 0.7 * 500 + 0.3 * 0.6 = 350 + 0.18 = 350.18;
[0087] The priority attribute of data block C = 0.7 * 100 + 0.3 * 0.4 = 70 + 0.12 = 70.12.
[0088] According to the calculated priority attributes, a candidate data sequence is generated: data block A > data block B > data block C. Assuming N is set to 2, the system determines data block A and data block B as reference data and synchronizes them preferentially.
[0089] Subsequently, the first storage device will dynamically adjust the synchronization rates of data blocks A and B according to the current network bandwidth and the load status of the storage device, and preferentially utilize system resources for synchronization. For example, if the network bandwidth is sufficient, data blocks A and B will be transmitted at a higher synchronization rate to ensure timeliness; while if the network bandwidth is tight, the system will adjust the synchronization rate to preferentially ensure the synchronization of data block A because both the data access frequency and business relevance of A are higher than those of B.
[0090] Through the above embodiments of the present application, data blocks with high access frequencies and high service correlation degrees can be preferentially processed, ensuring the real-time synchronization of critical data, thereby improving data consistency and service continuity. At the same time, the dynamic allocation of system resources is fully considered, avoiding resource waste and system overload.
[0091] In an alternative embodiment, determining the first N candidate data in the candidate data sequence as reference data except for the target data includes: in the case where the data volume of the reference data is greater than a preset data threshold, splitting the reference data to obtain a plurality of reference sub-data; generating a reference synchronous transmission task for each reference sub-data.
[0092] It should be noted that the preset data threshold can be a preset data volume size standard, used to determine whether data needs to be split into smaller parts for synchronization to optimize data transmission efficiency and avoid network congestion. The reference sub-data is the smaller data blocks obtained by the first storage device when the data volume of the reference data exceeds the preset data threshold. Each sub-data contains a part of the original data content, but has an independent data identifier and synchronization requirement. The reference synchronous transmission task is a synchronous task generated for each reference sub-data or reference data that does not exceed the threshold. These tasks contain all the necessary information for data transmission, such as data source, target address, transmission rate, priority, etc., and are used to guide the synchronization module to complete the data transmission work.
[0093] In an alternative embodiment, when the data volume of the data in the candidate data sequence exceeds the preset data threshold, in order to avoid network congestion or long-term data stagnation caused by synchronizing a single large file, the first storage device will automatically split such data to generate a plurality of smaller reference sub-data. Each reference sub-data will be assigned an independent synchronization task, which can not only utilize higher network bandwidth for multi-task parallel transmission, but also avoid system response sluggishness caused by a single large data synchronization.
[0094] Embodiment 4: Suppose in a large e-commerce company, a large amount of order data is generated during the daily transaction peak period, and this data needs to be synchronized to multiple storage devices in real time to ensure data consistency and high service availability. At a certain moment, the system detects that a batch of new order data needs to be synchronized, but due to the large data volume (exceeding 1GB), directly synchronizing may consume too much network resources and affect other synchronization tasks.
[0095] The first storage device arranges all the data to be synchronized into a sequence according to priority rules, such as access frequency, business importance, etc. Select the first N pieces of data as reference data, which are the order data blocks generated frequently recently. When it is detected that the data volume of a certain reference data (for example, the summary of order data generated within one hour) exceeds the preset data threshold (assumed to be 500MB), the system will automatically split the data into multiple reference sub-data smaller than the threshold. For example, 1GB of data can be split into 2 sub-data of 500MB each. Generate an independent reference synchronization transfer task for each reference sub-data. Each task carries the metadata of the sub-data (such as identifier, source, destination, priority, etc.), and specifies a reasonable synchronization rate to ensure that the data can be successfully transferred without affecting the overall system performance. These tasks are added to the synchronization task queue and processed in the order of their respective priorities. The first storage device dynamically adjusts the execution timing and synchronization rate of each task according to the current network bandwidth and the load status of the storage device. For example, when there is sufficient network resources, multiple tasks may be synchronized simultaneously, while when resources are scarce, they will be executed step by step in the order of priority.
[0096] Through the above implementation manners of the present application, for big data that exceeds the preset data threshold, it is further split into smaller reference sub-data, and a separate synchronization task is generated for each sub-data. This can not only make full use of network resources, but also ensure the efficiency and accuracy of data synchronization, while reducing the processing pressure on the source device and the target device. It optimizes the use of network resources and avoids the overall system delay caused by a single task. This strategy performs particularly well in processing high-concurrency data synchronization during peak periods, significantly improving the performance and reliability of the active-active storage system in actual business applications.
[0097] In an alternative implementation manner, determining the data synchronization transfer rate by using the target network bandwidth information and the target device resources includes: calculating a first ratio of the data volume of the target data to the total data volume, where the total data volume is used to indicate the sum of the data volume of the target data and the data volume of the reference data; determining the target sub-network bandwidth information corresponding to the target data by multiplying the target network bandwidth information by the first ratio; and determining the data synchronization transfer rate by using the target sub-network bandwidth information and the target device resources.
[0098] It should be noted that the data volume of the target data may refer to the actual size of the target data on the first storage device, including file size, the number of database records, the total volume of data blocks, etc. The total data volume may refer to the sum of the data volume of the target data and the reference data. The reference data may include other data items in the current synchronization queue, or refer to the total amount of data expected to be processed by the system within a certain time window. The total data volume reflects the overall picture of the data involved in the synchronization process of the storage device cluster and is an important basis for calculating the data synchronization priority and rate.
[0099] The target sub-network bandwidth information can represent the share of network bandwidth that can be used during the synchronization of the target data. The target sub-network bandwidth information is dynamically allocated according to the unique requirements of the target data and its importance in the entire system.
[0100] In an alternative embodiment, the first storage device quantifies the proportion of the data volume of the target data in the current or expected total data volume, i.e., the first ratio. This step helps to evaluate the importance or urgency of the target data. Next, the predicted target network bandwidth information is multiplied by the first ratio to obtain the bandwidth share that the target data can use, i.e., the target sub-network bandwidth information. Finally, based on the target sub-network bandwidth information and the current computing resources and storage resource status of the target storage device, the most appropriate data synchronization transfer rate is determined.
[0101] Specifically, the target sub-network bandwidth information can be calculated by the following formula:
[0102] R = C·D i / D total
[0103] where R is the target sub-network bandwidth information, D i is the data volume of the target data, D total is the total data volume, and C is the target network bandwidth information.
[0104] Figure 5 is a schematic diagram of another alternative data transmission method according to an embodiment of the present application; as Figure 5 shown, the first storage device 502 synchronously transfers multiple data (including target data and reference data) to the target storage device 504 in parallel. N transmission tasks (including the target synchronization transmission task 506 and N - 1 reference synchronization transmission tasks) can be executed in parallel simultaneously. When allocating network bandwidth resources, the data volumes of the target data and N - 1 reference data can be summed to obtain the total data volume, and the ratio of the data volume of each data to the total data volume is calculated and then multiplied by the target network bandwidth information to obtain the sub-network bandwidth information corresponding to each data.
[0105] Embodiment 5: Suppose a cloud computing service provider is running a storage device cluster to support a large-scale online trading system. During a peak trading period, the first storage device needs to synchronize a large amount of order data (i.e., "target data"), which is both hot data and a key component of the business.
[0106] If: the data volume of the target data (order data): 100 GB;
[0107] the data volume of the reference data (such as user information, historical transaction records, etc.): 200 GB;
[0108] Total data volume: 300 GB;
[0109] The first ratio (target data / total data volume) is 100 GB / 300 GB = 0.333.
[0110] Currently predicted target network bandwidth information: 10 Gbps;
[0111] Target sub-network bandwidth information: 10 Gbps * 0.333 = 3.33 Gbps.
[0112] This means that considering the proportion of the target data in the entire data synchronization task, the system can reserve 3.33 Gbps of network bandwidth for the efficient synchronization of order data.
[0113] If the load status of the first storage device (source device): CPU usage rate 70%, memory occupancy 60%; the load status of the second storage device (target device): CPU usage rate 40%, memory occupancy 30%, it means that the target storage device currently has sufficient computing and storage resources to handle the synchronization task of the target data.
[0114] Based on the above information, the data synchronization transmission rate can be set to 3.0 Gbps by using the target sub-network bandwidth information (3.33 Gbps) and the availability of the target device resources. This rate takes into account the floating space of the network bandwidth and the possible additional load of the target device to ensure the fast synchronization of data, while avoiding device overload and maintaining the stable operation of the system.
[0115] Through the above implementation manners of the present application, it is ensured that while meeting the data synchronization requirements, the use of network and storage resources can be optimized, resource waste or device overload can be avoided, thereby improving the overall performance and disaster recovery ability of the dual-active storage system.
[0116] In an optional implementation manner, predicting the target network bandwidth information between the first storage device and the target storage device based on the reference network bandwidth information includes: calculating the mean value of at least one first network information in the reference network bandwidth information set to obtain the first network bandwidth information mean value, where the first network information is used to indicate the historical network bandwidth information between the first storage device and the target storage device; performing a weighted sum of the second network bandwidth information and the first network bandwidth information mean value in the reference network information set to obtain the target network bandwidth information, where the second network bandwidth information is used to indicate the current network bandwidth information between the first storage device and the target storage device.
[0117] It should be noted that the reference network bandwidth information set is a collection of historical and current network bandwidth information, which is used to predict the future data synchronization transmission rate. It includes the bandwidth data of the network connection between the first storage device and the target storage device over a past period of time, as well as the currently measured real-time bandwidth information.
[0118] In the reference network bandwidth information set, the first network information specifically refers to the historical network bandwidth information. By calculating the mean value of these historical information, the mean value of the first network bandwidth information is obtained, which serves as the basis of the prediction model. The mean value calculation can adopt simple average, weighted average or other statistical methods, depending on the characteristics of the data and the accuracy requirements of the prediction. Different from the first network information, the second network bandwidth information reflects the current network bandwidth status between the first storage device and the target storage device. It is usually the result of real-time monitoring and is used to capture instantaneous network performance fluctuations.
[0119] The target network bandwidth information is calculated by the system based on the mean value of the first network bandwidth information and the second network bandwidth information through weighted summation, predicting the available network bandwidth amount between the first storage device and the target storage device in the future for a period of time. The calculation of the target network bandwidth information combines historical stability and current real-time nature, providing a more accurate and dynamic reference value for the determination of the data synchronization transmission rate.
[0120] In an alternative embodiment, the first network information (historical network bandwidth information) is selected from the reference network bandwidth information set, and the mean value of these information is calculated to obtain the mean value of the first network bandwidth information. This mean value reflects the average level of the historical network bandwidth and is a stable indicator of the long-term network performance. Then, the second network bandwidth information (current real-time network bandwidth information) is introduced and weighted summation is performed with the mean value of the first network bandwidth information to finally obtain the target network bandwidth information. The process of weighted summation allows the system to adjust the prediction value according to the actual situation of the current network, ensuring that the predicted network bandwidth information can be closer to the real-time network performance, thereby supporting a more reliable prediction of the data synchronization transmission rate.
[0121] Embodiment 6: Suppose a storage device cluster of a financial company needs to synchronize a large amount of transaction data during the peak period. To ensure the efficiency and stability of data synchronization, it is necessary to accurately predict the network bandwidth between the first storage device (i.e., the storage node of the main data center) and the target storage device (i.e., the storage node of the backup data center).
[0122] The system automatically records the network bandwidth data between the first storage device and the target storage device in the past week to form a reference network bandwidth information set. This data includes the peak and trough periods of the network, providing a comprehensive historical perspective for subsequent prediction. Calculate the mean value of the historical data in the reference network bandwidth information set to obtain the first network bandwidth information mean. Assume that the average network bandwidth in the past week is 10 Gbps.
[0123] At the same time, the system monitors the network bandwidth between the current first storage device and the target storage device in real time to obtain the second network bandwidth information. Assume that the current real-time bandwidth is 12 Gbps, indicating that the network performance is temporarily better than the historical average level.
[0124] Use the weighted summation formula to calculate the target network bandwidth information. The formula can be:
[0125] B pred = ε·B avg +(1 - ε)·B curr
[0126] where B pred is the target network bandwidth information, B avg is the historical average bandwidth (the first network bandwidth information mean), B curr is the current real-time bandwidth (the second network bandwidth information), and ε is the weighting factor, generally between 0 and 1. In this embodiment, assume that ε = 0.6, then the target network bandwidth information B pred = 0.6 * 10+(1 - 0.6) * 12 = 10.8 Gbps.
[0127] Through the above calculation, the target network bandwidth information predicted by the system is 10.8 Gbps, which means that when synchronizing data, the synchronization transfer rate can be set to a value close to 10.8 Gbps.
[0128] Through the above implementation manner of the present application, both the current network advantages are utilized and the historical stability is considered, thus ensuring the efficiency and reliability of data synchronization. This prediction and adjustment mechanism has a significant effect on improving the overall performance and user satisfaction of the dual-active storage system in scenarios with extremely high requirements for data synchronization rate such as finance.
[0129] It should be noted that in a dual-active storage system (storage device cluster), the efficiency of data synchronization and the utilization of network resources are crucial. In order to further improve the intelligence and flexibility of data synchronization, a comprehensive mechanism of dynamic perception of business requirements, adaptive scheduling of network status, and prediction of data prefetching can be introduced, aiming to optimize the processing of hot data while ensuring the effective synchronization of non-hot data and improving the overall performance and user experience of the system.
[0130] In an alternative embodiment, the system business status can be monitored in real time. By analyzing the traffic, access patterns, and priority rules, the synchronization priority and rate of data can be adjusted dynamically. For example, during peak trading periods, the synchronization priority of transaction data is automatically increased, while during the operation and maintenance period, the synchronization weight of logs and maintenance data is increased. Specifically, communication with the business application layer can be achieved through an API interface to collect real-time business data, including the number of accesses, transaction amounts, user behaviors, etc. The priority weight Wi of the data is calculated by combining DPI (Dynamic Priority Index) and ΔQ (data update frequency) and adjusted dynamically.
[0131] In an alternative embodiment, the segmentation strategy and transmission rate of data synchronization can be adjusted dynamically based on real-time network monitoring and prediction models. It can intelligently identify network congestion situations, dynamically adjust the segmentation size and rate of synchronization tasks, optimize network resource allocation, and reduce synchronization latency. By interacting with the network monitoring module, data such as network latency, packet loss rate, and bandwidth usage are collected to predict the future network bandwidth status, and the segmentation size and transmission rate of the reference sub-data are automatically adjusted according to the prediction results.
[0132] Then, through machine learning and historical data analysis, the data that may become hot in the next period can be predicted in advance, and it can be pre-loaded into the cache or synchronized to the target data center in advance to reduce the user waiting time and improve the user experience. Specifically, based on data access history, user behavior patterns, seasonal and time periodicity analysis, etc., the trend of data heat change can be predicted, and potential hot data can be pre-read or synchronized in advance to reduce synchronization latency and peak bandwidth requirements.
[0133] Example 7: During an e-commerce festival, the storage device cluster of an e-commerce platform needs to process a large amount of user access and transaction data. The system can detect that the access volume and update frequency of transaction data have increased significantly, and automatically adjust its priority weight to ensure the real-time synchronization of transaction data. At the same time, it is monitored that the network bandwidth between data centers fluctuates during peak periods. The synchronization strategy is adjusted through a prediction model to segment the reference sub-data of transaction data into smaller pieces to adapt to the network state, while reducing the synchronization rate of non-hot data to ensure the efficient utilization of network resources.
[0134] In addition, based on historical data and user behavior analysis, it is predicted that the information of hot-selling products during the festival may become hot data, and these data are pre-loaded into the cache or synchronized to the target data center in advance to ensure that users can quickly obtain data when accessing hot-selling products, greatly improving the user experience and reducing the data synchronization pressure during peak periods.
[0135] Through the above embodiments of the present application, by introducing mechanisms such as dynamic perception of business requirements, adaptive scheduling of network status, and data pre-reading prediction, intelligent optimization of data synchronization in a dual-active storage system is achieved. It not only improves the processing efficiency of hot data but also ensures the smooth synchronization of non-hot data, providing a stronger guarantee for the high availability and business continuity of the system.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0137] The embodiments of the present application also provide a data transmission device. Figure 6 It is a structural block diagram of an optional data transmission device according to an embodiment of the present application, as Figure 6 shown. The device includes:
[0138] A data processing module 602, configured to perform operations on the data in the first storage device in a storage device cluster to obtain target data after being operated. Among them, the storage device cluster includes a first storage device and at least one second storage device;
[0139] A storage device status determination module 604, configured to determine a target storage device for receiving the target data from at least one second storage device, and obtain reference network bandwidth information between the first storage device and the target storage device, as well as the first load status of the first storage device and the second load status of the second storage device;
[0140] A transmission rate prediction module 606, configured to predict the data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load status, and the second load status;
[0141] A data transmission module 608, configured to transmit the target data to the target storage device according to the data synchronization transmission rate.
[0142] Optionally, the above transmission rate prediction module 606 is further configured to: predict the target network bandwidth information between the first storage device and the target storage device based on the reference network bandwidth information; determine the target device resources corresponding to the target data according to the data access frequency of the target data and the first load status and the second load status, where the target device resources are used to indicate the computing resources and storage resources for the target data transmission; and determine the data synchronization transmission rate by using the target network bandwidth information and the target device resources.
[0143] Optionally, the above transmission rate prediction module 606 is further configured to: calculate a target resource weight matching the target data according to the data access frequency, service relevance of the target data, and the predicted transmission time of the target synchronization transmission task corresponding to the target data, where the predicted transmission time is used to indicate the expected time to complete the synchronization transmission task, and the service relevance is used to indicate the degree of association between the data and the service process; calculate the reference resource weights matching the respective reference synchronization transmission tasks according to the data access frequency, predicted transmission time, and service relevance of the reference data transmitted by the reference synchronization transmission tasks respectively; allocate the total device resources of the first storage device and the target storage device according to the target resource weight, at least one reference resource weight, the first load state, and the second load state to obtain the target device resources.
[0144] Optionally, the above storage device state determination module 604 is further configured to: calculate the priority attributes corresponding to the respective candidate data by weighted calculation of the data access frequency and service relevance of at least one candidate data, where the at least one candidate data includes the target data; sort the candidate data according to the magnitudes of the respective corresponding priority attributes to obtain a candidate data sequence, where the candidate data with a larger priority attribute is located at a more forward position in the candidate data sequence; determine the candidate data other than the target data that are in the first N positions in the candidate data sequence as the reference data, where N is an integer greater than 0.
[0145] Optionally, the above storage device state determination module 604 is further configured to: split the reference data into multiple reference sub-data in the case where the data volume of the reference data is greater than a preset data threshold; generate a reference synchronization transmission task for each reference sub-data.
[0146] Optionally, the above transmission rate prediction module 606 is further configured to: calculate a first ratio of the data volume of the target data to the total data volume, where the total data volume is used to indicate the sum of the data volume of the target data and the data volume of the reference data; determine the target sub-network bandwidth information corresponding to the target data as the product of the target network bandwidth information and the first ratio; determine the data synchronization transmission rate by using the target sub-network bandwidth information and the target device resources.
[0147] Optionally, the above transmission rate prediction module 606 is further configured to: calculate the mean value of at least one first network information in the reference network bandwidth information set to obtain a first network bandwidth information mean value, where the first network information is used to indicate the historical network bandwidth information between the first storage device and the target storage device; perform weighted summation of the second network bandwidth information in the reference network information set and the first network bandwidth information mean value to obtain the target network bandwidth information, where the second network bandwidth information is used to indicate the current network bandwidth information between the first storage device and the target storage device.
[0148] For the description of the features in the corresponding embodiments of the data transmission device, reference may be made to the relevant description in the corresponding embodiments of the data transmission method, which will not be elaborated herein one by one.
[0149] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the data transmission method.
[0150] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the data transmission method when running.
[0151] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, etc., various media that can store computer programs.
[0152] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the data transmission method.
[0153] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the data transmission method.
[0154] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0155] The above has introduced in detail a data transmission method, device, storage medium, and electronic device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data transmission method, characterized in that: Including: Performing an operation on the data in the first storage device in the storage device cluster to obtain the target data after the operation, where the storage device cluster includes the first storage device and at least one second storage device; Determining a target storage device for receiving the target data from the at least one second storage device, and obtaining the reference network bandwidth information between the first storage device and the target storage device, as well as the first load state of the first storage device and the second load state of the second storage device; Based on the reference network bandwidth information, the first load state and the second load state, predicting the data synchronization transmission rate between the first storage device and the target storage device; Transmitting the target data to the target storage device according to the data synchronization transmission rate.
2. The method according to claim 1, characterized in that: The predicting the data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load state and the second load state includes: Predicting the target network bandwidth information between the first storage device and the target storage device based on the reference network bandwidth information; Determining the target device resources corresponding to the target data according to the data access frequency of the target data, the first load state and the second load state, where the target device resources are used to indicate the computing resources and storage resources for the target data transmission; Using the target network bandwidth information and the target device resources to determine the data synchronization transmission rate.
3. The method according to claim 2, characterized in that: The determining the target device resources corresponding to the target data according to the data access frequency of the target data, the first load state and the second load state includes: Calculating the target resource weight matching the target data according to the data access frequency, service correlation degree of the target data and the predicted transmission time of the target synchronization transmission task corresponding to the target data, where the predicted transmission time is used to indicate the expected time to complete the synchronization transmission task, and the service correlation degree is used to indicate the correlation degree between the data and the service process; Calculating the reference resource weights matching the respective reference synchronization transmission tasks according to the data access frequency, predicted transmission time and service correlation degree of the reference data transmitted by the reference synchronization transmission tasks respectively; Allocating the total device resources of the first storage device and the target storage device according to the target resource weight, at least one of the reference resource weights, the first load state and the second load state to obtain the target device resources.
4. The method according to claim 2, characterized in that: Before the determining the target device resources corresponding to the target data according to the data access frequency of the target data, the first load state and the second load state, including: Calculating the weighted data access frequency and business relevance of each of at least one candidate data to obtain the priority attribute corresponding to each of the candidate data, where the at least one candidate data includes the target data; Sorting the candidate data according to the magnitude of the priority attribute corresponding to each of them to obtain a candidate data sequence, where the candidate data with a larger priority attribute is located at a more forward position in the candidate data sequence; Determining the first N candidate data in the candidate data sequence except the target data as reference data, where N is an integer greater than 0.
5. The method according to claim 4, wherein: The determining the first N candidate data in the candidate data sequence except the target data as reference data includes: When the amount of the reference data is greater than a preset data threshold, splitting the reference data into multiple reference sub-data; Generating a reference synchronous transmission task for each of the reference sub-data.
6. The method according to claim 2, wherein: The determining the data synchronous transmission rate by using the target network bandwidth information and the target device resources includes: Calculating a first ratio of the data amount of the target data to the total data amount, where the total data amount is used to indicate the sum of the data amount of the target data and the data amount of the reference data; Determining the target sub-network bandwidth information corresponding to the target data as the product of the target network bandwidth information and the first ratio; Determining the data synchronous transmission rate by using the target sub-network bandwidth information and the target device resources.
7. The method according to claim 2, wherein: The predicting the target network bandwidth information between the first storage device and the target storage device based on the reference network bandwidth information includes: Calculating the mean value of at least one first network information in the reference network bandwidth information set to obtain a first network bandwidth information mean value, where the first network information is used to indicate the historical network bandwidth information between the first storage device and the target storage device; Performing a weighted sum of the second network bandwidth information in the reference network information set and the first network bandwidth information mean value to obtain the target network bandwidth information, where the second network bandwidth information is used to indicate the current network bandwidth information between the first storage device and the target storage device.
8. A data transmission device, wherein: It includes: A data processing module, configured to perform operations on the data in the first storage device in a storage device cluster to obtain the target data after being operated, where the storage device cluster includes the first storage device and at least one second storage device; A storage device state determination module, configured to determine a target storage device for receiving the target data from the at least one second storage device, and obtain the reference network bandwidth information between the first storage device and the target storage device, as well as the first load state of the first storage device and the second load state of the second storage device; A transmission rate prediction module, configured to predict a data synchronization transmission rate between the first storage device and the target storage device based on the reference network bandwidth information, the first load status, and the second load status; A data transmission module, configured to transmit the target data to the target storage device according to the data synchronization transmission rate.
9. An electronic device, characterized in that it includes: a memory, configured to store a computer program; a processor, configured to implement the steps of the data transmission method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that a computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the data transmission method according to any one of claims 1 to 7.