A data maintenance method and application system for server leasing

By analyzing the historical business traffic data of server tenants, determining the business data distribution and contention intensity, obtaining SLA isolation requirements, and dynamically adjusting resource isolation strategies, we solve the problems of I/O latency and cache contention in multi-tenant environments, and achieve efficient data isolation and performance assurance.

CN120415926BActive Publication Date: 2025-10-17NANCHANG HOME TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510919564.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

When servers are leased to multiple tenants, existing technologies cannot effectively solve the problems of data isolation and performance independence between tenants. In particular, when business peaks overlap, this can cause a surge in I/O latency and intensify cache contention, violating SLA performance commitments.

Method used

By analyzing tenants' historical business traffic data, determining the business data distribution and the contention intensity of tenants' overlapping businesses, and obtaining SLA isolation requirements, the resource isolation strategy is dynamically adjusted based on this information to avoid delay avalanches caused by excessive isolation and queue accumulation, and reduce the probability of secondary contention.

Benefits of technology

It achieves effective data isolation and performance assurance in a multi-tenant environment, reduces bandwidth usage and latency jitter caused by full migration, and improves resource utilization efficiency and SLA satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120415926B_ABST
    Figure CN120415926B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data maintenance, in particular to a data maintenance method and application system for server leasing. The method comprises the following steps: obtaining historical business traffic data of tenants; analyzing the historical business traffic data to determine business data distribution; determining tenant overlapping businesses and contention intensity of the tenant overlapping businesses according to the business data distribution; obtaining SLA isolation requirements of the tenants; and determining a resource isolation strategy according to the tenant overlapping businesses and the contention intensity based on the SLA isolation requirements. According to the business data distribution, the tenant overlapping businesses and the contention intensity of the tenant overlapping businesses are determined, so that elastic degradation negotiation can be started in advance, and delay avalanche caused by queue accumulation can be avoided. High-contention businesses are avoided from being migrated to nodes with serious resource fragmentation, the probability of secondary contention is reduced, and bandwidth occupation and delay jitter caused by full migration are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data maintenance, in particular to a data maintenance method and application system for server leasing. BACKGROUND

[0002] In the field of server leasing, there may be a case where a server is leased to multiple tenants, which inevitably leads to multiple tenants sharing physical resources. At this time, virtualization technology is needed to ensure data isolation and performance independence between tenants.

[0003] However, the constraints of existing technologies for leasing a server to multiple tenants are relatively single. When the peak periods of multiple tenants' businesses overlap in time or resource usage mode, the competition for shared resources will cause problems such as I / O delay surge and cache contention intensification, thereby violating the performance commitment in the tenant service level agreement. SUMMARY

[0004] The present application provides a data maintenance method and application system for server leasing to solve the above problems.

[0005] In a first aspect, the present application provides a data maintenance method for server leasing, which comprises:

[0006] Obtaining historical business traffic data of a tenant; analyzing the historical business traffic data to determine the business data distribution;

[0007] According to the business data distribution, determining the tenant overlapping business and the contention intensity of the tenant overlapping business;

[0008] Obtaining the SLA isolation requirement of the tenant; based on the SLA isolation requirement, determining the resource isolation strategy according to the tenant overlapping business and the contention intensity.

[0009] According to the present application, the historical business traffic data of the tenant is obtained to avoid misjudging short-term traffic peaks as normal business cycles. The historical business traffic data is analyzed to determine the business data distribution, avoiding excessive isolation decisions due to simple time overlap. According to the business data distribution, the tenant overlapping business and the contention intensity of the tenant overlapping business are determined to start elastic degradation negotiation in advance, while avoiding delay avalanche caused by queue accumulation. The SLA isolation requirement of the tenant is obtained to convert natural language clauses into machine executable constraint rules, providing clear basis for dynamic adjustment. Based on the SLA isolation requirement, the resource isolation strategy is determined according to the tenant overlapping business and the contention intensity, avoiding migrating high contention business to nodes with serious resource fragmentation, reducing the probability of secondary contention, and reducing bandwidth occupation and delay jitter caused by full migration.

[0010] Optionally, the analyzing the historical traffic data, determining the traffic data distribution, comprises:

[0011] analyzing the historical traffic data, determining whether the historical traffic data contains burst traffic peaks;

[0012] if the historical traffic data contains burst traffic peaks, analyzing the burst traffic peaks, determining peak operation traffic, and marking the peak operation traffic as an abnormal event;

[0013] analyzing the abnormal event, determining abnormal traffic data, and removing the abnormal traffic data from the historical traffic data to obtain normal traffic data;

[0014] analyzing the normal traffic data by Fourier transform, determining a periodic waveform;

[0015] adopting a sliding window filtering algorithm to remove noise from the periodic waveform to obtain a reference waveform;

[0016] analyzing the reference waveform, determining a business cycle feature, and taking the business cycle feature as the traffic data distribution.

[0017] According to the scheme, the historical traffic data is analyzed to determine whether the historical traffic data contains burst traffic peaks, so as to avoid misjudgment of a business cycle due to burst traffic. If the historical traffic data contains burst traffic peaks, the burst traffic peaks are analyzed to determine peak operation traffic, and the peak operation traffic is marked as an abnormal event to provide an accurate range for data removal. The abnormal event is analyzed to determine abnormal traffic data, and the abnormal traffic data is removed from the historical traffic data to obtain normal traffic data, so as to ensure that the cycle analysis is based on a real and stable business load mode. The normal traffic data is analyzed by Fourier transform to determine a periodic waveform, so as to reveal inherent time regularity of tenant business. The sliding window filtering algorithm is adopted to remove noise from the periodic waveform to obtain a reference waveform, so as to avoid calculation deviation of cycle parameters caused by noise interference. The reference waveform is analyzed to determine a business cycle feature, and the business cycle feature is taken as the traffic data distribution, so as to avoid underestimation or over-reservation of resource contention intensity caused by misjudgment of the cycle.

[0018] Optionally, the determining the tenant overlapping business according to the traffic data distribution comprises:

[0019] analyzing the traffic data distribution, determining a business use peak period of each tenant;

[0020] determining a peak business distribution rule of each tenant based on the business use peak period;

[0021] determining the tenant overlapping business according to the peak business distribution rule.

[0022] By the scheme, the service data distribution is analyzed, the service peak period of each tenant is determined, interference of abnormal events on identification of the peak period is avoided, and it is ensured that the extracted service peak period only reflects the real and stable service load characteristics of the tenant. Based on the service peak period, the peak service distribution law of each tenant is determined, misassociation of non-periodic tenants is avoided, and a basis is provided for resource coupling analysis. According to the peak service distribution law, the tenant overlapping service is determined, and resource fragmentation and secondary contention caused by full migration are avoided.

[0023] Optionally, the determining the tenant overlapping service according to the peak service distribution law comprises:

[0024] determining a peak service type according to the peak service distribution law;

[0025] analyzing the service data distribution based on the peak service type, and determining whether any two peak service types are in a use time period overlap stage;

[0026] if yes, determining whether the any two peak service types in the use time period overlap stage belong to the same physical server according to the service data distribution;

[0027] if yes, determining a service overlap degree according to the use time period of the peak service;

[0028] comparing the service overlap degree with a preset overlap threshold, and if the service overlap degree is higher than the preset overlap threshold, determining that it is a tenant overlapping service.

[0029] By the scheme, the peak service type is determined according to the peak service distribution law, misjudgment of overlap detection caused by ambiguous service type is avoided. Based on the peak service type, the service data distribution is analyzed, it is determined whether any two peak service types are in a use time period overlap stage, invalid detection of non-overlapping time period is excluded, roughness of relying on simple time window division is overcome, and only the actual overlapping time period is focused on to reduce calculation redundancy. If yes, it is determined whether the any two peak service types in the use time period overlap stage belong to the same physical server according to the service data distribution, which helps to exclude hardware irrelevant interference between cross-server tenants and avoid the problem that an isolation strategy at a virtual machine level cannot cover underlying resource contention. If yes, the service overlap degree is determined according to the use time period of the peak service, which helps to quantify the actual contention intensity of shared resources of two tenants and avoid one-sidedness of relying on time or a single index. The service overlap degree is compared with the preset overlap threshold, and if the service overlap degree is higher than the preset overlap threshold, it is determined that it is a tenant overlapping service, which helps to adapt to nonlinear changes of resource contention intensity and avoid SLA violation caused by response lag.

[0030] Optionally, the historical service traffic data further comprises I / O data and SSD garbage collection data, and the determination of the contention intensity of the tenant overlapping services according to the service data distribution comprises:

[0031] analyzing the I / O data and the SSD garbage collection data to determine I / O queue depth and SSD garbage collection status;

[0032] determining an I / O delay growth rate according to the I / O queue depth and the SSD garbage collection status;

[0033] determining the contention intensity of the tenant overlapping services according to the service data distribution and the I / O delay growth rate.

[0034] Through the scheme, the I / O data and the SSD garbage collection data are analyzed to determine the I / O queue depth and the SSD garbage collection status, which intuitively reflects the instantaneous load pressure of the physical server storage I / O resource and clearly determines the influence period of the storage device self-maintenance operation on the I / O performance, thereby improving the accuracy of contention intensity quantification. According to the I / O queue depth and the SSD garbage collection status, the I / O delay growth rate is determined, which helps to eliminate the defects of dynamic interference response lag and provide early warning of potential SLA violation risks. According to the service data distribution and the I / O delay growth rate, the contention intensity of the tenant overlapping services is determined, which helps to eliminate the problem of fragmented resource utilization inefficiency, provides comparable basis for migration decisions, and reduces the secondary contention risk after migration.

[0035] Optionally, the determination of the resource isolation strategy according to the tenant overlapping services and the contention intensity based on the SLA isolation requirement comprises:

[0036] determining the basic performance between tenants according to the SLA isolation requirement;

[0037] determining a performance interference coefficient of the tenant overlapping services according to the contention intensity;

[0038] analyzing the service data distribution to determine the dynamic change of the tenant overlapping services;

[0039] determining the resource isolation strategy according to the dynamic change, the performance interference coefficient, and the basic performance.

[0040] According to the SLA isolation requirement, the base performance between tenants is determined, so as to avoid default caused by insufficient resource allocation, and prevent resource over-allocation or under-allocation caused by ambiguous indicators. According to the contention intensity, the performance interference coefficient of the tenant overlapping service is determined, the potential threat level of the current resource contention to the performance is reflected, and the limitation of the virtual machine level isolation strategy is avoided. The business data distribution is analyzed, the dynamic change of the tenant overlapping service is determined, the real business demand of the tenant is ensured, and the strategy misjudgment caused by accidental events is avoided. According to the dynamic change, the performance interference coefficient and the base performance, the resource isolation strategy is determined, the data amount of migration is reduced, and the secondary contention caused by full migration is avoided.

[0041] Optionally, the performance interference coefficient of the tenant overlapping service is determined according to the contention intensity, and the performance interference coefficient of the tenant overlapping service is determined according to the contention intensity.

[0042] The data access dependency relationship is determined based on the contention intensity.

[0043] The cross-tenant shared cache line and / or contention lock resource is determined according to the data access dependency relationship.

[0044] The hot spot interference area is determined according to the cross-tenant shared cache line and / or contention lock resource.

[0045] The performance interference coefficient of the tenant overlapping service is determined based on the business data distribution and the hot spot interference area.

[0046] According to the contention intensity, the business data distribution is analyzed, the data access dependency relationship is determined, and the resource allocation strategy is deviated from the real demand caused by abnormal peaks. According to the data access dependency relationship, the cross-tenant shared cache line and / or contention lock resource is determined, which helps to make up for the defects caused by only focusing on virtual resources. According to the cross-tenant shared cache line and / or contention lock resource, the hot spot interference area is determined, which helps to avoid that the virtual machine abstraction layer hides the real resource conflict position, ensures that the isolation strategy focuses on the conflict point, and avoids the resource fragmentation problem caused by full migration. Based on the business data distribution, the hot spot interference area is analyzed, and the performance interference coefficient of the tenant overlapping service is determined, which helps to realize the fine quantization of the contention intensity and provide accurate input for the migration decision.

[0047] Optionally, the resource isolation strategy is determined according to the dynamic change, the performance interference coefficient and the base performance, and the performance interference coefficient is compared with a preset interference intensity threshold value.

[0048] If the performance interference coefficient is higher than the preset interference intensity threshold value, the adjustment plan of the base performance is determined according to the dynamic change.

[0049] Analyze the SLA isolation requirements and determine whether the tenant accepts the adjustment plan;

[0050] If accepted, the adjustment plan is determined as the resource isolation policy.

[0051] This solution compares the performance interference coefficient with a preset interference intensity threshold. If the performance interference coefficient exceeds the preset interference intensity threshold, a basic performance adjustment plan is determined based on dynamic changes. This helps eliminate the problem of static thresholds being unable to adapt to nonlinear latency growth, avoids migrating highly contention services to saturated nodes, and reduces the probability of secondary contention. The SLA isolation requirements are analyzed to determine whether the tenant accepts the adjustment plan, enabling differentiated isolation policy negotiation. If accepted, the adjustment plan is determined as a resource isolation policy, helping to eliminate the problem of inefficient fragmented resource utilization and reduce secondary resource conflicts caused by full migration.

[0052] Optionally, the method further includes:

[0053] If not accepted, determining the resource fragment distribution state according to the business data distribution;

[0054] Determining a target node that minimizes the amount of data migration based on the resource fragmentation distribution state;

[0055] A double-write mechanism is used to migrate the target node.

[0056] If this solution is not accepted, the resource fragmentation distribution status is determined based on the business data distribution, which helps avoid the waste of fragmented resources caused by full data migration or random node selection, and reduces the amount of redundant data migration. Based on the resource fragmentation distribution status, the target node that minimizes the amount of data migration is determined to avoid the risk of secondary contention caused by unoptimized paths during migration, while ensuring that the target node resources can support the tenant's SLA requirements. Using a dual-write mechanism to migrate the target node helps eliminate the risk of business interruption caused by single-point write or full copy during migration, achieves seamless switching and avoids data loss, while reducing the additional impact of migration on I / O performance.

[0057] In a second aspect, the present application provides a data maintenance application system for server leasing, the system comprising:

[0058] A data analysis module is used to obtain the tenant's historical business traffic data; analyze the historical business traffic data to determine the business data distribution;

[0059] An overlap analysis module, configured to determine overlapping services of tenants and contention intensity of the overlapping services of the tenants according to the distribution of the service data;

[0060] A policy determination module is configured to acquire SLA isolation requirements of the tenant, and determine a resource isolation policy according to the contention intensity based on the SLA isolation requirements.

[0061] Optionally, when the data analysis module analyzes the historical traffic data and determines the traffic data distribution, the data analysis module is configured to:

[0062] analyze the historical traffic data to determine whether the historical traffic data contains burst traffic peaks;

[0063] if the historical traffic data contains burst traffic peaks, analyze the burst traffic peaks to determine peak operation traffic, and mark the peak operation traffic as an abnormal event;

[0064] analyze the abnormal event to determine abnormal traffic data, and remove the abnormal traffic data from the historical traffic data to obtain normal traffic data;

[0065] analyze the normal traffic data by Fourier transform to determine a periodic waveform;

[0066] perform noise elimination on the periodic waveform by using a sliding window filtering algorithm to obtain a reference waveform;

[0067] analyze the reference waveform to determine a traffic periodicity feature, and use the traffic periodicity feature as the traffic data distribution.

[0068] Optionally, when the overlap analysis module determines the tenant overlap service according to the traffic data distribution, the overlap analysis module is configured to:

[0069] analyze the traffic data distribution to determine a service use peak period of each tenant;

[0070] determine a peak service distribution rule of each tenant based on the service use peak period;

[0071] determine the tenant overlap service according to the peak service distribution rule.

[0072] Optionally, when the overlap analysis module determines the tenant overlap service according to the peak service distribution rule, the overlap analysis module is configured to:

[0073] determine a peak service type according to the peak service distribution rule;

[0074] analyze the traffic data distribution based on the peak service type to determine whether any two peak service types are in a use time period overlap stage;

[0075] if any two peak service types are in the use time period overlap stage, determine whether the any two peak service types belong to a same physical server according to the traffic data distribution;

[0076] If yes, according to the use period of peak service, the service overlap degree is determined;

[0077] The service overlap degree is compared with a preset overlap threshold value, and if the service overlap degree is higher than the preset overlap threshold value, the tenant overlap service is determined.

[0078] Optionally, the historical service traffic data further includes I / O data and SSD garbage collection data, and when determining the contention intensity of the tenant overlap service of the overlap analysis module according to the service data distribution, the I / O data and the SSD garbage collection data are used to:

[0079] The I / O data and the SSD garbage collection data are analyzed to determine I / O queue depth and SSD garbage collection state;

[0080] According to the I / O queue depth and the SSD garbage collection state, an I / O delay growth rate is determined;

[0081] According to the service data distribution and the I / O delay growth rate, the contention intensity of the tenant overlap service is determined.

[0082] Optionally, when the policy determination module determines the resource isolation strategy based on the SLA isolation requirement according to the tenant overlap service and the contention intensity, the policy determination module is used to:

[0083] According to the SLA isolation requirement, a basic performance between tenants is determined;

[0084] According to the contention intensity, a performance interference coefficient of the tenant overlap service is determined;

[0085] The service data distribution is analyzed to determine a dynamic change of the tenant overlap service;

[0086] According to the dynamic change, the performance interference coefficient, and the basic performance, a resource isolation strategy is determined.

[0087] Optionally, when the policy determination module determines the performance interference coefficient of the tenant overlap service according to the contention intensity, the policy determination module is used to:

[0088] Based on the contention intensity, the service data distribution is analyzed to determine a data access dependency relationship;

[0089] According to the data access dependency relationship, a cross-tenant shared cache line and / or contention lock resource is determined;

[0090] According to the cross-tenant shared cache line and / or contention lock resource, a hot spot interference area is determined;

[0091] Based on the service data distribution, the hotspot interference area is analyzed, and a performance interference coefficient of the tenant overlapping service is determined.

[0092] Optionally, when the policy determination module determines the resource isolation policy according to the dynamic change, the performance interference coefficient and the basic performance, the method further comprises:

[0093] The performance interference coefficient is compared with a preset interference intensity threshold value, and if the performance interference coefficient is higher than the preset interference intensity threshold value, an adjustment plan of the basic performance is determined according to the dynamic change;

[0094] The SLA isolation requirement is analyzed to determine whether the adjustment plan is accepted by the tenant;

[0095] If the adjustment plan is accepted, the adjustment plan is determined as the resource isolation policy.

[0096] Optionally, the data maintenance application system of the server lease further comprises a node migration module, configured to:

[0097] If the adjustment plan is not accepted, a resource fragmentation distribution state is determined according to the service data distribution;

[0098] A target node minimizing data migration amount is determined according to the resource fragmentation distribution state;

[0099] The target node is migrated by using a double-write mechanism. BRIEF DESCRIPTION OF DRAWINGS

[0100] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative labor on the basis of these drawings.

[0101] Figure 1 An application scenario schematic diagram provided by an embodiment of the present application;

[0102] Figure 2 A flowchart of a data maintenance method of a server lease provided by an embodiment of the present application;

[0103] Figure 3 A data maintenance application system structure schematic diagram of a server lease provided by an embodiment of the present application. DETAILED DESCRIPTION

[0104] In order to make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0105] In addition, the term "and / or" in the present application is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects unless otherwise specified.

[0106] The embodiments of the present application will be described in further detail below with reference to the drawings of the specification.

[0107] The prior art has a single constraint on a server leasing to multiple tenants. When the business peak periods of multiple tenants overlap in time or resource usage mode, the competition for shared resources can cause I / O delay to surge, cache contention to intensify, and other problems, thereby violating the performance commitment in the tenant service level agreement.

[0108] Based on this, the present application provides a data maintenance method and application system for server leasing. The historical business traffic data of tenants is obtained to avoid misjudging short-term traffic peaks as normal business cycles. The historical business traffic data is analyzed to determine the business data distribution, thereby avoiding excessive isolation decisions caused by pure time overlap. According to the business data distribution, the tenant overlapping business and the contention intensity of the tenant overlapping business are determined, and the elastic degradation negotiation is started in advance, while avoiding the delay avalanche caused by queue accumulation. The SLA isolation requirements of tenants are obtained, the natural language clauses are converted into machine executable constraint rules, and explicit basis is provided for dynamic adjustment. Based on the SLA isolation requirements, the resource isolation strategy is determined according to the tenant overlapping business and the contention intensity, thereby avoiding migrating high contention business to nodes with serious resource fragmentation, reducing the probability of secondary contention, and reducing the bandwidth occupation and delay jitter caused by full migration.

[0109] Figure 1 An application scenario provided by the present application is shown in the figure. The method provided by the present application is applied when maintaining the data of server leasing.

[0110] Specifically, the method provided in the application is applied to any server, the server interacts with a tenant device, historical business traffic data of the tenant is obtained and analyzed through the tenant device, and the distribution of the business data is determined. According to the distribution of the business data, the tenant overlapping business and the contention strength of the tenant overlapping business are determined, the elastic degradation negotiation is started in advance, and the delay avalanche caused by queue accumulation is avoided. Based on the SLA isolation requirement of the tenant, the resource isolation strategy is determined according to the tenant overlapping business and the contention strength, the high contention business is avoided to be migrated to the node with serious resource fragmentation, the secondary contention probability is reduced, and the bandwidth occupation and delay jitter caused by full migration are reduced. The specific implementation mode can refer to the following embodiments.

[0111] Figure 2 The flowchart of the data maintenance method of the server lease provided in an embodiment of the application, the method of the embodiment can be applied to the server in the above scene. As shown in the figure, the method comprises the following steps. Figure 2

[0112] S201, obtaining historical business traffic data of a tenant; analyzing the historical business traffic data to determine the distribution of the business data;

[0113] The historical business traffic data can be time-series resource usage records reflecting the business load characteristics generated by the tenant in the past running period, including I / O data and SSD garbage collection data. The distribution of the business data can be a set of distribution characteristics of the tenant business load in a multi-dimensional resource space.

[0114] Specifically, the virtual machine monitoring module, the Hypervisor layer and the tenant device interact to obtain the historical business traffic data of the tenant. Then, based on the historical business traffic data, the business traffic main frequency component is extracted through Fourier transform to determine the periodic peak period; then, the resource occupation proportion of CPU, storage I / O and network bandwidth in the peak period is calculated and normalized to a three-dimensional vector; subsequently, based on the virtual machine resource allocation information recorded by the Hypervisor layer and the physical hardware topology discovery mechanism obtained through ACPI table analysis and device tree traversal, a virtual machine-physical resource mapping table of virtual machines and underlying hardware resources is established, the underlying resources shared by the tenant are counted, and the distribution of the business data is determined.

[0115] S202, determining the tenant overlapping business and the contention strength of the tenant overlapping business according to the distribution of the business data;

[0116] The tenant overlapping business can be a resource contention business generated by two or more tenants. The contention strength can be a quantitative index of the competition degree of shared physical resources.

[0117] ​Specifically, according to the service data distribution, the period component of each tenant baseline model is compared to identify the time window overlap interval; then, a time series alignment algorithm is used to quantify the time window overlap degree according to the time window overlap interval, and the resource demand conflict degree is calculated; further, a percentile method in statistics is used in combination with a lag back difference mechanism in control theory to set a conflict degree threshold, and if the conflict degree threshold is exceeded, it is determined that there is tenant overlapping service;

[0118] Based on the resource demand conflict degree and the service data distribution, the initial state is calculated; if it is detected that the initial state of the target physical machine SSD is in the garbage collection period, the storage contention intensity weighting coefficient is increased; the spin lock waiting time is obtained through the performance counter, and when the lock contention increases, the CPU contention intensity is improved; and thus the final tenant overlapping service and the contention intensity of the tenant overlapping service are determined.

[0119] S203, obtain the SLA isolation requirement of the tenant; based on the SLA isolation requirement, determine the resource isolation strategy according to the tenant overlapping service and the contention intensity.

[0120] The SLA isolation requirement can be a set of performance isolation constraint conditions agreed by the tenant in the service level agreement.

[0121] The resource isolation strategy can be a dynamic decision scheme generated based on the contention intensity and the SLA isolation requirement.

[0122] Specifically, based on the domain-specific grammar rules in natural language processing, regular expressions are used to extract performance constraint conditions, elasticity clauses and other parameters in the tenant contract to construct the SLA isolation requirement. According to the memory dirty page rate of the virtual machine to be migrated, the storage volume capacity, and the network bandwidth margin, the migration time estimation and performance impact score of each candidate target node are calculated; then, the node with continuous physical memory block and balanced SSD remaining erase times is preferentially selected; an improved greedy algorithm is used to select the scheme that minimizes the increment of target node resource fragmentation degree each time; further, the migration path is detected for conflict to avoid forming new high contention node pairs; finally, when it is predicted that the basic performance indicators must be temporarily violated, a negotiation request containing alternative schemes is sent to the tenant console, and the final resource isolation strategy is determined according to the real-time feedback of the tenant.

[0123] By the scheme, the historical business traffic data of the tenant is acquired, and short-term traffic spikes are avoided from being misjudged as normal business periods. The historical business traffic data is analyzed to determine the business data distribution, and over-isolation decisions caused by pure time overlap are avoided. According to the business data distribution, the tenant overlapping business and the contention strength of the tenant overlapping business are determined, the elastic degradation negotiation is started in advance, and delay avalanches caused by queue accumulation are avoided. The SLA isolation requirements of the tenant are acquired, the natural language clauses are converted into machine executable constraint rules, and explicit basis is provided for dynamic adjustment. Based on the SLA isolation requirements, the resource isolation strategy is determined according to the tenant overlapping business and the contention strength, high contention business is avoided from being migrated to a node with serious resource fragmentation, the probability of secondary contention is reduced, and bandwidth occupation and delay jitter caused by full migration are reduced.

[0124] In some embodiments, the historical business traffic data is analyzed to determine whether it contains burst traffic spikes; if it does, the burst traffic spikes are analyzed to determine spike operation business, and the spike operation business is marked as an abnormal event; the abnormal event is analyzed to determine abnormal traffic data, and the abnormal traffic data is excluded from the historical business traffic data to obtain normal business traffic data; the normal business traffic data is analyzed by Fourier transform to determine a periodic waveform; a sliding window filtering algorithm is used to remove noise from the periodic waveform to obtain a reference waveform; the reference waveform is analyzed to determine business period characteristics, and the business period characteristics are used as the business data distribution.

[0125] The burst traffic spike can be a transient traffic surge that significantly deviates from the historical mean in the historical business traffic data. The spike operation business can be a business process or task associated with the burst traffic spike. The abnormal event can be a business operation instance caused by the burst traffic spike. The abnormal traffic data can be a raw traffic data segment containing the time window corresponding to the abnormal event. The normal business traffic data can be a traffic data set filled by interpolation after excluding the abnormal traffic data. The Fourier transform can be a mathematical method for converting time domain traffic data into frequency domain components. The periodic waveform can be a time domain waveform reconstructed from the main frequency components. The sliding window filtering algorithm can be a median filtering method based on window width. The reference waveform can be a smooth periodic waveform. The business period characteristics can be a set of periodic parameters such as period length, peak interval, amplitude range, etc.

[0126] Specifically, a sliding window statistics is performed on historical traffic data, and a standard deviation of traffic values in each time window is calculated; then, a spike determination threshold is set based on the traffic standard deviation of the sliding window statistics, and subsequently, a time window with a traffic value exceeding the spike determination threshold is marked as a burst traffic spike. Furthermore, a process ID and a request type log of the burst traffic spike are extracted; then, non-periodic tasks in a tenant service deployment list are matched; if the spike area operation matches the list, the corresponding service is determined as a spike operation service, and is marked as an abnormal event. The traffic time sequence segment corresponding to the abnormal event is removed from the original data; subsequently, linear interpolation is performed on the missing data segment to fill in the missing data segment, and normal service traffic data is generated. Fourier transform is applied to the normal service traffic data to generate a frequency spectrum; then, based on the Fourier transform frequency spectrum energy cumulative distribution characteristics, a preset proportion is set, and a main frequency component with an energy proportion exceeding the preset proportion in the frequency spectrum is identified; subsequently, a periodic waveform is reconstructed based on the main frequency component. A sliding window filtering algorithm is used, and the window width is defined as an integer multiple of the periodic component; then, the traffic median value is calculated in each window, and the original value of the window center point is replaced; subsequently, the points are slid one by one until the entire waveform is covered, and a smoothed reference waveform is generated. Furthermore, the peak interval and amplitude variation law of the reference waveform are analyzed; if the standard deviation of the adjacent peak time intervals is less than the spike determination threshold, it is determined as a stable periodic characteristic; finally, the periodic parameters such as the period length, peak bandwidth, and valley baseline are output as the service data distribution.

[0127] By the scheme, historical service traffic data is analyzed to determine whether it contains a burst traffic spike, avoiding misjudgment of service periods caused by burst traffic. If it contains, the burst traffic spike is analyzed to determine a spike operation service, and the spike operation service is marked as an abnormal event, providing an accurate range for data elimination. The abnormal event is analyzed to determine abnormal traffic data, and the abnormal traffic data is eliminated from the historical service traffic data to obtain normal service traffic data, ensuring that the periodic analysis is based on real and stable service load patterns. Fourier transform is used to analyze the normal service traffic data to determine a periodic waveform, revealing the inherent time regularity of tenant services. A sliding window filtering algorithm is used to eliminate noise from the periodic waveform to obtain a reference waveform, avoiding calculation deviation of periodic parameters caused by noise interference. The reference waveform is analyzed to determine a service period characteristic, and the service period characteristic is output as the service data distribution, avoiding underestimation or over-reservation of resource contention intensity caused by period misjudgment.

[0128] In some embodiments, the service data distribution is analyzed to determine a service use peak period of each tenant; based on the service use peak period, a peak service distribution law of each tenant is determined; and according to the peak service distribution law, a tenant overlapping service is determined.

[0129] The service usage peak period can be a continuous time section exceeding the minimum steady-state flow value of the benchmark waveform and reaching the peak bandwidth. The peak service distribution law can be a multi-dimensional feature set formed by statistically analyzing the periodic intensity, peak duration, and resource dependency type of the tenant service usage peak period.

[0130] Specifically, based on the service data distribution, the continuous time section exceeding the minimum steady-state flow value of the benchmark waveform and reaching the peak bandwidth in the identified waveform is defined as the service usage peak period of each tenant through the benchmark waveform of each tenant. Subsequently, the service usage peak period is time-axis aligned to identify the periodic repetition mode of the service usage peak period. Then, by calculating the standard deviation and mean of the maximum and minimum flow values in each cycle peak interval, if the standard deviation and mean are too low, it is marked as a uniform distribution peak, otherwise it is marked as a fluctuation distribution peak. The peak service distribution law of the flow value in the statistical service usage peak period is determined by integrating the uniform distribution peak and the fluctuation distribution peak. The coupling weight of each tenant service type and physical resource is obtained through the resource configuration requirement marked in the tenant service deployment list. Based on any two tenants, if there is a time overlap window in the service usage peak period of any two tenants, the overlap intensity index is calculated. Then, combined with the upper limit of the physical node resource capacity and the lower limit of the tenant SLA performance commitment, the preset threshold is determined through historical contention data regression analysis. If the overlap intensity index exceeds the preset threshold, the tenant overlapping service is determined.

[0131] Through the scheme, the service data distribution is analyzed to determine the service usage peak period of each tenant, avoiding the interference of abnormal events on the identification of the peak interval, and ensuring that the extracted service usage peak period only reflects the real and stable service load characteristics of the tenant. Based on the service usage peak period, the peak service distribution law of each tenant is determined to avoid misassociation of non-periodic tenants and provide a basis for resource coupling analysis. According to the peak service distribution law, the tenant overlapping service is determined to avoid resource fragmentation and secondary contention caused by full migration.

[0132] In some embodiments, according to the peak service distribution law, the peak service type is determined; based on the peak service type, the service data distribution is analyzed to determine whether any two peak service types are in the overlapping stage of the usage period; if so, according to the service data distribution, it is determined whether any two peak service types in the overlapping stage of the usage period belong to the same physical server; if so, according to the usage period of the peak service, the service overlap degree is determined; the service overlap degree is compared with the preset overlap threshold, if the service overlap degree is higher than the preset overlap threshold, it is determined as the tenant overlapping service.

[0133] The peak service type can be a resource dependency characteristic category exhibited by the tenant in the peak service distribution law. The usage time period overlap stage can be a time section in which the service usage peak periods of two tenants exist on a time axis. The physical server can be an entity hardware device that carries a tenant virtual machine or a container instance. The peak service can be a high-load service stage. The usage time period can be an active period of tenant service on a time axis. The service overlap degree can be an index quantifying the comprehensive contention intensity of two tenants on the same physical server. The preset overlap threshold can be a critical value preset for determining whether the service overlap degree triggers a resource isolation action. The preset overlap threshold is pre-stored in the server and called when in use.

[0134] Specifically, based on the peak service distribution law of the tenant, the resource dependency type of the peak service distribution law is extracted, and the peak service type is determined in combination with the resource configuration requirement in the tenant service deployment manifest. The usage peak period of each identified tenant service is periodically aligned on a time axis, and then, based on the service data distribution, the peak time periods of several tenants are scanned through a sliding window to detect whether any two peak service types are in a usage time period overlap stage. If in the usage time period overlap stage, based on the service data distribution, whether the two peak service types that exist in overlap belong to the same physical server is determined through the physical server identifier recorded in the tenant service deployment manifest. When belonging to the same physical server, the time overlap window length is determined according to the proportion of the overlap period to the peak duration of the two services. Then, based on the coupling weight of the resource dependency type, the weight difference absolute value is calculated. Subsequently, the proportion of the sum of the storage I / O throughput or network bandwidth usage in the overlap period to the upper limit of the corresponding resource capacity of the server. Finally, the time overlap window length, the weight difference absolute value, and the resource usage proportion are weighted and summed to obtain the service overlap degree. Further, based on the correlation between the resource contention intensity and the performance degradation in the historical SLA violation data, a preset overlap threshold is set, and then the service overlap degree is compared with the preset overlap threshold. If the service overlap degree is higher than the preset overlap threshold, it is determined that the tenant overlap service exists.

[0135] According to the peak service distribution law, the peak service type is determined, and the misjudgment caused by the ambiguous service type is avoided. Based on the peak service type, the service data distribution is analyzed, and it is determined whether any two peak service types are in the overlapping stage of the use period. The invalid detection of the non-overlapping period is excluded, and the roughness of relying on the simple time window division is overcome. Only the actual overlapping period is concerned to reduce the calculation redundancy. If yes, according to the service data distribution, it is determined whether any two peak service types in the overlapping stage of the use period belong to the same physical server, which helps to exclude the hardware-independent interference between cross-server tenants and avoid the problem that the isolation strategy at the virtual machine level cannot cover the underlying resource competition. If yes, according to the use period of the peak service, the service overlap degree is determined, which helps to quantify the actual competition intensity of the two tenants for the shared resources and avoid the one-sidedness of relying only on time or a single indicator. The service overlap degree is compared with the preset overlap threshold value, and if the service overlap degree is higher than the preset overlap threshold value, it is determined that the tenant has overlapping services, which helps to adapt to the nonlinear change of the resource competition intensity and avoid the SLA violation caused by the response lag.

[0136] In some embodiments, I / O data and SSD garbage collection data are analyzed to determine I / O queue depth and SSD garbage collection status; according to the I / O queue depth and the SSD garbage collection status, the I / O delay growth rate is determined; according to the service data distribution and the I / O delay growth rate, the competition intensity of the tenant overlapping service is determined.

[0137] The I / O data can be input / output operation records. The SSD garbage collection data can be event information recorded by the SSD when performing the garbage collection operation. The I / O queue depth can be the number of uncompleted I / O requests. The SSD garbage collection status can be the active state of the SSD garbage collection operation. The I / O delay growth rate can be the relative change rate of the storage I / O operation response delay in adjacent time windows.

[0138] Specifically, I / O data and SSD garbage collection data are extracted from historical service traffic data; then, based on the difference between the accumulation rate and the processing rate of the request queue of the I / O data, the I / O queue depth is calculated; at the same time, the SSD garbage collection log is analyzed, and based on the correlation between the triggering frequency and the duration, the current SSD garbage collection state is determined. According to the combination relationship of the I / O queue depth and the SSD garbage collection state, an I / O delay dynamic model is established: first, when the SSD garbage collection state is in the burst high-load cleaning state, the additional delay coefficient caused by garbage collection is superimposed; second, when the I / O queue depth exceeds the physical channel carrying capacity, the delay weight is increased by the overload ratio; thus, the I / O delay dynamic model is established; based on the output of the I / O delay dynamic model, the I / O delay growth rate per unit time is calculated. Combined with the service data distribution, the actual resource usage baseline of each tenant in the overlapping period is extracted; then, based on the correlation mapping between the historical resource usage baseline of the tenant and the performance indicators in the SLA, the dynamic resource usage mode is combined with the static performance commitment to set the SLA allowed threshold, and then the I / O delay growth rate is coupled with the actual resource usage baseline for calculation. If the I / O delay growth rate exceeds the SLA allowed threshold corresponding to the actual resource usage baseline, the contention intensity value is increased by the exceeding ratio; if the SSD garbage collection state is high-load cleaning and overlaps with the tenant peak period, the hardware-level interference weight is further superimposed; finally, the contention intensity of the tenant overlapping service is output.

[0139] Through the scheme, I / O data and SSD garbage collection data are analyzed, I / O queue depth and SSD garbage collection state are determined, the instantaneous load pressure of the physical server storage I / O resource is intuitively reflected, and the influence period of the storage device self-maintenance operation on the I / O performance is clear, which improves the accuracy of the contention intensity quantification. According to the I / O queue depth and the SSD garbage collection state, the I / O delay growth rate is determined, which helps to eliminate the defects of dynamic interference response lag and early warning of potential SLA violation risk. According to the service data distribution and the I / O delay growth rate, the contention intensity of the tenant overlapping service is determined, which helps to eliminate the problem of fragmented resource utilization inefficiency, provides comparable basis for migration decision, and reduces the secondary contention risk after migration.

[0140] In some embodiments, according to the SLA isolation requirement, the basic performance between tenants is determined; according to the contention intensity, the performance interference coefficient of the tenant overlapping service is determined; the dynamic change of the tenant overlapping service is determined by analyzing the service data distribution; according to the dynamic change, the performance interference coefficient, and the basic performance, the resource isolation strategy is determined.

[0141] The basic performance can be a minimum performance guarantee index. The performance interference coefficient can be a numerical weight used to represent an actual interference level of resource contention on tenant performance. The dynamic change can be a fluctuation feature of a resource usage baseline in a tenant overlapping business period.

[0142] Specifically, SLA isolation requirements of the tenant are analyzed, and a static performance constraint index is extracted. The static performance constraint index is mapped to physical resource capability, and a "performance-resource" corresponding relationship is established, thereby generating the basic performance between tenants. Based on contention intensity, a resource dependency matrix is defined according to the business type of the tenant, wherein the resource dependency matrix is generated based on historical business traffic feature statistical analysis. Then, the contention intensity and the resource dependency matrix are point multiplied, and a performance interference coefficient is output. Time series analysis is performed on the business data distribution, and the peak overlapping period of the tenant is extracted. The sliding window method is used to extract the change trend in the peak overlapping period, thereby generating the dynamic change of the tenant overlapping business. The dynamic change is mapped to a resource demand increment, and the performance interference coefficient is added to the correction amount of the basic performance, thereby generating a dynamic resource allocation quota. If migration needs to be triggered, according to the fragmentation resource distribution characteristics of the physical server, the target node with the shortest migration path and the highest fragmentation matching degree is selected, and a resource isolation strategy is generated.

[0143] According to the SLA isolation requirements, the basic performance between tenants is determined, which avoids the default of resource allocation and prevents the problems of over-allocation or under-allocation of resources caused by ambiguous indicators. According to the contention intensity, the performance interference coefficient of the tenant overlapping business is determined, which reflects the potential threat level of current resource contention on performance, and avoids the limitations of virtual machine level isolation strategy. The business data distribution is analyzed to determine the dynamic change of the tenant overlapping business, which ensures to reflect the real business demand of the tenant and avoids the misjudgment of the strategy caused by occasional events. According to the dynamic change, the performance interference coefficient and the basic performance, the resource isolation strategy is determined, which reduces the amount of migration data and avoids the secondary contention caused by full migration.

[0144] In some embodiments, based on the contention intensity, the business data distribution is analyzed to determine the data access dependency relationship; according to the data access dependency relationship, the cross-tenant shared cache line and / or contention lock resource is determined; according to the cross-tenant shared cache line and / or contention lock resource, the hot spot interference area is determined; based on the business data distribution, the hot spot interference area is analyzed to determine the performance interference coefficient of the tenant overlapping business.

[0145] The data access dependency relationship can be a space-time correlation access mode of tenant business on a physical resource. The cross-tenant shared cache line can be a hardware-level resource contention of businesses of multiple tenants due to access to the same physical memory cache line. The contention lock resource can be a resource mutual exclusion waiting caused by competition for the same hardware or software lock when multiple tenant processes are concurrently operated. The hotspot interference area can be a high-density conflict area caused by the cross-tenant shared cache line and / or the contention lock resource.

[0146] Specifically, based on the contention intensity, the physical resource address distribution characteristics of tenant business access are extracted through the distribution of business data after abnormal traffic is removed; then, according to the physical resource address distribution characteristics, a data access mode matrix across the time dimension is generated, and the data access dependency relationship of tenant business to the physical resource is identified. According to the data access dependency relationship, the cross-tenant shared cache line of different tenants accessing the same physical storage page is extracted; the lock operation log is monitored, and the contention lock resource of the same lock resource being alternately held by multiple tenant processes is counted. The cross-tenant shared cache line and / or the contention lock resource are clustered according to the physical resource location, and a hardware-level interference heat map is generated; according to the resource conflict density in the hardware-level interference heat map, the hotspot interference area is marked. Based on the contention intensity, the hardware-level weight of the hotspot interference area is superimposed; then, according to the resource demand fluctuation of the overlapping period in the business data distribution, the interference coefficient calculation proportion is dynamically adjusted, and thus the final performance interference coefficient is output.

[0147] Through the scheme, based on the contention intensity, the business data distribution is analyzed, the data access dependency relationship is determined, and the resource allocation strategy is avoided to deviate from the real demand due to abnormal peaks. According to the data access dependency relationship, the cross-tenant shared cache line and / or the contention lock resource are determined, which helps to make up for the defects caused by only focusing on virtual resources. According to the cross-tenant shared cache line and / or the contention lock resource, the hotspot interference area is determined, which helps to avoid that the virtual machine abstraction layer hides the real resource conflict position, ensures that the isolation strategy focuses on the conflict point, and at the same time avoids the resource fragmentation problem caused by full migration. Based on the business data distribution, the hotspot interference area is analyzed, the performance interference coefficient of the tenant overlapping business is determined, which helps to realize the fine quantization of the contention intensity and provides accurate input for the migration decision.

[0148] In some embodiments, the performance interference coefficient is compared with a preset interference intensity threshold value, if the performance interference coefficient is higher than the preset interference intensity threshold value, the adjustment plan of the basic performance is determined according to the dynamic change; the SLA isolation requirement is analyzed, and it is determined whether the tenant accepts the adjustment plan; if yes, the adjustment plan is determined as the resource isolation strategy.

[0149] The preset interference intensity threshold value can be a preset dynamic threshold value for determining whether to trigger a resource adjustment action. The preset interference intensity threshold value is pre-stored in a server and is called when used. The adjustment plan can be a strategy set generated according to dynamic changes when the performance interference coefficient exceeds the preset interference intensity threshold value.

[0150] Specifically, the performance interference coefficient is compared with the preset interference intensity threshold value calculated based on the basic performance, and it is determined whether to trigger resource adjustment. If the performance interference coefficient exceeds the preset interference intensity threshold value, the contention intensity evolution direction of the future time window is predicted based on the contention intensity recorded in the hardware-level interference heat map and using a time series prediction algorithm, so as to generate an adjustment plan containing a resource degradation amplitude or a migration target node. Then, a dynamic adjustment acceptance clause defined in the SLA isolation requirement is extracted, and the degradation amplitude and duration in the adjustment plan are matched and verified with the clause to determine whether the tenant accepts the adjustment plan. If the tenant accepts the adjustment plan, an optimal migration path is calculated based on the basic performance, and a resource isolation strategy containing a migration data volume limit and a lock resource reorganization rule is generated.

[0151] By the scheme, the performance interference coefficient is compared with the preset interference intensity threshold value. If the performance interference coefficient is higher than the preset interference intensity threshold value, an adjustment plan of the basic performance is determined according to dynamic changes, which helps to eliminate the problem that a static threshold value cannot adapt to nonlinear delay growth, avoids migrating a high-contention service to a node whose load is saturated, and reduces the probability of secondary contention. The SLA isolation requirement is analyzed to determine whether the tenant accepts the adjustment plan, and a differentiated isolation strategy negotiation is realized. If the adjustment plan is accepted, the adjustment plan is determined as a resource isolation strategy, which helps to eliminate the problem of fragmented resource utilization inefficiency and reduce secondary resource conflicts caused by full migration.

[0152] In some embodiments, if the adjustment plan is not accepted, a resource fragmentation distribution state is determined according to a service data distribution situation, a target node minimizing data migration amount is determined according to the resource fragmentation distribution state, and a double-write mechanism is used to migrate the target node.

[0153] The resource fragmentation distribution state can be a scattered distribution feature of available resources on a physical node. The data migration amount minimizing can be only migrating a hot spot interference area data block causing resource contention in a resource migration process. The target node can be a physical node selected according to the resource fragmentation distribution state. The double-write mechanism can be synchronously writing corresponding data blocks in the source node and the target node in the migration process.

[0154] Specifically, if the adjustment plan is not accepted, the fragmentation characteristics of the available resources on the physical node are analyzed according to the business data distribution, and a resource fragmentation distribution state is generated. Then, based on the resource fragmentation distribution state, a candidate node that meets the basic performance requirement and has the highest complementarity with the current node fragmentation is matched, so as to filter out a target node with the smallest data migration amount. During the migration process, a double-write operation is synchronously performed on the target node and the source node, and the access entry is switched when the migration is completed and the target node passes the consistency check.

[0155] By the scheme, if the adjustment plan is not accepted, the resource fragmentation distribution state is determined according to the business data distribution, which helps to avoid the waste of fragmented resources caused by full data migration or random node selection, and reduce the redundant data migration amount. According to the resource fragmentation distribution state, the target node with the smallest data migration amount is determined, which avoids the secondary contention risk caused by the unoptimized path during migration, and ensures that the target node resources can support the tenant SLA requirements. The double-write mechanism is adopted to migrate the target node, which helps to eliminate the business interruption risk caused by single-point writing or full replication during migration, realizes seamless switching and avoids data loss, and reduces the additional impact of migration on I / O performance.

[0156] Figure 3 A structural diagram of a data maintenance application system of a server lease according to an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the data maintenance application system 300 of the server lease according to the embodiment includes a data analysis module 301, an overlap analysis module 302, and a strategy determination module 303. Figure 3

[0157] The data analysis module 301 is configured to obtain historical business traffic data of a tenant, analyze the historical business traffic data, and determine a business data distribution.

[0158] The overlap analysis module 302 is configured to determine a tenant overlap business and a contention intensity of the tenant overlap business according to the business data distribution.

[0159] The strategy determination module 303 is configured to obtain an SLA isolation requirement of a tenant, and determine a resource isolation strategy according to the contention intensity based on the SLA isolation requirement.

[0160] Optionally, when the data analysis module 301 analyzes the historical business traffic data and determines a business data distribution, the data analysis module 301 is configured to:

[0161] analyze the historical business traffic data to determine whether the historical business traffic data contains a burst traffic peak;

[0162] if the historical business traffic data contains the burst traffic peak, analyze the burst traffic peak to determine a peak operation business, and mark the peak operation business as an abnormal event; ​

[0163] analyze the abnormal event, determine abnormal traffic data, and eliminate the abnormal traffic data from the historical traffic data to obtain normal traffic data;

[0164] analyze the normal traffic data by Fourier transform to determine a periodic waveform;

[0165] perform noise elimination on the periodic waveform by using a sliding window filtering algorithm to obtain a reference waveform;

[0166] analyze the reference waveform to determine a service period feature, and use the service period feature as the service data distribution.

[0167] Optionally, when the overlap analysis module 302 determines tenant overlap services according to the service data distribution, it is used for:

[0168] analyzing the service data distribution to determine a service use peak period of each tenant;

[0169] determining a peak service distribution rule of each tenant based on the service use peak period;

[0170] determining tenant overlap services according to the peak service distribution rule.

[0171] Optionally, when the overlap analysis module 302 determines tenant overlap services according to the peak service distribution rule, it is used for:

[0172] determining a peak service type according to the peak service distribution rule;

[0173] analyzing the service data distribution based on the peak service type to determine whether any two peak service types are in a use period overlap stage;

[0174] if yes, determining whether the any two peak service types in the use period overlap stage belong to the same physical server according to the service data distribution;

[0175] if yes, determining a service overlap degree according to the use period of the peak service;

[0176] comparing the service overlap degree with a preset overlap threshold value, and if the service overlap degree is higher than the preset overlap threshold value, determining that it is a tenant overlap service.

[0177] Optionally, the historical traffic data further includes I / O data and SSD garbage collection data, and when the overlap analysis module 302 determines the contention intensity of tenant overlap services according to the service data distribution, it is used for:

[0178] analyze the I / O data and the SSD garbage collection data to determine an I / O queue depth and an SSD garbage collection state;

[0179] determine an I / O delay growth rate according to the I / O queue depth and the SSD garbage collection state;

[0180] determine a contention intensity of the tenant overlapping services according to the service data distribution and the I / O delay growth rate.

[0181] Optionally, when the policy determination module 303 determines the resource isolation policy according to the tenant overlapping services and the contention intensity based on the SLA isolation requirement, the policy determination module 303 is configured to:

[0182] determine a basic performance between tenants according to the SLA isolation requirement;

[0183] determine a performance interference coefficient of the tenant overlapping services according to the contention intensity;

[0184] analyze the service data distribution to determine a dynamic change of the tenant overlapping services;

[0185] determine the resource isolation policy according to the dynamic change, the performance interference coefficient and the basic performance.

[0186] Optionally, when the policy determination module 303 determines the performance interference coefficient of the tenant overlapping services according to the contention intensity, the policy determination module 303 is configured to:

[0187] analyze the service data distribution based on the contention intensity to determine a data access dependency relationship;

[0188] determine a cross-tenant shared cache line and / or contention lock resource according to the data access dependency relationship;

[0189] determine a hotspot interference area according to the cross-tenant shared cache line and / or contention lock resource;

[0190] determine the performance interference coefficient of the tenant overlapping services based on the service data distribution and the hotspot interference area.

[0191] Optionally, when the policy determination module 303 determines the resource isolation policy according to the dynamic change, the performance interference coefficient and the basic performance, the policy determination module 303 is configured to:

[0192] compare the performance interference coefficient with a preset interference intensity threshold value, and if the performance interference coefficient is higher than the preset interference intensity threshold value, determine an adjustment plan of the basic performance according to the dynamic change;

[0193] analyzing the SLA isolation requirement, determining whether the tenant accepts the adjustment plan;

[0194] If accepted, the adjustment plan is determined as the resource isolation strategy.

[0195] Optionally, the data maintenance application system of the server lease further comprises a node migration module 304, configured to:

[0196] If not accepted, according to the business data distribution, a resource fragmentation distribution state is determined;

[0197] According to the resource fragmentation distribution state, a target node minimizing data migration amount is determined;

[0198] The target node is migrated by using a double-write mechanism.

[0199] The system of the embodiment can be used to execute the method of any of the above embodiments, and has similar implementation principles and technical effects, which will not be described here again.

Claims

1. A data maintenance method for server leasing, characterized in that: include: Obtain tenants' historical business traffic data; Analyze the historical business traffic data to determine the business data distribution; Determining, based on the distribution of the service data, tenant overlapping services and contention intensity of the tenant overlapping services; Obtain the tenant's SLA isolation requirements; Determining a resource isolation strategy based on the SLA isolation requirements and the tenant overlapping services and the contention intensity; The determining of tenant overlapping services according to the distribution of the service data includes: Analyze the distribution of the business data and determine the peak business usage period of each tenant; Determine the peak service distribution pattern of each tenant based on the peak service usage period; Determining the peak service type according to the peak service distribution pattern; Based on the peak service types, analyzing the service data distribution to determine whether any two peak service types have overlapping usage periods; If yes, then determining whether any two peak service types in the overlapping usage periods belong to the same physical server based on the service data distribution; If yes, determine the service overlap based on the peak service usage period; Comparing the service overlap degree with a preset overlap threshold, and determining that the service is a tenant overlapping service if the service overlap degree is higher than the preset overlap threshold; The historical service traffic data also includes I / O data and SSD garbage collection data. Determining the contention intensity of the tenant's overlapping services based on the service data distribution includes: Analyzing the I / O data and SSD garbage collection data to determine the I / O queue depth and SSD garbage collection status; Determining an I / O latency growth rate based on the I / O queue depth and the SSD garbage collection status; The contention intensity of the tenant overlapping services is determined according to the service data distribution and the I / O delay growth rate.

2. The method according to claim 1, characterized in that The analyzing the historical service flow data to determine the service data distribution includes: Analyzing the historical service traffic data to determine whether it contains sudden traffic spikes; If included, analyzing the sudden traffic peak, determining the peak-running service, and marking the peak-running service as an abnormal event; Analyzing the abnormal event, determining abnormal traffic data, and removing the abnormal traffic data from the historical business traffic data to obtain normal business traffic data; Analyzing the normal service flow data through Fourier transform to determine a periodic waveform; Using a sliding window filtering algorithm to remove noise from the periodic waveform to obtain a reference waveform; The reference waveform is analyzed to determine service cycle characteristics, and the service cycle characteristics are used as the service data distribution.

3. The method according to claim 1, characterized in that The determining of a resource isolation strategy based on the SLA isolation requirement and according to the tenant overlapping services and the contention intensity includes: Determine the basic performance between tenants based on the SLA isolation requirements; determining a performance interference coefficient of the tenant's overlapping services according to the contention intensity; Analyze the business data distribution and determine the dynamic changes of the tenants' overlapping businesses; A resource isolation strategy is determined based on the dynamic change, the performance interference coefficient, and the basic performance.

4. The method according to claim 3, characterized in that The determining, according to the contention intensity, a performance interference coefficient of the tenant overlapping service includes: Analyzing the distribution of the service data based on the contention intensity to determine data access dependencies; Determining cross-tenant shared cache lines and / or contention lock resources based on the data access dependency; Determining a hotspot interference area based on the cross-tenant shared cache lines and / or contention lock resources; Based on the distribution of the service data, the hotspot interference area is analyzed to determine the performance interference coefficient of the tenant's overlapping service.

5. The method according to claim 3, characterized in that The determining of the resource isolation strategy according to the dynamic change, the performance interference coefficient, and the basic performance includes: comparing the performance interference coefficient with a preset interference intensity threshold, and if the performance interference coefficient is higher than the preset interference intensity threshold, determining an adjustment plan for the basic performance according to the dynamic change; Analyze the SLA isolation requirements and determine whether the tenant accepts the adjustment plan; If accepted, the adjustment plan is determined as the resource isolation policy.

6. The method according to claim 5, characterized in that The method further comprises: If not accepted, determining the resource fragment distribution state according to the business data distribution; Determining a target node that minimizes the amount of data migration based on the resource fragmentation distribution state; A double-write mechanism is used to migrate the target node.

7. A data maintenance application system for server leasing, characterized in that: Applicable to executing the method according to any one of claims 1 to 6, comprising: A data analysis module is used to obtain the tenant's historical business traffic data; analyze the historical business traffic data to determine the business data distribution; An overlap analysis module, configured to determine overlapping services of tenants and contention intensity of the overlapping services of the tenants according to the distribution of the service data; The policy determination module is used to obtain the SLA isolation requirements of the tenant; based on the SLA isolation requirements and according to the contention intensity, determine the resource isolation policy.

Citation Information

Patent Citations

  • Multi-service and multi-tenant-oriented micro-service method and system

    CN119917289A

  • User-oriented multi-tenant security enhancement service isolation system

    CN120151065A