Distributed data storage method and device based on multiple machine rooms

By deploying proxy services and caching containers in each data center and selecting an appropriate list of data centers for data caching, the problem of business interruption caused by single data center failures was solved, and high availability and fault tolerance for multi-data center data storage were achieved.

CN120909499APending Publication Date: 2025-11-07BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510843167.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional distributed coordination services are typically deployed in a single data center. When the data center fails, business application services cannot read the stored data, affecting the continuity and availability of business application services.

Method used

Deploy proxy services and cache containers in each data center, and select local data center and at least one remote data center from the data center list to configure cache information, thereby achieving multi-data center replica storage, ensuring that data is cached in multiple data centers and avoiding single points of failure.

Benefits of technology

It improves the continuity and availability of business application services, simplifies configuration management through the collaborative work of proxy services and caching containers, and is suitable for large-scale, multi-regional distributed systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909499A_ABST
    Figure CN120909499A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed data storage method and device based on multiple machine rooms, relates to the technical field of data processing, and mainly aims to realize multi-machine-room copy storage of data, so that the stored data can be merged and read from the multiple machine rooms, and the continuity and availability of business application services are improved. According to the main technical scheme, when an agency service of a target machine room receives configuration information of an application service, a to-be-cached machine room is determined from a machine room list corresponding to the agency service, the machine room list is constructed based on the fault rate of each machine room, and the to-be-cached machine room comprises a local machine room and at least one remote machine room; and writing the configuration information into a cache container in the to-be-cached machine room in a specified data structure. The method and the device are used for data storage of business application services.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a distributed data storage method and device based on multiple machine rooms. BACKGROUND

[0002] With the rapid development of Internet technology, enterprises and service providers are facing growing data storage and access demands. Distributed data storage systems have become a core component of modern data centers and cloud computing platforms due to their high availability, high scalability and high performance. Distributed data storage systems distribute data in multiple geographically distributed machine rooms, which not only improves the fault tolerance and data reliability of the system, but also improves user experience through load balancing and local access.

[0003] Currently, the existing technology usually applies distributed coordination services to configuration management, naming services, distributed locks and other scenarios of distributed systems, which can provide high availability and fault tolerance by maintaining a distributed and consistent data storage. However, traditional distributed coordination services are usually deployed within a single machine room, and when the machine room fails, the business application service will not be able to read the corresponding storage data, affecting the continuity and availability of the business application service. SUMMARY

[0004] In view of the above problems, the present application provides a distributed data storage method and device based on multiple machine rooms, the main purpose of which is to realize multi-machine room replica storage of data, so that the stored data can be read from multiple machine rooms, thereby improving the continuity and availability of the business application service.

[0005] To solve the above technical problems, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a distributed data storage method based on multiple machine rooms, comprising:

[0007] When the proxy service of the target machine room receives the configuration information of the application service, determine the to-be-cached machine room from the machine room list corresponding to the proxy service, the machine room list is constructed based on the failure rate of each machine room, and the to-be-cached machine room includes the local machine room and at least one remote machine room;

[0008] Write the configuration information into the cache container in the to-be-cached machine room in a specified data structure.

[0009] In a second aspect, the present application provides a distributed data storage device based on multiple machine rooms, comprising:

[0010] The first determining unit is configured to determine a to-be-cached machine room from a machine room list corresponding to the proxy service of the target machine room when the proxy service receives the configuration information of the application service, the machine room list being constructed based on failure rates of respective machine rooms, and the to-be-cached machine room including a local machine room and at least one remote machine room.

[0011] The processing unit is configured to write the configuration information into a cache container in the to-be-cached machine room obtained by the first determining unit in a specified data structure.

[0012] To achieve the above object, according to a third aspect of the present application, a storage medium is provided, which comprises a stored program, wherein the storage medium controls a device where the storage medium is located to execute the multi-machine-room-based distributed data storage method of the first aspect when the program is run.

[0013] To achieve the above object, according to a fourth aspect of the present application, a processor is provided, which is configured to run a program, wherein the processor executes the multi-machine-room-based distributed data storage method of the first aspect when the program is run.

[0014] By the above technical solution, the multi-machine-room-based distributed data storage method and device provided by the present application can determine a to-be-cached machine room from a machine room list corresponding to a proxy service when the proxy service receives configuration information of an application service, the machine room list being constructed based on failure rates of respective machine rooms, and the to-be-cached machine room including a local machine room and at least one remote machine room, and write the configuration information into a cache container in the to-be-cached machine room in a specified data structure. Through the technical solution provided by the present application, the configuration information can be cached in multiple machine rooms, so that even if a machine room fails, other machine rooms can still provide the configuration information, and the storage data can be read in combination, thereby avoiding the problem of single-point failure, improving the fault tolerance, and further improving the data reliability by preferentially selecting machine rooms with lower failure rates for caching, thereby ensuring the continuity and availability of business application services. In addition, the cooperation of the proxy service and the cache container effectively simplifies the configuration management. The present application has obvious advantages in large-scale, multi-geographical distributed systems.

[0015] The above description is only a summary of the technical solution of the present application. In order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented in accordance with the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0016] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not intended to limit the scope of the present application. The same reference numbers in different drawings identify the same components or elements throughout the text. In the drawings:

[0017] Figure 1 A flow chart of a method for distributed data storage based on multiple machine rooms is shown according to an embodiment of the present application;

[0018] Figure 2 A flow chart of another method for distributed data storage based on multiple machine rooms is shown according to an embodiment of the present application;

[0019] Figure 3 A block diagram of a device for distributed data storage based on multiple machine rooms is shown according to an embodiment of the present application;

[0020] Figure 4 A block diagram of another device for distributed data storage based on multiple machine rooms is shown according to an embodiment of the present application;

[0021] Figure 5 A schematic diagram of a system for distributed data storage based on multiple machine rooms is shown according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] Exemplary embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While example embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms without being limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0023] At present, the existing technology usually applies distributed coordination services to configuration management, naming service, distributed lock and other scenarios of distributed systems, which can provide high availability and fault tolerance by maintaining a distributed and consistent data storage. However, the traditional distributed coordination service is usually deployed in a single machine room, and when the machine room fails, the business application service will not be able to read the corresponding storage data, affecting the continuity and availability of the business application service.

[0024] The inventor finds that the logic of using a distributed coordination service to implement data storage in a single machine room can be abandoned, and a proxy service and a cache container are respectively deployed in each machine room in advance, and a machine room list is set for each proxy service, the machine room list including a local machine room and at least one off-site machine room with a lower failure rate screened according to a failure rate, the proxy service selects the local machine room and the at least one off-site machine room from the machine room list, and writes configuration information into the cache container of the local machine room and the at least one off-site machine room respectively. In this way, the configuration information can be cached in multiple machine rooms, even if a machine room fails, other machine rooms can still provide configuration information, the combined reading of stored data is realized, the problem of single point failure is avoided, and the continuity and availability of business application services are ensured.

[0025] Based on the above considerations, the embodiment of the present application provides a distributed data storage method based on multiple machine rooms, which can realize multiple machine room copy storage of data, so that the stored data can be read from multiple machine rooms, thereby improving the continuity and availability of business application services. The embodiment of the present application is applied to a distributed system. The specific execution steps are as shown in Figure 1

[0026] 101. When the proxy service in the target machine room receives the configuration information of the application service, the proxy service determines the to-be-cached machine room from the machine room list corresponding to the proxy service.

[0027] The machine room list is constructed based on the failure rate of each machine room, and the to-be-cached machine room includes a local machine room and at least one off-site machine room.

[0028] ​It should be noted that in the present embodiment, the distributed system comprises multiple machine rooms, each of which has a unique identifier, such as a machine room ID, and each of which is deployed with a proxy service and a cache container, the proxy service being responsible for receiving configuration information of an application service and selecting a to-be-cached machine room according to a machine room list, and the cache container being used for storing configuration information. Each proxy service has a corresponding machine room list, which is constructed based on the failure rates of various machine rooms in the system, including the local machine room and at least one off-site machine room with a lower failure rate. The failure rate can be calculated from historical failure records of various off-site machine rooms, or can be estimated according to the hardware configuration and environmental conditions of various off-site machine rooms, or can be determined by combining the two. The present embodiment does not limit this. The machine room list can be stored in the configuration file or system database of the proxy service. It should be noted that the machine room list can be a dedicated machine room list corresponding to one proxy service, or a common machine room list corresponding to all proxy services, and the machine room list will be updated regularly, which can be updated according to a preset period, or by detecting whether each machine room triggers a failure. For example, one or more monitoring services are deployed in advance to continuously monitor the machine room status of each machine room in the distributed system. Specifically, the availability of the server, the network connection state, the power supply state, and other indicators related to the health status of the machine room can be collected, and a series of rules and thresholds can be set based on the selected indicators to determine when the machine room is considered to have failed, for example, when a certain proportion of servers are offline, the network delay is too high, or communication is completely lost, the machine room is considered to have failed. At this time, the monitoring service immediately notifies the proxy service responsible for maintaining the machine room list, and the proxy service updates the machine room list after receiving the failure notification. In the present step, the proxy service receives configuration information from its corresponding application service, which can include but is not limited to: configuration parameters, version information, service address, digital label, etc. For example, application service A sends a configuration information to the proxy service of target machine room B, which contains the latest configuration parameters of service A. The proxy service obtains the machine room list from its configuration file or system database, and selects the local machine room and at least one off-site machine room therefrom. The off-site machine room can be selected according to the failure rate, such as the off-site machine room with the lowest failure rate, or the off-site machine room with a failure rate lower than a failure rate threshold. The off-site machine room can also be selected in combination with the failure rate, the geographical location and the recent load condition, such as setting a corresponding scoring system or membership matrix for the failure rate, the geographical location and the recent load condition, calculating the scores or memberships of the failure rate, the geographical location and the recent load condition, and then combining the pre-set weight coefficients for the failure rate, the geographical location and the recent load condition to calculate the comprehensive scores of each off-site machine room, and selecting the off-site machine room with a comprehensive score lower than a threshold.In addition, since the limit of the off-site machine room in the embodiment is at least one, and the number of off-site machine rooms required by different use requirements may be different, for example, in order to ensure high availability, 3-5 off-site machine rooms may be required, and in order to reduce cost or ensure higher read-write performance, 1-2 off-site machine rooms may be required. Therefore, different upper limits can also be set according to different use requirements, so that the determination of the to-be-cached machine room is more in line with user requirements.

[0029] 102. Write the configuration information in the specified data structure into the cache container in the to-be-cached machine room.

[0030] In this step, the proxy service converts the received configuration information into a specified data structure, which can be in the form of key-value pairs such as hash key, string key, etc., to ensure data standardization and parsability. The proxy service writes the structured configuration information into the cache container of the to-be-cached machine room, i.e. the cache container of the local machine room and the cache container of at least one selected off-site machine room, which can be an in-memory database, a file system or a distributed cache system (such as Redis). For example, the proxy service writes the configuration information into the cache container of the local machine room B and the cache container of the off-site machine room C. By locating each cache container in a different machine room, multiple copies of data storage can be ensured.

[0031] In subsequent queries, the application service reads data from multiple cache containers associated with it (including the local machine room and other selected off-site machine rooms), and combines these scattered data fragments into a complete configuration information and returns it to the application service. For example, application service A requests the latest configuration information, and application service A reads the configuration information from the cache container of the local machine room B and the cache container of the off-site machine room C, and combines them into a complete configuration information and returns it to the application service.

[0032] Based on the above Figure 1As can be seen from the implementation mode of the above technical solution, the application provides a distributed data storage method based on multiple machine rooms. When the proxy service of the target machine room receives the configuration information of the application service, the to-be-cached machine room is determined from the machine room list corresponding to the proxy service. The machine room list is constructed based on the failure rates of the machine rooms. The to-be-cached machine room includes the local machine room and at least one remote machine room. The configuration information is written into the cache container in the to-be-cached machine room in a specified data structure. Through the technical solution provided by the application, the configuration information can be cached in multiple machine rooms. Even if a machine room fails, other machine rooms can still provide configuration information, realizing the combined reading of stored data, thereby avoiding the problem of single point failure, improving the fault tolerance, and further improving the data reliability, thereby ensuring the continuity and availability of the business application service. In addition, through the cooperative work of the proxy service and the cache container, the configuration management is effectively simplified. It has obvious advantages in large-scale, multi-geographical distributed systems.

[0033] Further, the preferred embodiment of the application is based on the above Figure 1 Detailed description of the process of distributed data storage based on multiple machine rooms, as shown in Figure 2 201-206.

[0034] 201, obtain the historical failure rate of each remote machine room.

[0035] In this step, the historical failure records of each remote machine room are collected in advance from the monitoring system or the log. The monitoring system and the log usually record the failure events of the machine room, including the failure time, the failure type, the duration, etc.

[0036] The collected historical failure records are processed to calculate the historical failure rate of each remote machine room. The calculation formula of the historical failure rate is: historical failure rate = number of failures / total running time. For example, remote machine room C has occurred 5 times in the past year, and the total running time is 8760 hours (one year), then the historical failure rate of machine room C is: historical failure rate = 5 / 8760 ≈ 0.00057. After determining the historical failure rate, it is stored in the system database for subsequent steps.

[0037] 202, predict the configuration failure rate of each remote machine room according to the hardware configuration and environmental conditions of each remote machine room.

[0038] In this step, the hardware configuration information of each off-site machine room is collected in advance, including server model, storage device, etc. At the same time, the environmental condition information of each off-site machine room is collected, including temperature, humidity, power supply state, etc. For example, the environmental conditions of machine room C include: average temperature 25℃, average humidity 60%, and power supply state stable.

[0039] It should be noted that the hardware configuration at least includes one of the CPU core number, CPU frequency, memory capacity and read / write speed of the storage device, and the environmental condition at least includes one of the temperature, humidity and power supply state.

[0040] Based on historical data and expert experience, etc., a failure rate prediction model can be established to predict the configuration failure rate of the machine room. The failure rate prediction model can be a simple linear regression model, or a complex machine learning model, or a corresponding scoring system can be constructed for at least one of the CPU core number, CPU frequency, memory capacity, read / write speed and one of the temperature, humidity and power supply state, i.e. a score is given to each hardware configuration and environmental condition, and then the configuration failure rate is predicted by summing and averaging or weighted summation. This embodiment is not limited.

[0041] It needs to be explained that the input and output of the linear regression model and the machine learning model are similar, the input is various characteristic variables of hardware configuration and environmental conditions, including but not limited to CPU core number, CPU frequency, memory capacity, read-write speed, temperature, humidity and power supply state, etc., and the output is a numerical value representing the configuration failure rate, usually a real number between 0 and 1, where 0 means no failure risk at all, and close to 1 means a very high failure probability. The specific implementation process of the linear regression model is: obtaining the hardware configuration information and environmental condition data of each computer room from historical records, and marking the corresponding actual failure occurrence, cleaning, converting and standardizing the data, and extracting CPU core number, CPU frequency, memory capacity, read-write speed, temperature, humidity and power supply state, etc. from the data, using linear regression algorithm to fit the known data points, finding the best fitting straight line, such as minimizing mean square error (MSE), evaluating the model performance through cross-validation to ensure its good generalization ability, and using the trained linear regression model to calculate the corresponding configuration failure rate estimate value for new or unknown data sets. The specific implementation process of the machine learning model is: selecting a suitable machine learning algorithm according to the characteristics of the problem, such as random forest, gradient boosting tree (GBDT), support vector machine (SVM) or neural network, etc. Similar to linear regression, obtain the hardware configuration information and environmental condition data of each computer room from historical records, and mark the corresponding actual failure occurrence, clean, convert and standardize the data, and extract CPU core number, CPU frequency, memory capacity, read-write speed, temperature, humidity and power supply state, etc. from the data, use the selected machine learning algorithm to train the model, and adjust the hyperparameters to optimize the performance, use multiple indicators (such as accuracy, recall, F1 score, etc.) to evaluate the model effect, and prevent overfitting through K-fold cross-validation, etc. Use the trained machine learning model to predict new or unknown data sets. Compared with linear regression, machine learning model can capture more complex nonlinear relationships and provide better prediction accuracy.

[0042] For the above description of the scoring system, the specific execution process of predicting the configuration failure rate of each off-site computer room using the hardware configuration and environmental conditions of each off-site computer room is: determining the hardware score corresponding to the hardware configuration and the environmental score corresponding to the environmental condition using the preset scoring system; calculating the configuration failure rate of each off-site computer room according to the hardware score and the environmental score.

[0043] In this step, the scoring system is based on the degree of influence of hardware configuration and environmental conditions on the machine room failure. Specifically, historical failure records of each machine room are collected in advance, including failure time, failure type, duration, etc., as well as hardware configuration information and environmental condition information of each machine room, including server model, CPU core number, CPU frequency, memory capacity, read / write speed, temperature, humidity, power supply status, etc.

[0044] Among them, the more CPU cores, the stronger the processing capability, and the failure rate may decrease; the higher the frequency, the faster the processing speed, but the power consumption and heat may increase, affecting stability; the larger the memory, the smoother the system runs, and the failure rate may decrease; the faster the read / write speed, the higher the data transmission efficiency, but high-speed devices may be more expensive and have higher maintenance costs; too high or too low temperature may affect the normal operation of the device; too high humidity may cause device short circuit, and too low humidity may cause static problems; unstable power supply may cause device power failure, affecting system stability. Therefore, CPU core number, CPU frequency, memory capacity, read / write speed are taken as hardware configuration factors, and temperature, humidity, power supply status are taken as environmental condition factors. Analyze historical failure records to find the failure rate of each factor in different value range, and take the failure rate of each factor in different value range as the degree of influence of each factor on machine room failure. For example, analyzing historical data finds that the failure rate of machine room with CPU core number of 24 cores and above is the lowest, which can be given 10 points; and the failure rate of machine room with CPU core number below 4 cores is the highest, which can be given 2 points. For example, as follows:

[0045] CPU core number: 24 cores and above: 10 points; 16-23 cores: 8 points; 8-15 cores: 6 points; 4-7 cores: 4 points; 4 cores and below: 2 points.

[0046] CPU frequency: 2.4 GHz and above: 8 points; 2.0-2.3 GHz: 6 points; 1.6-1.9 GHz: 4 points; 1.6 GHz and below: 2 points.

[0047] Memory capacity: 128 GB and above: 9 points; 64-127 GB: 7 points; 32-63 GB: 5 points; 16-31 GB: 3 points; 16 GB and below: 1 point;

[0048] Read / write speed: 3500 MB / s and above: 9 points; 2000-3499 MB / s: 7 points; 1000-1999 MB / s: 5 points; 500-999 MB / s: 3 points; 500 MB / s and below: 1 point.

[0049] Temperature: 20-25°C: 8 points; 25-30°C: 6 points; 30-35°C: 4 points; 35°C and above: 2 points.

[0050] Humidity: 40-60%: 8 points; 60-70%: 6 points; 70-80%: 4 points; above 80%: 2 points.

[0051] Power supply status: Stable (no power outage records): 10 points; Occasional power outages (no more than 3 times per year): 6 points; Frequent power outages (more than 3 times per year): 2 points.

[0052] Based on the specific scores and corresponding weights of the selected hardware and environmental influencing factors from the aforementioned factors, the corresponding hardware and environmental scores can be calculated. The final score can then be determined based on these scores, specifically through weighted summation or average. A relationship table between the final score and the configuration failure rate can be pre-built and maintained; this table can be used to determine the configuration failure rate. Alternatively, a formula for estimating the configuration failure rate can be set, such as Configuration Failure Rate = 1 - Overall Score / 10. By scientifically and rationally dividing the scoring ranges for each factor and selecting the corresponding environmental and hardware influencing factors to calculate the final configuration failure rate, the configuration failure rate of remote data centers can be assessed more accurately.

[0053] 203. Determine the overall failure rate of each remote data center using historical failure rates and configuration failure rates.

[0054] In this step, the comprehensive failure rate of each off-site machine room is calculated by combining the historical failure rate and the configuration failure rate. The comprehensive failure rate can be calculated by weighted average, for example, the comprehensive failure rate of the computer room C is calculated using the weighted average method: comprehensive failure rate = a x historical failure rate + (1-a) x configuration failure rate. Wherein, a is a weight coefficient, used to represent the importance of the historical failure rate, and correspondingly, 1-a is used to represent the importance of the configuration failure rate. The weight coefficient can be calculated by the correlation algorithm such as Pearson correlation coefficient to calculate the relationship strength between the historical failure rate and the configuration failure rate. If they are highly positively correlated, it means that the relationship between the historical failure rate and the configuration failure rate is strong or positive, which means that the more the number of failures that have occurred in the past, the higher the possibility of failure under the current hardware configuration and environmental conditions. Therefore, in this case, it can be considered that the configuration failure rate is an important factor for predicting future failures, which can well reflect the potential risk, so the configuration failure rate can be given a larger weight (i.e. smaller a), and the historical failure rate is given a smaller weight. On the contrary, if they are not correlated or highly negatively correlated, it means that the relationship between the historical failure rate and the configuration failure rate is weak or negative, in which case the historical failure rate may not be able to predict future failure risks, or the current hardware configuration and environmental conditions have been significantly improved, so that the past failure mode is no longer applicable. At this time, the historical failure rate should be given a larger weight, and the configuration failure rate should be given a smaller weight (i.e. larger a). It can also be determined according to the risk tolerance of the enterprise or the service level agreement (SLA), such as for those enterprises that pay more attention to preventing potential problems, they may tend to give the configuration failure rate a higher weight (i.e. smaller a) in order to take measures to avoid risks in advance, while for those enterprises that are willing to accept a certain degree of uncertainty in exchange for cost savings, they may choose a larger a. If the enterprise promises a specific service level to the outside, the setting of a should help to ensure that these SLAs can be met. For example, if a certain SLA emphasizes high availability, the importance of the configuration failure rate should be appropriately increased to enhance the fault tolerance of the system. The comprehensive failure rate can also be determined according to the comparison result and the specific difference between the historical failure rate and the configuration failure rate, for example, when the difference between the historical failure rate and the configuration failure rate is small, the larger failure rate of the historical failure rate and the configuration failure rate can be selected as the final comprehensive failure rate. This processing avoids unnecessary complex calculation, and selects a higher failure rate to ensure a more conservative estimate, which can improve the reliability of the system. When the difference between the historical failure rate and the configuration failure rate is large, the smaller failure rate of the historical failure rate and the configuration failure rate can be selected as the superposition failure rate, and a superposition rule is defined to process the superposition failure rate and the larger failure rate to obtain the final comprehensive failure rate.Such processing considers the dual influence of historical failure rate and configuration failure rate, provides a more comprehensive risk assessment, and through superposition, can more accurately reflect the influence of the current configuration on the failure rate. After determining the comprehensive failure rate, it is also stored in the system database for subsequent steps to use.

[0055] For the above description of determining the comprehensive failure rate according to the size comparison result and the specific difference between the two. The specific execution process of determining the comprehensive failure rate of each off-site machine room using the historical failure rate and the configuration failure rate is as follows: if the historical failure rate exceeds the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is within the preset difference interval, the historical failure rate is taken as the comprehensive failure rate; if the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is within the preset difference interval, the configuration failure rate is taken as the comprehensive failure rate; if the historical failure rate exceeds the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, the configuration failure rate is taken as the superposition failure rate, and the comprehensive failure rate is calculated based on the superposition failure rate and the historical failure rate; if the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, the historical failure rate is taken as the superposition failure rate, and the comprehensive failure rate is calculated based on the superposition failure rate and the configuration failure rate.

[0056] In this step, the difference between the historical failure rate and the configuration failure rate of all machine rooms can be calculated in advance, and a distribution diagram of the difference is drawn, for example, the difference distribution is as follows: 0.0002: 2 times; 0.0003: 1 time; -0.0002: 1 time. According to the distribution of the difference, a reasonable interval is selected, which can be a range around the median or average of the difference. For example, the median of the difference is 0.0002, and [0.0001, 0.0003] can be selected as a reasonable difference interval. After determining the difference interval, the difference between the historical failure rate and the configuration failure rate can be calculated and compared with the difference interval to determine whether it is within the difference interval. When it is within the difference interval, it means that the difference between the two is not large, and the larger one of the two can be selected as the comprehensive failure rate, while when it is not within the difference interval, the smaller one of the two can be considered as an additional risk factor, that is, as a superposition failure rate, which can be superimposed with the larger one of the two using addition or multiplication, etc. to form a more accurate comprehensive failure rate. For example, using addition: comprehensive failure rate = historical failure rate + configuration failure rate, or using multiplication: comprehensive failure rate = historical failure rate x (1 + configuration failure rate).

[0057] 204、Based on the off-site machine room and the local machine room with a comprehensive failure rate less than a preset failure rate threshold, a machine room list corresponding to the agent service is generated.

[0058] In this step, a failure rate threshold can be set in advance according to the system requirements and tolerances. System requirements include but are not limited to high availability, low latency, etc., and tolerances include but are not limited to fault tolerance, performance tolerance, recovery time, etc. Among them, high availability means that the system needs to maintain normal operation in most cases, even if some components fail, low latency means that the system needs to respond quickly to user requests and reduce waiting time, fault tolerance means the maximum failure rate that the system can tolerate, i.e. the maximum number of failures allowed within a certain time, performance tolerance means the performance degradation of the system when a fault occurs, i.e. the maximum performance degradation allowed, and recovery time is the time required for the system to recover from a fault to normal operation. Specifically, according to system requirements and business requirements, a target failure rate threshold is set, for example, the target failure rate threshold is 0.0005, i.e. a maximum of 5 failures per 1000 hours, the average failure rate and the maximum failure rate of each machine room are calculated, and the fault tolerance and performance tolerance of the system are considered to ensure that the set failure rate threshold can meet the business requirements and ensure the high availability and reliability of the system. For example, the maximum failure rate that the system can tolerate is 0.0005, but in the case of high load, the performance degradation cannot exceed 10%.

[0059] From all the off-site machine rooms, the off-site machine rooms whose comprehensive failure rate is less than or equal to the preset failure rate threshold are selected as the off-site machine rooms corresponding to the local machine room proxy service, for example, assuming that the comprehensive failure rate of machine room C is 0.000489, which is less than the threshold 0.0005, then machine room C meets the conditions. A machine room list is generated based on the selected off-site machine rooms and the local machine room, and stored in the configuration file or system database of the proxy service.

[0060] It should be noted that in this embodiment, one proxy service corresponds to one dedicated machine room list, which can ensure that each proxy service can access the machine room that best meets its own needs, thereby improving the reliability of the system and ensuring the continuity and availability of the business application service.

[0061] 205、When the proxy service of the target machine room receives the configuration information of the application service, the machine room to be cached is determined from the machine room list corresponding to the proxy service.

[0062] This step combines the description of step 101 in the above method, and the same content will not be repeated here. It should be noted that the specific execution process of determining the machine room to be cached from the machine room list corresponding to the proxy service is: obtaining the geographic location and recent load situation of each machine room in the machine room list; determining the machine room to be cached in the machine room list according to the geographic location, recent load situation and comprehensive failure rate.

[0063] Among them, the recent load situation is used to represent the comprehensive load usage rate of each machine room in a preset historical time period.

[0064] In this step, when the proxy service starts, its exclusive machine room list is read from the configuration file or system database of the proxy service, and the geographic location information of each off-site machine room is obtained respectively, including latitude and longitude, city, country, etc. For example, off-site machine room A is located in Beijing, with latitude and longitude (39.9042, 116.4074), off-site machine room B is located in Shanghai, with latitude and longitude (31.2304, 121.4737), etc. According to the geographic location of the local machine room where the proxy service is located, the distance between each off-site machine room and the proxy service is calculated. Specifically, the distance can be calculated using the Haversine formula. The geographic location information is pre-stored in the system database for easy query at any time.

[0065] A historical time period is defined in advance for calculating the comprehensive load usage rate of each off-site machine room, which represents the recent load situation of the off-site machine room. For example, the historical time period is selected as the past 24 hours. The load data of each off-site machine room in the past 24 hours is obtained from the monitoring system, including CPU usage rate, memory usage rate, network bandwidth usage rate, etc. A comprehensive load usage rate is calculated by combining CPU usage rate, memory usage rate, network bandwidth usage rate, etc. Specifically, it can be obtained by weighted summation. For example, assuming that the weights of CPU usage rate, memory usage rate and network bandwidth usage rate are 0.4, 0.3 and 0.3 respectively, the average CPU usage rate of machine room A in the past 24 hours is 60%, the memory usage rate is 70%, and the network bandwidth usage rate is 50%, then the comprehensive load usage rate = 60% x 0.4 + 70% x 0.3 + 50% x 0.3 = 60%.

[0066] The comprehensive failure rate of each off-site machine room is obtained from the system database, and the geographic location and recent load situation are determined as described above. The specific execution process for determining the to-be-cached machine room in the machine room list is as follows: according to the pre-set usage requirements, the upper limit of the number of off-site machine rooms and the membership matrix are determined; the membership degrees corresponding to the geographic location, recent load situation and comprehensive failure rate are determined respectively according to the membership matrix, and the comprehensive score of each off-site machine room is calculated according to the membership degrees and pre-set weight coefficients using the maximum and minimum composition operator; the to-be-cached machine room is determined based on the upper limit of the number and the comprehensive score.

[0067] In this step, different usage requirement types are defined according to the system's needs and resource limitations. The usage requirement types can include high availability requirements, cost requirements and performance requirements. The upper limit of the number of off-site machine rooms and the corresponding membership matrix are set for each type of usage requirement. Specifically, the association between different usage requirement types and the upper limit of the number and the membership matrix can be pre-constructed, and the upper limit of the number and the membership matrix corresponding to different usage requirement types can be determined through the association.

[0068] For the upper limit of the number of off-site machine rooms, high availability requirements: multiple off-site machine rooms (such as 3-5) need to be selected, and the upper limit is set to 3 to ensure that it can still work normally even if multiple machine rooms fail; cost requirements: select fewer off-site machine rooms (such as 1-2), set the upper limit to 2 to reduce deployment and maintenance costs; performance requirements: select the closest machine room (possibly only one off-site machine room), set the upper limit to 1 to reduce network latency. On this basis, the upper limit can also be dynamically adjusted according to the real-time state (such as the current machine room failure situation). For example, if 2 machine rooms have failed, temporarily increase the upper limit.

[0069] For the membership matrix, the membership matrix is used to represent the membership of each off-site machine room under different use requirements, so as to filter out off-site machine rooms that meet the corresponding use requirements in the machine room list. The membership matrix is an n x m matrix, n is the number of machine rooms (such as local machine room and N off-site machine rooms), and m is the number of evaluation factors (such as geographical location membership, load membership, and failure rate membership). Each element r ij represents the membership of the i-th machine room to the j-th factor. The value of the element ranges from 0 to 1, and the higher the value, the higher the "fitness" of the machine room on that factor. The geographical location membership represents the "fitness" of a certain off-site machine room to the selection of the local machine room, which can be determined according to the distance between each off-site machine room and the proxy service. For example, the closer the distance, the higher the membership (close to 1), the farther the distance, the lower the membership (close to 0), and the following formula can be used: geographical location membership = 1 / distance + 1. The recent load condition membership represents the "fitness" of the current load of the machine room to the selection of the machine room, which can be calculated according to the comprehensive load usage rate. For example, the lower the load, the higher the membership (close to 1), the higher the load, the lower the membership (close to 0), and the following formula can be used: load membership = 1-comprehensive load usage rate. The comprehensive failure rate membership represents the "reliability" of the machine room failure risk to the selection of the machine room, which can be calculated according to the comprehensive failure rate. For example, the lower the failure rate, the higher the membership (close to 1), the higher the failure rate, the lower the membership (close to 0). The following formula can be used: failure rate membership = 1-comprehensive failure rate. At the same time, according to the demand and priority of the system, the weight coefficients of each factor are defined. For example, the weight of geographical location is 0.3, the weight of load is 0.4, and the weight of failure rate is 0.3. According to the membership and the preset weight coefficients, the maximum and minimum synthesis operator is used to calculate the comprehensive score of each off-site machine room.

[0070] It should be noted that the maximum minimum composition operator is an operation method in fuzzy logic, which is used to combine the membership degrees of multiple factors with weights to obtain a comprehensive score. The core idea is: maximum (Max) operation: used to select the factor with higher membership degree, representing "satisfying at least one condition"; minimum (Min) operation: used to select the factor with lower membership degree, representing "all conditions need to be met". In this step, the weighted minimum / maximum combined membership degree and weight are usually used. Specifically, w 地理 , w 负载 , w 故障率 respectively represent the weight coefficients of each factor, μ 地理 , μ 负载 , μ 故障率 respectively represent the membership degrees of each factor, and the specific expression of using the maximum minimum composition algorithm is:

[0071] When taking the minimum value:

[0072] Comprehensive score = min(w 地理 · μ 地理 , w 负载 · μ 负载 , w 故障率 · μ 故障率 );

[0073] Or when taking the maximum value:

[0074] Comprehensive score = max(w 地理 · μ 地理 , w 负载 · μ 负载 , w 故障率 · μ 故障率 ). After determining the comprehensive scores of each off-site machine room through the above expression, the off-site machine rooms are sorted according to the comprehensive scores, and the top N machine rooms with the highest comprehensive scores are selected as the to-be-cached machine rooms according to the upper limit of the number, N is not greater than the upper limit of the number and is a positive integer.

[0075] It should be noted that since the concepts of "high availability", "low cost" and "high performance" in this embodiment are difficult to describe with precise numerical values, the membership matrix is used to quantify these fuzzy concepts, and multiple factors (such as geographic location, recent load situation, comprehensive failure rate, etc.) are considered at the same time, so that the final decision to determine the to-be-cached machine room through the comprehensive score is more scientific, reasonable, comprehensive and accurate.

[0076] 206, write the configuration information into the cache container in the to-be-cached machine room in a specified data structure, so as to facilitate subsequent query and read.

[0077] This step combines the description of step 102 in the above method, and the same content will not be repeated here.

[0078] Further, in combination with the above Figures 1-2 To achieve the method embodiments shown, the embodiments of the present application provide a distributed data storage system schematic diagram based on multiple machine rooms. As shown in the specific Figure 5 The system adopts a distributed architecture and contains multiple machine rooms (such as machine room 1, machine room 2, and machine room n), each of which has the same component structure, and high availability and fault tolerance are achieved through cross-regional data synchronization. The overall layout is divided into three main parts: the top part: proxy service and business application service, responsible for data interaction and registration; the middle part: Redis database and data structure, storing configuration information; the bottom part: read business application service and UI interface, responsible for data reading and display. The specific interaction process is as follows:

[0079] (1) Proxy service and business application service

[0080] Proxy service: responsible for writing configuration data of business application service into local machine room Redis, and randomly selecting other off-site machine room Redis for data synchronization. The arrow direction (from right to left) indicates that the data flow flows from the business application service to the proxy service.

[0081] Business application service: provides machine room information, configuration, theme, environment, version number, digital tag, and other data to the proxy service, and completes the registration and write operation of the configuration through the proxy service.

[0082] (2) Redis database and data structure

[0083] Storage type:

[0084] String type: stores simple configuration in the form of key-value pair, for example: string key-1→value-1.

[0085] Hash type: stores complex configuration in table form, for example: |field-1|value-1|.

[0086] Data synchronization mechanism: the proxy service randomly selects other off-site machine room Redis for data write through the dashed arrow, ensuring data redundancy.

[0087] (3) Reading and display process

[0088] Read business application service: reads data from local or off-site Redis, and ensures data consistency through merging processing. Connects to the UI interface through the arrow, emphasizing the importance of reading operation.

[0089] UI interface: displays the data results after reading, providing a front-end interface for user and system interaction.

[0090] The system realizes high-availability management of cross-regional configuration data through a distributed architecture and efficient storage of Redis. The proxy service, as the core coordination component, ensures the redundancy and synchronization of data in multiple machine rooms, while the merging mechanism of the read business application service and the intuitive display of the UI interface further ensure the stability of the system and the user experience. All designs are centered around "data redundancy", "failover", and "scalability", which meet the typical needs of distributed systems.

[0091] Further, as an implementation of the method embodiment described above Figures 1-2 , the embodiment of the present application provides a distributed data storage device based on multiple machine rooms, which is used to realize multiple machine room copy storage of data, so that the stored data can be read from multiple machine rooms, thereby improving the continuity and availability of business application services. The embodiment of the device corresponds to the aforementioned method embodiment. For ease of reading, the details of the aforementioned method embodiment will not be described one by one, but it should be clear that the device in this embodiment can correspondingly implement all the contents of the aforementioned method embodiment.

[0092] Specifically, as shown in Figure 3 , the device comprises:

[0093] The first determination unit 31 is configured to determine a to-be-cached machine room from a machine room list corresponding to the proxy service when the proxy service of the target machine room receives configuration information of an application service, wherein the machine room list is constructed based on the failure rate of each machine room, and the to-be-cached machine room includes a local machine room and at least one remote machine room.

[0094] The processing unit 32 is configured to write the configuration information into a cache container in the to-be-cached machine room obtained by the first determination unit 31 in a specified data structure.

[0095] Further, as shown in Figure 4 , the device further comprises:

[0096] The acquisition unit 33 is configured to acquire historical failure rates of each remote machine room before the determination unit 31.

[0097] The prediction unit 34 is configured to predict configuration failure rates of each remote machine room according to hardware configurations and environmental conditions of each remote machine room.

[0098] The second determination unit 35 is configured to determine comprehensive failure rates of each remote machine room by using the historical failure rates obtained by the acquisition unit 33 and the configuration failure rates obtained by the prediction unit 34.

[0099] The generating unit 36 is configured to generate the machine room list corresponding to the proxy service based on the off-site machine room and the local machine room whose comprehensive failure rate obtained by the second determining unit 35 is less than the preset failure rate threshold.

[0100] Further, as shown in Figure 4

[0101] The hardware configuration at least includes one of the number of CPU cores, CPU frequency, memory capacity and read-write speed, and the environment condition at least includes one of temperature, humidity and power supply state.

[0102] Further, as shown in Figure 4

[0103] The first determining module 341 is configured to determine the hardware score corresponding to the hardware configuration and the environment score corresponding to the environment condition by using a preset scoring system, wherein the scoring system is constructed based on the influence degree of the hardware configuration and the environment condition on the failure of the machine room.

[0104] The calculating module 342 is configured to calculate the configuration failure rate of each off-site machine room according to the hardware score and the environment score obtained by the first determining module 341.

[0105] Further, as shown in Figure 4

[0106] If the historical failure rate exceeds the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is within the preset difference interval, the historical failure rate is taken as the comprehensive failure rate.

[0107] If the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is within the preset difference interval, the configuration failure rate is taken as the comprehensive failure rate.

[0108] If the historical failure rate exceeds the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, the configuration failure rate is taken as the superimposed failure rate, and the comprehensive failure rate is calculated based on the superimposed failure rate and the historical failure rate.

[0109] If the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, the historical failure rate is taken as the superimposed failure rate, and the comprehensive failure rate is calculated based on the superimposed failure rate and the configuration failure rate.

[0110] Further, as shown in​​​Figure 4 As shown, the first determining unit 31 includes:

[0111] The acquisition module 311 is used to acquire the geographical location and recent load information of each data center in the data center list. The recent load information is used to characterize the overall load utilization rate of each data center within a preset historical time period.

[0112] The second determining module 312 is used to determine the data center to be cached in the data center list based on the geographical location, recent load status and comprehensive failure rate obtained by the obtaining module 311.

[0113] Furthermore, such as Figure 4 As shown, the second determining module 312 is specifically used for,

[0114] The maximum number of remote data centers and the membership matrix are determined based on preset usage requirements, which include high availability requirements, cost requirements, and performance requirements.

[0115] The membership degree corresponding to the geographical location, the recent load status and the comprehensive failure rate are determined according to the membership degree matrix, and the comprehensive score of each remote data center is calculated using the maximum and minimum composition operator based on the membership degree and the preset weight coefficient.

[0116] The data centers to be cached are determined based on the upper limit of the quantity and the comprehensive score.

[0117] Furthermore, embodiments of this application also provide a storage medium for storing a computer program, wherein the computer program, when running, controls the device where the storage medium is located to execute the above-described... Figures 1-2 The distributed data storage method based on multiple data centers described in the article.

[0118] Furthermore, embodiments of this application also provide a processor for running a program, wherein the program executes the above-described... Figures 1-2 The distributed data storage method based on multiple data centers described in the article.

[0119] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0120] It is understood that the relevant features in the above methods and apparatus can be referenced interchangeably. Furthermore, the terms "first," "second," etc., in the above embodiments are used to distinguish between embodiments and do not represent the superiority or inferiority of any particular embodiment.

[0121] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, the device and the unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0122] The algorithms and displays presented herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with programs in accordance with the teachings herein, or it can prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description above. In addition, the present application is not intended to be limited to any particular programming language. It will be appreciated that there are many programming languages that can be used to implement the teachings of the present application as described herein, and any such programming language can be used in connection with the various embodiments.

[0123] In addition, the storage can include non-transitory storage such as a storage chip, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory, among others.

[0124] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer program code.

[0125] The present application is described in reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 The flowchart illustrations and / or block diagrams in accordance with the embodiments of the application have been described above. Figure 1 The means can comprise any apparatus or means effecting the functions specified in the flowchart illustrations and / or block diagrams block or blocks.

[0126] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 The flowchart illustrations and / or block diagrams in accordance with the embodiments of the application have been described above.Figure 1 the function(s) specified in the block or blocks.

[0127] These computer program instructions can also be loaded into computer or other programmable data processing devices to cause a series of operational steps to be performed on the computer or other programmable devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable devices provide steps for implementing the flowchart block(s) or flowchart flow(s) and / or portions thereof. Figure 1 the flowchart block(s) or flowchart flow(s) and / or portions thereof. Figure 1 the function(s) specified in the block or blocks.

[0128] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0129] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) about which the computer stores information such as computer program code and data structures; and non-volatile memory, such as read only memory (ROM), EPROM, EEPROM, or flash memory, about which the computer stores information, such as computer program code and data structures. Memory is an example of computer readable media.

[0130] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.

[0131] It should also be noted that the terms "comprising," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements in the list, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0132] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code thereon for use by or in connection with an instruction execution system. For the purposes of this description, a computer-usable or computer readable storage medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0133] The foregoing is merely illustrative of the embodiments of this application, and is not intended to limit the application. Numerous variations and modifications can be possible to the embodiments without departing from the spirit and scope of the application. Any equivalent modifications or variations, made within the spirit and scope of the application, should be considered within the scope of the application.

Claims

1. A multi-machine room based distributed data storage method, characterized in that, The method comprises: When the proxy service of the target machine room receives the configuration information of the application service, determining a to-be-cached machine room from a machine room list corresponding to the proxy service, the machine room list being constructed based on failure rates of respective machine rooms, the to-be-cached machine room including a local machine room and at least one off-site machine room; writing the configuration information into a cache container in the to-be-cached machine room in a specified data structure.

2. The method of claim 1, wherein, Before the to-be-cached machine room is determined from the machine room list corresponding to the proxy service, the method further comprises: obtaining historical failure rates of respective off-site machine rooms; predicting configuration failure rates of respective off-site machine rooms according to hardware configurations and environmental conditions of the respective off-site machine rooms; determining comprehensive failure rates of the respective off-site machine rooms by using the historical failure rates and the configuration failure rates; generating the machine room list corresponding to the proxy service based on off-site machine rooms and the local machine room whose comprehensive failure rates are less than a preset failure rate threshold.

3. The method of claim 2, wherein: the hardware configurations at least include one of a CPU core number, a CPU frequency, a memory capacity, and a read-write speed, and the environmental conditions at least include one of a temperature, a humidity, and a power supply state.

4. The method of claim 3, wherein, The configuration failure rates of the respective off-site machine rooms are predicted according to the hardware configurations and the environmental conditions of the respective off-site machine rooms, comprising: determining hardware scores corresponding to the hardware configurations and environmental scores corresponding to the environmental conditions by using a preset scoring system, the scoring system being constructed based on respective influences of the hardware configurations and the environmental conditions on machine room failure; calculating the configuration failure rates of the respective off-site machine rooms according to the hardware scores and the environmental scores.

5. The method of claim 1, wherein, The comprehensive failure rates of the respective off-site machine rooms are determined by using the historical failure rates and the configuration failure rates, comprising: if the historical failure rate exceeds the configuration failure rate and a difference between the historical failure rate and the configuration failure rate is within a preset difference interval, taking the historical failure rate as the comprehensive failure rate; if the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is within the preset difference interval, taking the configuration failure rate as the comprehensive failure rate; if the historical failure rate exceeds the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, taking the configuration failure rate as a superimposed failure rate, and calculating the comprehensive failure rate based on the superimposed failure rate and the historical failure rate; if the historical failure rate does not exceed the configuration failure rate and the difference between the historical failure rate and the configuration failure rate is not within the preset difference interval, taking the historical failure rate as a superimposed failure rate, and calculating the comprehensive failure rate based on the superimposed failure rate and the configuration failure rate.

6. The method according to any one of claims 2-5, characterized in that, The to-be-cached machine room is determined from the machine room list corresponding to the proxy service, comprising: obtaining geographical positions and recent load conditions of respective machine rooms in the machine room list, the recent load conditions being used to represent comprehensive load usage rates of respective machine rooms in a preset historical time period; The to-be-cached machine room is determined in the machine room list according to the geographic location, the recent load condition and the comprehensive failure rate.

7. The method of claim 6, wherein, The to-be-cached machine room is determined in the machine room list according to the geographic location, the recent load condition and the comprehensive failure rate. A number upper limit and a membership matrix of the off-site machine rooms are determined according to preset use requirements, the use requirements including high availability requirements, cost requirements and performance requirements; Membership degrees corresponding to the geographic location, the recent load condition and the comprehensive failure rate respectively are determined according to the membership matrix, and a comprehensive score of each off-site machine room is calculated according to the membership degrees and preset weight coefficients by using a maximum minimum synthesis operator; The to-be-cached machine room is determined based on the number upper limit and the comprehensive score.

8. A multi-machine room based distributed data storage apparatus, characterized by, The apparatus comprises: A first determining unit configured to determine a to-be-cached machine room from a machine room list corresponding to a proxy service of a target machine room when the proxy service receives configuration information of an application service, the machine room list being constructed based on failure rates of machine rooms, the to-be-cached machine room including a local machine room and at least one off-site machine room; A processing unit configured to write the configuration information into a cache container in the to-be-cached machine room obtained by the first determining unit in a specified data structure.

9. A storage medium, characterized by The storage medium includes a stored program, wherein the program, when executed, controls a device in which the storage medium is located to perform the multi-machine room-based distributed data storage method of any one of claims 1 to 7.

10. A processor, comprising: The processor is configured to execute a program, wherein the program, when executed, performs the multi-machine room-based distributed data storage method of any one of claims 1 to 7.