Machine room management method and management device based on digital twinning
By building a digital twin model of the computer room and real-time data mapping, the problem of inefficiency in traditional computer room management is solved, real-time monitoring and intelligent operation and maintenance of the computer room operation status is realized, and the stability and reliability of the computer room are improved.
Patent Information
- Application Number
- CN202510819717.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Traditional computer room management methods rely on manual inspection and regular maintenance, which are inefficient, difficult to grasp the operating status of the equipment in real time, and cannot detect potential faults in a timely manner. The automated monitoring system lacks the ability to comprehensively analyze and early warning of the overall operating status of the computer room.
Build a digital twin model of the computer room, establish a state mapping relationship by obtaining real-time perceptual data, simulate the operating status of the equipment, conduct abnormal characteristics analysis, generate operation and maintenance adjustment instructions to drive equipment status adjustments, and realize intelligent operation and maintenance.
It improves the real-time and accuracy of computer room management, can quickly determine abnormal equipment and its impact range, improves the stability and reliability of computer room, and realizes fault prevention and intelligent operation and maintenance.
Smart Images

Figure CN120338770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital twins, and more particularly, to a computer room management method and management device based on digital twins. Background Art
[0002] With the rapid development of information technology, as the core place for data storage and processing, the stability and reliability of computer rooms are of crucial importance. Traditional computer room management methods mainly rely on manual inspections and regular maintenance. This method is not only inefficient but also difficult to grasp the operating status of each device in the computer room in real time, and it is impossible to detect and handle potential fault hazards in a timely manner. Although some computer rooms have started to introduce automated monitoring systems, these systems often can only provide real-time data of the devices, lacking the comprehensive analysis and simulation capabilities of the overall operating status of the computer room, and it is difficult to give early warnings and interventions before faults occur. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, the present invention provides a computer room management method based on digital twins, and the method includes: Construct a digital twin model of the physical entity of the computer room, where the digital twin model includes the spatial layout structure information of each device in the computer room and the associated relationship information of the operating parameters between devices; Obtain the real-time perception data of each device in the computer room, and establish a state mapping relationship based on the real-time perception data, the spatial layout structure information, and the operating parameter association relationship information in the digital twin model; Simulate the operating status of the device in the digital twin model according to the state mapping relationship, and generate a dynamic operation representation reflecting the real-time operation of the physical computer room; Perform abnormal feature analysis and processing on the dynamic operation representation to determine the abnormal associated devices and the influence range of the abnormal associated devices; Generate an operation and maintenance adjustment instruction according to the abnormal associated devices and the influence range, and drive the corresponding devices in the physical computer room to perform operation status adjustment operations through the operation and maintenance adjustment instruction.
[0004] In another aspect, the present invention further provides a computer room management device, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0005] Based on the above aspects, the present invention realizes accurate modeling of the spatial layout structure and the correlation relationship of operation parameters of each device in the computer room by constructing a digital twin model of the physical entities in the computer room. Further, by obtaining the real-time perception data of each device in the computer room and establishing a state mapping relationship in the digital twin model, it is possible to simulate the operation state of the computer room in real time, generate a dynamic operation representation reflecting the real-time operation of the physical computer room, which not only improves the timeliness and accuracy of computer room management, but also provides an intuitive and comprehensive decision-making basis for the operation and maintenance management of the computer room. By performing abnormal feature analysis and processing on the dynamic operation representation, it is possible to quickly determine the abnormal associated devices and their influence ranges, and then generate operation and maintenance adjustment instructions to drive the corresponding devices in the physical computer room to perform operation state adjustment operations, realizing intelligent operation and maintenance and fault prevention of the computer room, and significantly improving the stability and reliability of the computer room. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 It is a schematic flowchart of the execution process of the computer room management method based on digital twin provided by an embodiment of the present invention.
[0007] Figure 2 It is a schematic diagram of exemplary hardware and software components of the computer room management device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0008] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 It is a schematic flowchart of the computer room management method based on digital twin provided by an embodiment of the present invention. The computer room management method based on digital twin will be introduced in detail below.
[0009] Step S110: Construct a digital twin model of the physical entities in the computer room, where the digital twin model includes the spatial layout structure information of each device in the computer room and the correlation relationship information of the operation parameters between devices.
[0010] In this embodiment, taking a comprehensive enterprise computer room as an example, which includes various types of devices such as servers, storage devices, network devices, and power supply devices. To construct the digital twin model, relevant information of each device in the computer room needs to be obtained. For specific details, please refer to the following steps.
[0011] Step S111: Collect spatial data of the physical entities in the computer room, and obtain the installation position coordinates of each device, the device shape size parameters, and the route information of the connection lines between devices as the spatial layout structure information.
[0012] In the enterprise computer room, first, the three-dimensional laser scanning technology is used to scan the entire computer room space. Three-dimensional laser scanning can accurately obtain the three-dimensional space information of each device in the computer room. For each device, such as a server, the scanning device will record its installation position coordinates in the three-dimensional coordinate system of the computer room. Assuming that the installation position coordinates of the server are represented by a set of coordinate values (X, Y, Z), the coordinate values of different servers will vary according to their actual positions in the computer room. At the same time, the scanning device will also obtain the external dimension parameters of the server, including information such as length, width, and height. These parameters are recorded in a set measurement unit to ensure the consistency of dimensions. For the connection lines between devices, a line tracing device is used for detection. The line tracing device can trace along connection lines such as network cables and power cables and record the path information of the lines. For example, for the network cable connection between a server and a switch, the line tracing device will record the coordinates of each path point that the network cable passes through starting from the interface of the server and finally reaching the interface of the switch. The coordinate information of these path points constitutes the path information of the connection line.
[0013] Step S112: Conduct a correlation analysis on the historical operation data of each device in the computer room, and extract the variation rules of the device operation parameters. The operation parameters include device temperature parameters, power consumption parameters, and data processing rate parameters.
[0014] To accurately extract the variation rules of the device operation parameters, it is necessary to conduct a comprehensive and in-depth analysis of the historical operation data of each device in the computer room. First, collect various historical operation data of the devices during continuous operation cycles.
[0015] Step S1121: Collect the historical records of temperature parameters, power consumption parameters, and data processing rate parameters of each device in the computer room during continuous operation cycles.
[0016] In the enterprise computer room, through temperature sensors installed on the surface and inside of the devices, the temperature data of the devices is collected in real time and stored in the data storage system to form historical records of temperature parameters. Similarly, through power consumption monitoring devices, the power consumption data of the devices during operation is recorded to construct historical records of power consumption parameters. For the data processing rate parameters, using the built-in performance monitoring tools or network traffic monitoring devices of the devices, collect the data processing rate information of the devices to form historical records of data processing rate parameters. The above historical records will be arranged in chronological order for subsequent analysis and processing.
[0017] Step S1122: Conduct a time series analysis on the historical records of temperature parameters to identify the variation trend of temperature parameters with the operation duration of the devices. This variation trend includes the time nodes of the temperature rising stage, stable stage, and falling stage.
[0018] After obtaining the historical record of temperature parameters, it is processed using time series analysis methods. By analyzing the curve of temperature data changing over time, the changing trend of the device temperature can be identified. During the device startup phase, the temperature usually rises with the increase in the running duration, forming a temperature rising phase. When the device reaches the set stable operating state, the temperature fluctuates within a relatively stable range, which is the temperature stable phase. When the device stops running or an abnormal situation occurs, the temperature gradually decreases, entering the temperature decreasing phase. By analyzing the historical data, the time nodes of these phases can be determined. For example, through the statistical analysis of a large amount of historical data, it is found that after a certain model of server is normally started, it enters the temperature stable phase after a period of time T1, and after it stops running, the temperature drops to near the ambient temperature after a time T2.
[0019] Step S1123: Conduct a correlation analysis on the historical record of power consumption parameters, calculate the correlation coefficient between the power consumption parameters and the device data processing rate parameters, and determine the linear or non-linear relationship between the power consumption parameters and the change in the data processing rate.
[0020] When conducting a correlation analysis on the historical record of power consumption parameters and the historical record of data processing rate parameters, a set statistical analysis method is adopted. First, pair the power consumption data and the data processing rate data at the same time point to form data pairs. Then, use the correlation coefficient calculation method to calculate the correlation coefficient between the power consumption parameters and the data processing rate parameters. If the correlation coefficient is close to 1 or -1, it indicates that there is a strong linear relationship between the two; if the correlation coefficient deviates from 1 or -1, it is necessary to further analyze whether there is a non-linear relationship. For example, through the calculation and analysis of a large amount of data, it is found that there is a linear relationship between the power consumption parameters and the data processing rate parameters of a certain type of server, and it can be represented by a linear equation, where the power consumption increases linearly with the increase in the data processing rate.
[0021] Step S1124: Conduct a periodic analysis on the historical record of data processing rate parameters, and extract the peak frequency and valley frequency of the data processing rate in different business time periods.
[0022] When performing periodic analysis on the historical records of data processing rate parameters, methods such as spectral analysis are adopted. First, the data processing rate data is arranged in chronological order, and then a spectral analysis tool is used to process the data to find the periodic components in the data. By analyzing these periodic components, the peak frequency and valley frequency of the data processing rate in different business time periods can be determined. For example, during the day on weekdays in an enterprise, business activities are relatively frequent, and there may be multiple peaks in the data processing rate. By analysis, the frequencies at which these peaks occur can be determined. While at night or on holidays, business activities are relatively less, and the data processing rate will show valleys, and the frequencies at which the valleys occur can also be determined.
[0023] Step S1125: Integrate the change trend of the temperature parameter, the correlation between the power consumption parameter and the data processing rate, and the periodic characteristics of the data processing rate to generate a description of the change law of the device operation parameters. This description of the change law includes the time-dependent characteristics of parameter changes and the mutual influence characteristics between parameters.
[0024] After completing the individual analysis of the temperature parameter, the power consumption parameter, and the data processing rate parameter, it is necessary to integrate these analysis results. First, comprehensively consider the change trend of the temperature parameter, the correlation between the power consumption parameter and the data processing rate, and the periodic characteristics of the data processing rate. For example, when the data processing rate is at the peak stage, according to the correlation between the power consumption parameter and the data processing rate, it can be predicted that the power consumption of the device will increase correspondingly; at the same time, due to the increase in power consumption, according to the change trend of the temperature parameter, the temperature of the device may also rise. Through the above comprehensive analysis, a description of the change law of the device operation parameters is generated. This description includes the time-dependent characteristics of parameter changes, that is, the changes of different parameters at different time points, and the mutual influence characteristics between parameters, that is, how the change of one parameter affects the changes of other parameters.
[0025] Step S113: Establish a description of the causal relationship between different device operation parameters. This description of the causal relationship includes the triggering conditions for parameter changes and the information on the transmission direction of parameter changes.
[0026] In an enterprise computer room, there are inter - influencing relationships among the operating parameters of different devices. For example, the operating state of a server affects the network traffic of the switch connected to it, and the network traffic of the switch in turn affects the stability of the entire network. To accurately describe these relationships, it is necessary to establish a causal relationship description between the operating parameters of different devices. First, through the analysis of a large amount of historical data, determine the triggering conditions for parameter changes. For example, when the data processing rate of the server exceeds the set threshold, it can trigger an increase in the network traffic of the switch connected to it. Then, determine the transfer direction information of parameter changes. For example, the change in the data processing rate of the server is transmitted through the network connection to the switch, affecting the network traffic of the switch. Through the above methods, a causal relationship description between the operating parameters of different devices is established.
[0027] Step S114: Integrate and model the correlation relationship information between the spatial layout structure information and the device operating parameters to generate a digital twin model basic framework containing a three - dimensional space coordinate system and parameter correlation rules.
[0028] After obtaining the correlation relationship information between the spatial layout structure information and the device operating parameters, it is necessary to integrate and model these two parts of information. First, correlate the information such as the device installation position coordinates and external dimension parameters in the spatial layout structure information with the correlation relationship information of the device operating parameters. For example, correlate the installation position coordinates of the server with the operating parameters such as the temperature, power consumption, and data processing rate of the server. Then, in the three - dimensional space coordinate system, construct a three - dimensional model of the device according to the device installation position coordinates and external dimension parameters. At the same time, convert the correlation relationship information of the device operating parameters into parameter correlation rules and embed them into the three - dimensional model. For example, when the temperature of the server rises, according to the parameter correlation rules, the cooling power of the air - conditioning device connected to it will increase accordingly. Through the above methods, a digital twin model basic framework containing a three - dimensional space coordinate system and parameter correlation rules is generated.
[0029] Step S115: Conduct a compatibility check on the spatial position information and operating parameter information of newly connected devices through the digital twin model basic framework, and update the spatial layout structure information and operating parameter correlation relationship information of the digital twin model.
[0030] During the operation of an enterprise computer room, new devices may be connected. When a new device is connected, it is necessary to use the basic framework of the digital twin model to perform compatibility verification on the spatial location information and operating parameter information of the new device. First, check whether the spatial location of the new device conflicts with the spatial layout structure of the existing devices. For example, whether the installation location of the new device will affect the heat dissipation or maintenance channels of other devices. Then, check whether the operating parameter information of the new device is compatible with the associated relationship of the operating parameters of the existing devices. For example, whether parameters such as the data processing rate and power consumption of the new device will affect the existing network and power supply systems. If the spatial location information and operating parameter information of the new device pass the compatibility verification, the information of the new device will be added to the digital twin model, and the spatial layout structure information and operating parameter associated relationship information of the digital twin model will be updated. If the compatibility verification fails, it is necessary to adjust the installation location or operating parameters of the new device until the verification is passed.
[0031] Step S120: Obtain the real-time perception data of each device in the computer room, and establish a state mapping relationship based on this real-time perception data and the spatial layout structure information and operating parameter associated relationship information in the digital twin model.
[0032] After constructing the digital twin model, it is necessary to obtain the real-time perception data of each device in the computer room and establish a state mapping relationship between these data and the digital twin model. In an enterprise computer room, the real-time perception data can reflect the current operating state of the device, and the state mapping relationship can accurately map the operating state of the physical device into the digital twin model.
[0033] Step S121: Collect the real-time temperature value, real-time power consumption value, and real-time data processing rate value of the device through the sensor components deployed on the device surface and connection lines as the real-time perception data.
[0034] In an enterprise computer room, in order to accurately obtain the real-time operating state of the device, sensor components are deployed on the device surface and connection lines. For example, a temperature sensor is installed on the outer shell of the server to collect the temperature value of the server in real time; a power consumption sensor is installed on the power line to collect the power consumption value of the server in real time; a traffic sensor is installed at the network interface to collect the data processing rate value of the server in real time. The above sensor components will transmit the collected data to the data acquisition system in real time to form real-time perception data.
[0035] Step S122: Perform format standardization processing on the real-time perception data to generate standardized perception data with the same data format as the operating parameter data in the digital twin model.
[0036] Since the real-time perception data collected by different sensor components may have different data formats, in order to facilitate matching and processing with the operation parameter data in the digital twin model, it is necessary to perform format standardization processing on the real-time perception data. For example, the temperature data collected by temperature sensors of different brands may have different precisions and units, and these data need to be uniformly converted to the precision and units specified in the digital twin model. By performing format standardization processing on the real-time perception data, standardized perception data consistent with the operation parameter data format in the digital twin model is generated.
[0037] Step S123: Extract the spatial layout structure information in the digital twin model, and determine the device identifiers and device spatial position coordinates corresponding to each sensor component.
[0038] Extract the spatial layout structure information from the digital twin model, including information such as the installation position coordinates and external dimension parameters of the device. Then, according to the installation position of the sensor component, determine the device identifier and device spatial position coordinates corresponding to each sensor component. For example, the device identifier corresponding to the temperature sensor installed on Server A is Server A, and its device spatial position coordinates are the installation position coordinates of Server A.
[0039] Step S124: Bind the standardized perception data with the corresponding device identifier and device spatial position coordinates to generate multi-dimensional status data including the device identifier, spatial position coordinates, and standardized perception data.
[0040] After completing the format standardization processing of the real-time perception data and the association between the sensor component and the device, bind the standardized perception data with the corresponding device identifier and device spatial position coordinates. For example, bind the standardized temperature value, standardized power consumption value, and standardized data processing rate value of Server A with the device identifier and device spatial position coordinates of Server A to generate multi-dimensional status data including the device identifier, spatial position coordinates, and standardized perception data of Server A. Through the above method, the real-time perception data is integrated with the spatial information of the device to form more comprehensive multi-dimensional status data.
[0041] Step S125: Based on the operation parameter association relationship information in the digital twin model, establish a parameter mapping rule between the multi-dimensional status data of different devices. This parameter mapping rule includes the parameter synchronization frequency and the parameter change threshold.
[0042] In order to accurately reflect the operation state relationship between different devices, it is necessary to establish a parameter mapping rule between the multi-dimensional status data of different devices based on the operation parameter association relationship information in the digital twin model.
[0043] Step S1251: Extract the causal relationship description of the operating parameters between devices from the operating parameter correlation relationship information of the digital twin model.
[0044] Extract the causal relationship description of the operating parameters between devices from the operating parameter correlation relationship information of the digital twin model. For example, extract the causal relationship description of the operating parameters between the server and the switch from the digital twin model, that is, how the change in the data processing rate of the server affects the network traffic of the switch.
[0045] Step S1252: Determine the synchronization order of the master device parameters and the slave device parameters according to the parameter change trigger condition in the causal relationship description. The master device parameters are the starting parameters that trigger the parameter change, and the slave device parameters are the following parameters affected by the master device parameters.
[0046] Determine the synchronization order of the master device parameters and the slave device parameters according to the parameter change trigger condition in the causal relationship description. For example, in the relationship between the server and the switch, the data processing rate of the server is the master device parameter, and the network traffic of the switch is the slave device parameter. When the data processing rate of the server changes, it can trigger the corresponding change in the network traffic of the switch. According to this causal relationship, determine the synchronization order of the master device parameters and the slave device parameters, that is, the change in the data processing rate of the server occurs first, and the change in the network traffic of the switch occurs later.
[0047] Step S1253: Set the parameter synchronization frequency based on the synchronization order of the master device parameters and the slave device parameters. This parameter synchronization frequency is the time interval required for the slave device parameters to complete the update after the master device parameters are updated.
[0048] After determining the synchronization order of the master device parameters and the slave device parameters, set the parameter synchronization frequency. For example, according to the actual operating conditions of the server and the switch, set that after the data processing rate of the server is updated, the network traffic of the switch needs to be updated within the set time interval. This time interval is the parameter synchronization frequency. By setting the parameter synchronization frequency, ensure that the operating parameters between different devices can be synchronized in time and accurately reflect the operating status of the devices.
[0049] Step S1254: Analyze the parameter change transfer direction information in the causal relationship description, determine the maximum acceptable fluctuation range of the parameter change, and use this maximum acceptable fluctuation range as the parameter change threshold.
[0050] Analyze the parameter change transfer direction information in the cause-and-effect relationship description to determine the maximum acceptable fluctuation range of parameter changes. For example, when the data processing rate of the server changes, analyze the maximum acceptable fluctuation range of the network traffic of the switch. If the network traffic of the switch fluctuates beyond this range, it may affect the network stability. Use this maximum acceptable fluctuation range as the parameter change threshold for subsequent state mapping relationship establishment and anomaly detection.
[0051] Step S1255: Combine the parameter synchronization order, parameter synchronization frequency, and parameter change threshold to generate a parameter mapping rule between multi-dimensional state data of different devices.
[0052] After determining the parameter synchronization order, parameter synchronization frequency, and parameter change threshold, combine these three elements to generate a parameter mapping rule between multi-dimensional state data of different devices. For example, combine the parameter synchronization order, parameter synchronization frequency, and parameter change threshold between the server and the switch to form a parameter mapping rule between the server and the switch. In the above way, an accurate mapping relationship between multi-dimensional state data of different devices is established.
[0053] Step S126: Map the multi-dimensional state data to the corresponding device nodes of the digital twin model according to the parameter mapping rule to generate a state mapping relationship reflecting the one-to-one correspondence between the physical device and the digital twin model node.
[0054] After establishing the parameter mapping rule, map the multi-dimensional state data to the corresponding device nodes of the digital twin model according to this parameter mapping rule. For example, map the multi-dimensional state data of server A to the corresponding node of server A in the digital twin model according to the parameter mapping rule. In the above way, a state mapping relationship reflecting the one-to-one correspondence between the physical device and the digital twin model node is generated. This state mapping relationship can accurately transfer the operating state of the physical device to the digital twin model.
[0055] Step S130: Simulate the device operating state in the digital twin model according to the state mapping relationship to generate a dynamic operation representation reflecting the real-time operation of the physical computer room.
[0056] After establishing the state mapping relationship, the operating state of the device can be simulated in the digital twin model using this relationship to generate a dynamic operation representation reflecting the real-time operation of the physical computer room.
[0057] Step S131: Extract the multi-dimensional state data corresponding to each digital twin model node from the state mapping relationship.
[0058] Extract the multi-dimensional state data corresponding to each digital twin model node from the state mapping relationship. For example, extract the multi-dimensional state data corresponding to server A node in the digital twin model from the state mapping relationship, including information such as the device identifier, spatial position coordinates, standardized temperature value, standardized power consumption value, and standardized data processing rate value of server A.
[0059] Step S132: Input the multi-dimensional state data into the operation simulation module of the corresponding digital twin model node according to the parameter synchronization frequency in the state mapping relationship.
[0060] Input the multi-dimensional state data into the operation simulation module of the corresponding digital twin model node according to the parameter synchronization frequency in the state mapping relationship. For example, according to the parameter synchronization frequency between the server and the switch, input the multi-dimensional state data of the server into the operation simulation module of the server node in the digital twin model within a specified time interval. By the above method, ensure that the device operation state in the digital twin model can keep up with the actual operation state of the physical device in a timely manner.
[0061] Step S133: The operation simulation module performs parameter change trend deduction on the input multi-dimensional state data based on the operation parameter correlation relationship information in the digital twin model, and generates a simulation value of the device operation state.
[0062] After receiving the multi-dimensional state data, the operation simulation module performs parameter change trend deduction on the input multi-dimensional state data based on the operation parameter correlation relationship information in the digital twin model.
[0063] Step S1331: Input the real-time temperature value, real-time power consumption value, and real-time data processing rate value in the multi-dimensional state data into the parameter input layer of the operation simulation module.
[0064] Input the real-time temperature value, real-time power consumption value, and real-time data processing rate value in the multi-dimensional state data into the parameter input layer of the operation simulation module. For example, input the real-time temperature value, real-time power consumption value, and real-time data processing rate value of the server into the parameter input layer of the operation simulation module. The parameter input layer will perform preliminary processing on these input real-time parameter values, such as checking the integrity and accuracy of the data, to ensure that the input data meets the requirements of subsequent processing. Since the real-time perception data has been format standardized before, the real-time temperature value, real-time power consumption value, and real-time data processing rate value input here are consistent with the relevant data stored in the digital twin model in terms of dimension and format, ensuring the accuracy and consistency of subsequent calculations.
[0065] Step S1332: The parameter input layer of the operation simulation module matches the input real-time parameter values with the description of the device operation parameter change law stored in the digital twin model to determine the current stage of device operation. The stages include the temperature rising stage, the stable stage, or the falling stage, the peak stage or the trough stage of the data processing rate.
[0066] After receiving the real-time parameter values, the parameter input layer can match them in detail with the description of the device operation parameter change law stored in the digital twin model. Taking a server as an example, for the real-time temperature value, it can be compared with the characteristics of the temperature rising stage, the stable stage, and the falling stage in the temperature parameter change law description. If the real-time temperature value shows a continuous rising trend and falls within the characteristic range of the temperature rising stage, then it can be determined that the server is currently in the temperature rising stage; if the real-time temperature value fluctuates within a relatively stable range and conforms to the characteristics of the stable stage, then the server is in the stable stage; if the real-time temperature value starts to fall and meets the characteristic conditions of the falling stage, then the server is in the temperature falling stage. For the real-time data processing rate value, it can be matched with the peak stage and the trough stage in the periodic characteristics of the data processing rate. If the real-time data processing rate value reaches or approaches the numerical range corresponding to the peak frequency obtained from previous analysis, then the server is in the peak stage of the data processing rate; if the real-time data processing rate value is within the numerical range corresponding to the trough frequency, then the server is in the trough stage of the data processing rate.
[0067] Step S1333: Based on the device operation stage, call the corresponding change trend model in the description of the device operation parameter change law. The change trend models include the temperature change trend model, the correlation model between power consumption and data processing rate, and the data processing rate periodic model.
[0068] After determining the current operation stage of the device, the operation simulation module will call the corresponding change trend model according to this stage. If the server is in the temperature rising stage, the specific model for the temperature rising stage in the temperature change trend model can be called. This model may be fitted based on a large amount of historical data and can describe the change trend of temperature over time in this stage. For the correlation model between power consumption and data processing rate, if the server is in the peak stage of the data processing rate, since there is a certain correlation between power consumption and data processing rate, the part of the correlation model applicable to the peak stage can be called to more accurately predict the change of power consumption. Similarly, for the data processing rate periodic model, according to the current data processing rate stage of the server, the corresponding periodic characteristic part in the model can be called to predict the future change of the data processing rate.
[0069] Step S1334: Deduce the temperature prediction value at a future time point through the temperature change trend model, deduce the power consumption prediction value at a future time point through the correlation model between power consumption and data processing rate, and deduce the data processing rate prediction value at a future time point through the data processing rate periodic model.
[0070] Using the called change trend models, predict the temperature, power consumption, and data processing rate at a future time point respectively. For temperature prediction, the temperature change trend model will, based on the current real-time temperature value and the temperature change law at this stage, consider the time factor and calculate the temperature prediction value at a future time point. For example, according to the change law during the temperature rising stage, combined with the current temperature rising rate and time interval, predict the temperature value after a period of time. For power consumption prediction, the correlation model between power consumption and data processing rate will, based on the current data processing rate value and the correlation between the two, predict the impact of the change in data processing rate at a future time point on power consumption, thereby obtaining the power consumption prediction value at a future time point. The data processing rate periodic model will, based on the stage where the current data processing rate is located and the periodic characteristics, predict the data processing rate prediction value at a future time point. For example, according to the periodic fluctuation law of the data processing rate, predict the time point when the next peak or trough appears and the corresponding rate value.
[0071] Step S1335: Combine the temperature prediction value, power consumption prediction value, and data processing rate prediction value to generate a simulated value of the device operating state.
[0072] After obtaining the temperature prediction value, power consumption prediction value, and data processing rate prediction value at a future time point, combine these three prediction values to form a simulated value of the device operating state. This simulated value contains information on multiple operating parameters of the device at a future time point and can more comprehensively reflect the operating state of the device. For example, combine the temperature prediction value, power consumption prediction value, and data processing rate prediction value of a server at a certain future moment to form a multi-dimensional simulated value vector, and this simulated value vector represents the simulated operating state of the server at that moment.
[0073] Step S134: Perform consistency verification on the simulated value and the real-time perception data in the multi-dimensional state data. If the verification passes, retain the simulated value; if the verification fails, perform parameter change trend deduction again.
[0074] To ensure the accuracy and reliability of the simulated values, it is necessary to perform consistency verification on the generated simulated values and the real-time perception data in the multi-dimensional state data. Taking the server as an example, the predicted temperature value obtained by simulation can be compared with the real-time temperature value, the predicted power consumption value with the real-time power consumption value, and the predicted data processing rate value with the real-time data processing rate value respectively. During the comparison process, the parameter change threshold set previously can be used to determine whether the difference between the two is within the acceptable range. If the differences between the simulated values and the real-time perception data are all within the parameter change threshold range, it indicates that the simulated values are relatively consistent with the actual situation, and the verification passes. At this time, the simulated values are retained; if the difference between the simulated value of a certain parameter and the real-time perception data exceeds the parameter change threshold, it is considered that the verification fails. In this case, it is necessary to start from step S133 again, perform parameter change trend deduction on the input multi-dimensional state data again, and adjust the relevant parameters and models in the prediction process to obtain more accurate simulated values.
[0075] Step S135: Visualize and render the simulated values that have passed the consistency verification in the three-dimensional space coordinate system of the digital twin model according to the device spatial position coordinates, and generate a three-dimensional visualization scene including the color identification of the device operation state, the dynamic curve of parameter changes, and the connection line state indication.
[0076] After the simulated values pass the consistency verification, they can be visually rendered in the three-dimensional space coordinate system of the digital twin model according to the device spatial position coordinates. Taking devices such as servers and switches in the computer room as an example, for the simulated values of each device, corresponding visual processing can be performed according to its different operating parameter states. For the color identification of the device operation state, different colors can be assigned according to different situations of operating parameters such as the temperature and power consumption of the device. For example, when the temperature of the server is within the normal range, the server may be displayed in green in the three-dimensional visualization scene; if the temperature is close to or exceeds the safety threshold, the server may be displayed in red to intuitively remind the management personnel that there may be abnormalities in the device. For the dynamic curve of parameter changes, the dynamic curves of parameters such as temperature, power consumption, and data processing rate changing with time can be drawn based on the simulated values and real-time perception data of the device. The above curves can visually display the change trend of the device operating parameters in the three-dimensional visualization scene, facilitating the management personnel to analyze and predict. For the connection line state indication, the connection lines can be visually represented according to the connection relationship between devices and network traffic conditions. If the network traffic of the connection line is normal, the line may be displayed in blue; if abnormal situations such as network congestion occur, the line may be displayed in yellow or red. Through the above visual rendering, a three-dimensional visualization scene including the color identification of the device operation state, the dynamic curve of parameter changes, and the connection line state indication is generated, enabling the management personnel to more intuitively understand the real-time operation situation of the physical computer room.
[0077] Step S136: Synchronize the 3D visualization scene with the real-time updated simulation values in a time series to generate a dynamically running representation that continuously changes over time.
[0078] After generating the 3D visualization scene, it is necessary to synchronize it with the real-time updated simulation values in a time series. Since the operating state of the device changes continuously over time, the simulation values will also be updated in real time. To ensure that the 3D visualization scene can accurately reflect the real-time operating state of the device, it is necessary to synchronize the information such as the device operating state, the dynamic curve of parameter changes, and the connection line state in the 3D visualization scene with the real-time updated simulation values. Specifically, at a set time interval, for example, every once in a while (this time interval is related to the parameter synchronization frequency set previously), the latest simulation values are updated into the 3D visualization scene. In this way, the information such as the color identification of the device operating state, the dynamic curve of parameter changes, and the connection line state indication in the 3D visualization scene will change in real time as the simulation values are updated, thereby generating a dynamically running representation that continuously changes over time. This dynamically running representation can reflect the operating state and changes of the devices in the physical computer room in real time and accurately.
[0079] Step S140: Analyze and process the abnormal characteristics of the dynamically running representation to determine the abnormal associated devices and the influence scope of the abnormal associated devices.
[0080] After obtaining the dynamically running representation that reflects the real-time operation of the physical computer room, it is necessary to analyze and process its abnormal characteristics to discover possible abnormal situations during the device operation and determine the abnormal associated devices and the influence scope of the abnormal associated devices.
[0081] Step S141: Extract the time series data of the simulation values of each device and the color identification information in the 3D visualization scene from the dynamically running representation.
[0082] Extract the time series data of the simulation values of each device and the color identification information in the 3D visualization scene from the dynamically running representation. For the time series data of the simulation values, the simulation values of each device at different time points can be extracted in chronological order to form a multi-dimensional time series data set. For example, for a server, the temperature simulation value, power consumption simulation value, and data processing rate simulation value at each moment within a period of time can be extracted, and these values arranged in chronological order constitute the time series data of the simulation values of the server. For the color identification information in the 3D visualization scene, the color identification corresponding to each device can be extracted, and the color identification represents the operating state of the device. For example, the server being displayed in red may indicate that its operating state is abnormal, while green indicates normal.
[0083] Step S142: Based on the operation parameter correlation information in the digital twin model, set the normal change range of the operation parameters of each device and the normal display rule of the color identification.
[0084] According to the operation parameter correlation information in the digital twin model, set the normal change range for the operation parameters of each device. Taking the server as an example, for the temperature parameter, a normal temperature change range can be set according to the temperature change law obtained from previous analysis. This range takes into account different operation stages of the device and environmental factors, etc. For the power consumption parameter and the data processing rate parameter, corresponding normal change ranges will also be set respectively. At the same time, set the normal display rule for the color identification in the 3D visualization scene. For example, it is stipulated that when the operation parameters of the device are all within the normal change range, the device is displayed in green; when a certain parameter is close to or exceeds the normal change range, the device is displayed in yellow; when the parameter seriously exceeds the normal change range, the device is displayed in red. The above normal change range and the normal display rule of the color identification are important bases for judging whether the device is abnormal.
[0085] Step S143: Compare the simulated value time series data with the normal change range to identify abnormal parameter values that exceed the normal change range.
[0086] Compare the simulated value time series data of each device extracted with the set normal change range. Taking the server as an example, for the temperature simulated value in the simulated value time series data, it can be checked whether the temperature simulated value at each time point is within the normal temperature change range. If the temperature simulated value at a certain time point exceeds the normal change range, then this temperature simulated value is identified as an abnormal parameter value. Similarly, for the power consumption simulated value and the data processing rate simulated value, similar comparison operations will also be carried out. Through the above comparison, the parameter values that may be abnormal during the operation of the device can be found.
[0087] Step S144: Analyze the matching situation between the color identification in the 3D visualization scene and the normal display rule to identify device nodes with abnormal color identification.
[0088] Conduct a detailed analysis of the color identification of each device in the 3D visualization scene and the normal display rule. Check whether the color identification of each device conforms to the normal display rule. For example, if the operation parameters of the server are all within the normal change range, but the color identification is displayed in red, this indicates that the color identification does not match the normal display rule, and the color identification of this server node is abnormal. Through the above analysis, device nodes with abnormal color identification can be identified, and these nodes may imply potential abnormal situations of the device.
[0089] Step S145: Determine the device nodes with abnormal parameter values or abnormal color identification as the initial abnormal devices.
[0090] After identifying abnormal parameter values that exceed the normal variation range and device nodes with color - identification anomalies, determine the device nodes with abnormal parameter values or color - identification anomalies as initial abnormal devices. For example, if the temperature analog value of server A exceeds the normal variation range, or the color identification of server B is abnormal, then server A and server B are determined as initial abnormal devices. The above - mentioned initial abnormal devices are the starting point for further analyzing abnormal associated devices and the scope of influence.
[0091] Step S146: According to the operation - parameter correlation relationship information in the digital - twin model, analyze the parameter - change trigger conditions and transfer - direction information of the initial abnormal device, and deduce the set of associated devices affected by the parameter change of the initial abnormal device.
[0092] After determining the initial abnormal device, it is necessary to analyze the parameter - change trigger conditions and transfer - direction information of the initial abnormal device according to the operation - parameter correlation relationship information in the digital - twin model, so as to deduce the set of associated devices affected by it.
[0093] Step S1461: Extract the causal - relationship description of the initial abnormal device from the operation - parameter correlation relationship information in the digital - twin model.
[0094] Extract the causal - relationship description of the initial abnormal device from the operation - parameter correlation relationship information in the digital - twin model. Taking the initial abnormal device server A as an example, the causal - relationship description of the operation parameters between server A and other devices can be extracted, including how the parameter change of server A affects the operation parameters of other devices, and the feedback effect of the parameter change of other devices on server A. The above - mentioned causal - relationship description contains the parameter - change trigger conditions and transfer - direction information, which is an important basis for deducing the set of associated devices.
[0095] Step S1462: Analyze the parameter - change trigger conditions in the causal - relationship description, and determine whether the abnormal parameter value of the initial abnormal device meets the conditions for triggering the parameter change of other devices.
[0096] Analyze the parameter - change trigger conditions in the extracted causal - relationship description. For the abnormal parameter value of server A, check whether it meets the conditions for triggering the parameter change of other devices. For example, if the data - processing rate of server A increases abnormally, according to the causal - relationship description, when the data - processing rate increases to a set threshold, it can trigger an increase in the network traffic of the switch connected to it. At this time, it is necessary to judge whether the abnormal increase value of the data - processing rate of server A reaches this threshold. If it reaches, it meets the trigger condition and may affect the parameter change of other devices; if it does not reach, it will not affect other devices temporarily.
[0097] Step S1463: If the trigger condition is met, determine the first-level associated device for parameter change transmission according to the parameter change transmission direction information in the causality description. The first-level associated device is the device that has a direct parameter association relationship with the initial abnormal device.
[0098] If the abnormal parameter value of the initial abnormal device meets the condition for triggering parameter changes in other devices, then determine the first-level associated device for parameter change transmission according to the parameter change transmission direction information in the causality description. Taking Server A as an example, if the abnormal increase in the data processing rate of Server A meets the trigger condition, and according to the causality description, its parameter change will directly affect the connected switch, then this switch is the first-level associated device for parameter change transmission. The first-level associated device has a direct parameter association relationship with the initial abnormal device, and their operating parameters will be affected by the parameter change of the initial abnormal device first.
[0099] Step S1464: Conduct the same analysis on the causality description of the first-level associated device to determine the second-level associated device for parameter change transmission. The second-level associated device is the device that has a parameter association relationship with the first-level associated device.
[0100] After determining the first-level associated device, conduct the same analysis on the causality description of the first-level associated device. Taking the switch as an example, the causality description of the switch can be checked to determine whether its parameter change will trigger parameter changes in other devices. If the network traffic of the switch increases due to the abnormal increase in the data processing rate of Server A, and according to the causality description of the switch, when the network traffic increases to a certain extent, it can affect the network connection status of other servers connected to it, then these affected other servers are the second-level associated devices for parameter change transmission. The second-level associated device has a parameter association relationship with the first-level associated device, and their operating parameters will be indirectly affected by the parameter change of the initial abnormal device.
[0101] Step S1465: Repeat the above analysis process until no new associated devices are identified or the preset transmission level limit is reached.
[0102] The above analysis process of the causal relationship description of associated devices will be continuously repeated to determine more levels of associated devices. The determination of each level of associated devices is based on the parameter changes and causal relationship descriptions of the previous level of associated devices. For example, after determining the second-level associated devices, the causal relationship descriptions of the second-level associated devices can be analyzed to determine the third-level associated devices, and so on. This process will continue until no new associated devices are identified, that is, all affected devices have been determined; or the preset transfer level limit is reached. To avoid the analysis process from being too complex and time-consuming, a transfer level upper limit can be set. When the transfer level upper limit is reached, the derivation process of associated devices is stopped.
[0103] Step S1466: Merge the initial abnormal device, the first-level associated devices, and subsequent levels of associated devices to generate a set of associated devices affected by the parameter changes of the initial abnormal device.
[0104] After completing the derivation of all associated devices, the initial abnormal device, the first-level associated devices, and subsequent levels of associated devices are merged together to form a set of associated devices affected by the parameter changes of the initial abnormal device. This set of associated devices includes all devices directly or indirectly affected by the parameter changes of the initial abnormal device.
[0105] Step S147: Count the number of devices in the set of associated devices and their spatial distribution locations, and generate impact range description information including a list of associated device identifiers and a spatial distribution area.
[0106] Perform statistical analysis on the set of associated devices, count the number of devices in the set, and determine the spatial distribution location of each device. Taking a computer room as an example, the coordinate positions of each associated device in the three-dimensional space coordinate system of the computer room can be recorded. Then, based on the number of devices and their spatial distribution locations, impact range description information is generated. The impact range description information includes a list of associated device identifiers, which lists the identifiers of all associated devices, facilitating managers to quickly identify the affected devices. At the same time, based on the spatial distribution locations of the devices, a spatial distribution area can be determined, which represents the distribution range of the affected devices in the computer room. For example, the associated devices may be concentrated in a certain area of the computer room or scattered in multiple areas. By determining the spatial distribution area, the impact range of the abnormal situation can be intuitively understood.
[0107] Step S148: Merge the initial abnormal device with the set of associated devices and determine them as abnormal associated devices, and use the impact range description information as the impact range of the abnormal associated devices.
[0108] Merge the initial abnormal device with the associated device set to form a larger device set, which is the abnormal associated device. At the same time, use the previously generated impact scope description information as the impact scope of the abnormal associated device, thereby clarifying which devices in the computer room are affected by the abnormal situation and the impact scope of the abnormal situation.
[0109] Step S150: Generate an operation and maintenance adjustment instruction based on the abnormal associated device and the impact scope, and drive the corresponding device in the physical computer room to perform an operation status adjustment operation through the operation and maintenance adjustment instruction.
[0110] After determining the abnormal associated device and the impact scope of the abnormal associated device, it is necessary to generate an operation and maintenance adjustment instruction based on the above information, and drive the corresponding device in the physical computer room to perform an operation status adjustment operation to restore the normal operation status of the device.
[0111] For example, step S151: Extract the device identifier of the abnormal associated device and the corresponding abnormal parameter value.
[0112] Extract the device identifier of each device and the corresponding abnormal parameter value from the abnormal associated device set. Taking a server as an example, the device identifier of the server and the corresponding abnormal temperature value, abnormal power consumption value, or abnormal data processing rate value, etc. can be extracted. The device identifier is used to accurately identify the abnormal associated device, and the abnormal parameter value is the key basis for formulating subsequent adjustment strategies. The above abnormal parameter values are the parameter values that exceed the normal change range determined during the abnormal feature analysis process, ensuring the accuracy and pertinence of the data.
[0113] Step S152: Obtain the adjustment strategy library corresponding to the abnormal parameter value from the operation parameter association relationship information in the digital twin model. The adjustment strategy library includes the parameter correction target value and the adjustment operation steps.
[0114] According to the extracted abnormal parameter value, obtain the corresponding adjustment strategy library from the operation parameter association relationship information in the digital twin model. Taking the abnormal temperature value of a server as an example, the digital twin model stores adjustment strategies for different temperature abnormal situations. The adjustment strategy library includes the parameter correction target value, that is, the normal range value to which the abnormal parameter value is to be adjusted. For example, adjust the abnormally elevated server temperature to the normal temperature range. At the same time, it also includes adjustment operation steps, which detail how to operate the device to achieve parameter correction. For the situation of abnormally elevated server temperature, the adjustment operation steps may include increasing the air conditioning cooling power in the area where the server is located, checking whether the cooling fan of the server is operating normally, etc.
[0115] Step S153: Determine the priority order of the adjustment operations according to the number of associated devices and the spatial distribution area in the impact range description information. The priority order is formulated based on the rules that the areas with a large number of associated devices are adjusted first and the areas with a concentrated spatial distribution are adjusted first.
[0116] Determine the priority order of the adjustment operations based on the number of associated devices and the spatial distribution area in the impact range description information. If there are a large number of associated devices in a certain area, it indicates that this area is more affected by the abnormal situation and may have more serious consequences for the operation of the entire computer room. Therefore, it is necessary to adjust the devices in this area first. For example, if a corner of the computer room has concentrated a large number of associated servers and the operating parameters of these servers are abnormal, then this area should be the first object to be adjusted. For the areas with a concentrated spatial distribution, adjusting them first can solve the problem more efficiently and reduce the waste of resources and time consumption during the adjustment process. According to such rules, sort the adjustment operations in different areas to determine a reasonable priority order.
[0117] Step S154: Integrate the device identifier, abnormal parameter value, parameter correction target value, adjustment operation steps, and priority order of the abnormal associated devices to generate an operation and maintenance adjustment instruction containing device control instructions and operation guidance information.
[0118] Comprehensively integrate the device identifier, abnormal parameter value, parameter correction target value, adjustment operation steps, and priority order of the abnormal associated devices. Taking a server as an example, combine the device identifier of the server, abnormal parameter values such as temperature and power consumption, corresponding parameter correction target values, such as the normal temperature range and power consumption standard, adjustment operation steps, such as turning on the standby cooling device and adjusting the power supply, and the adjustment priority order of the area where the server is located. The integrated information will generate an operation and maintenance adjustment instruction containing device control instructions and operation guidance information. The device control instructions are used to directly control the operating state of the device. For example, send an instruction to the server to reduce the power. The operation guidance information provides specific operation steps and precautions for the operation and maintenance personnel to guide them in performing device adjustment operations.
[0119] Step S155: Send the operation and maintenance adjustment instruction to the controller of the abnormal associated device through the device control system in the computer room.
[0120] There is a dedicated equipment control system in the computer room, and this system is responsible for accurately sending the generated operation and maintenance adjustment instructions to the controllers of the abnormally associated devices. Taking the server as an example, the equipment control system will transmit the operation and maintenance adjustment instructions to the controller of the server through network communication and other means according to the device identifier of the server. During the transmission process, the accuracy and integrity of the operation and maintenance adjustment instructions can be ensured, preventing data loss or errors. A stable communication connection is established between the equipment control system and the controllers of the abnormally associated devices, ensuring that the instructions can be conveyed in a timely and effective manner.
[0121] Step S156: After receiving the operation and maintenance adjustment instruction, the controller executes the parameter correction operation according to the adjustment operation steps and the priority order, and adjusts the device operation parameters to the parameter correction target value.
[0122] After receiving the operation and maintenance adjustment instruction, the controller of the abnormally associated device strictly executes the parameter correction operation according to the adjustment operation steps and the priority order. Taking the server as an example, if the operation and maintenance adjustment instruction requires reducing the power consumption of the server, the controller will gradually adjust the operation mode of the server according to the preset program, such as closing unnecessary service processes, reducing the operating frequency of the CPU, etc., to adjust the power consumption of the server to the parameter correction target value. During the operation process, the priority order can be followed to process the adjustment tasks with higher priorities first, ensuring that the abnormal problems can be solved efficiently. At the same time, the controller will monitor the changes in the device operation parameters in real time and adjust the operation steps in a timely manner according to the parameter feedback, ensuring the accuracy and stability of the adjustment process.
[0123] Step S157: During the execution of the adjustment operation, the adjusted device operation parameters are obtained in real time through the state mapping relationship, and the dynamic operation representation is synchronously updated in the digital twin model until there are no abnormal features in the dynamic operation representation.
[0124] During the execution of the adjustment operation, the adjusted device operation parameters are obtained in real time by using the previously established state mapping relationship. Taking the server as an example, the temperature, power consumption, data processing rate, etc. of the server after adjustment are collected in real time through the sensors installed on the server. Then, these adjusted parameters are transmitted to the digital twin model through the state mapping relationship to synchronously update the dynamic operation representation in the digital twin model. The digital twin model will regenerate the simulated operation state of the device according to the new parameter values and perform visual rendering. During this process, it can be continuously checked whether there are still abnormal features in the dynamic operation representation. If there are still abnormal features, it means that the adjustment operation has not achieved the expected effect, and the reason needs to be further analyzed. It may be necessary to readjust the adjustment strategy or operation steps and continue the adjustment operation until there are no abnormal features in the dynamic operation representation, that is, the operation parameters of the device return to the normal range and the operation state of the entire computer room returns to normal.
[0125] Figure 2 FIG. 2 shows a schematic diagram of exemplary hardware and software components of a computer room management device 100 provided by some embodiments of the present application that can implement the idea of the present application. For example, the processor 120 can be used on the computer room management device 100 and is used to execute the functions in the present application.
[0126] The computer room management device 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the digital twin-based computer room management method of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0127] For example, the computer room management device 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROMs, or RAMs, or any combination thereof. Exemplarily, the computer room management device 100 can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The computer room management device 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0128] For ease of explanation, only one processor is described in the computer room management device 100. However, it should be noted that the computer room management device 100 in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the computer room management device 100 executes steps A and B, it should be understood that steps A and B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0129] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the digital twin-based computer room management method as described above is implemented.
[0130] It should be noted that, in order to simplify the presentation of the disclosure of the present invention and thus help the understanding of one or more embodiments of the present invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.
Claims
1. A computer room management method based on digital twin, characterized in that, The method includes: Constructing a digital twin model of the physical entities in the computer room, where the digital twin model contains the spatial layout structure information of each device in the computer room and the correlation relationship information of the operating parameters between the devices; Obtaining the real-time perception data of each device in the computer room, and establishing a state mapping relationship based on the real-time perception data, the spatial layout structure information, and the operating parameter correlation relationship information in the digital twin model; Simulating the operating state of the device in the digital twin model according to the state mapping relationship, and generating a dynamic operation representation reflecting the real-time operation of the physical computer room; Performing abnormal feature analysis and processing on the dynamic operation representation to determine the abnormal associated devices and the influence scope of the abnormal associated devices; Generating an operation and maintenance adjustment instruction according to the abnormal associated devices and the influence scope, and driving the corresponding devices in the physical computer room to execute the operation state adjustment operation through the operation and maintenance adjustment instruction.
2. The method for managing a computer room based on digital twin according to claim 1, wherein, The constructing of the digital twin model of the physical entities in the computer room, where the digital twin model contains the spatial layout structure information of each device in the computer room and the correlation relationship information of the operating parameters between the devices, includes: Collecting spatial data of the physical entities in the computer room, and obtaining the installation position coordinates of each device, the device shape size parameters, and the routing path information of the connection lines between the devices as the spatial layout structure information; Performing correlation analysis on the historical operation data of each device in the computer room, and extracting the variation rules of the device operation parameters, where the operation parameters include the device temperature parameter, the power consumption parameter, and the data processing rate parameter; Establishing a causal relationship description between the operation parameters of different devices, where the causal relationship description contains the triggering conditions of the parameter change and the information of the parameter change transmission direction; Fusing and modeling the spatial layout structure information and the correlation relationship information of the device operation parameters to generate a digital twin model basic framework containing a three-dimensional space coordinate system and parameter correlation rules; Performing compatibility verification on the spatial position information and the operation parameter information of the newly connected devices through the digital twin model basic framework, and updating the spatial layout structure information and the operation parameter correlation relationship information of the digital twin model.
3. The digital twin-based computer room management method according to claim 2, wherein, The performing of the correlation analysis on the historical operation data of each device in the computer room, and extracting the variation rules of the device operation parameters, where the operation parameters include the device temperature parameter, the power consumption parameter, and the data processing rate parameter, includes: Collecting the historical records of the temperature parameter, the power consumption parameter, and the data processing rate parameter of each device in the computer room during continuous operation cycles; Performing time series analysis on the historical record of the temperature parameter to identify the variation trend of the temperature parameter with the device operation duration, where the variation trend contains the time nodes of the temperature rising stage, the stable stage, and the falling stage; Performing correlation analysis on the historical record of the power consumption parameter, calculating the correlation coefficient between the power consumption parameter and the device data processing rate parameter, and determining the linear relationship or non-linear relationship of the power consumption parameter varying with the data processing rate; Performing periodic analysis on the historical record of the data processing rate parameter, and extracting the peak frequency and the valley frequency of the data processing rate in different service time periods; Integrate the change trend of the temperature parameter, the correlation between the power consumption parameter and the data processing rate, and the periodic characteristics of the data processing rate to generate a description of the change law of the device operation parameters. The description of the change law includes the time-dependent characteristics of parameter changes and the mutual influence characteristics between parameters.
4. The method for computer room management based on digital twin according to claim 1, wherein, Obtain the real-time perception data of each device in the computer room, and establish a state mapping relationship based on the real-time perception data and the spatial layout structure information and operation parameter correlation relationship information in the digital twin model, including: Collect the real-time temperature value, real-time power consumption value, and real-time data processing rate value of the device through the sensor components deployed on the device surface and the connection lines as the real-time perception data; Perform format standardization processing on the real-time perception data to generate standardized perception data consistent with the operation parameter data format in the digital twin model; Extract the spatial layout structure information in the digital twin model, and determine the device identifier and device spatial position coordinates corresponding to each sensor component; Bind the standardized perception data to the corresponding device identifier and device spatial position coordinates to generate multi-dimensional state data including the device identifier, spatial position coordinates, and standardized perception data; Based on the operation parameter correlation relationship information in the digital twin model, establish a parameter mapping rule between the multi-dimensional state data of different devices. The parameter mapping rule includes a parameter synchronization frequency and a parameter change threshold; Map the multi-dimensional state data to the corresponding device nodes of the digital twin model according to the parameter mapping rule to generate a state mapping relationship reflecting the one-to-one correspondence between the physical device and the digital twin model node.
5. The digital-twin-based computer room management method according to claim 4, wherein Based on the operation parameter correlation relationship information in the digital twin model, establish a parameter mapping rule between the multi-dimensional state data of different devices. The parameter mapping rule includes a parameter synchronization frequency and a parameter change threshold, including: Extract the causal relationship description of the operation parameters between devices from the operation parameter correlation relationship information in the digital twin model; According to the parameter change trigger condition in the causal relationship description, determine the synchronization order of the master device parameter and the slave device parameter. The master device parameter is the starting parameter that triggers the parameter change, and the slave device parameter is the following parameter affected by the master device parameter; Based on the synchronization order of the master device parameter and the slave device parameter, set the parameter synchronization frequency. The parameter synchronization frequency is the time interval required for the slave device parameter to complete the update after the master device parameter is updated; Analyze the parameter change transfer direction information in the causal relationship description, determine the maximum acceptable fluctuation range of the parameter change, and use the maximum acceptable fluctuation range as the parameter change threshold; Combine the parameter synchronization order, parameter synchronization frequency, and parameter change threshold to generate a parameter mapping rule between the multi-dimensional state data of different devices.
6. The method for managing a computer room based on digital twin according to claim 5, wherein Simulate the device operation state in the digital twin model according to the state mapping relationship to generate a dynamic operation representation reflecting the real-time operation of the physical computer room, including: Extract the multi-dimensional state data corresponding to each digital twin model node from the state mapping relationship; Synchronize the frequency of the parameters in the state mapping relationship, and input the multi-dimensional state data into the operation simulation module of the corresponding digital twin model node; Based on the parameter correlation relationship information in the digital twin model, the operation simulation module deduces the parameter change trend of the input multi-dimensional state data, and generates a simulation value of the device operation state; Perform consistency verification on the simulation value and the real-time perception data in the multi-dimensional state data. If the verification passes, retain the simulation value. If the verification fails, re-deduce the parameter change trend; Visually render the simulation value that passes the consistency verification in the three-dimensional space coordinate system of the digital twin model according to the device spatial position coordinates, and generate a three-dimensional visualization scene including the color identification of the device operation state, the dynamic curve of parameter change, and the connection line state indication; Synchronize the three-dimensional visualization scene with the real-time updated simulation value to generate a dynamic operation representation that continuously changes over time.
7. The method for managing a computer room based on digital twin according to claim 6, wherein Based on the parameter correlation relationship information in the digital twin model, the operation simulation module deduces the parameter change trend of the input multi-dimensional state data, and generates a simulation value of the device operation state, including: Input the real-time temperature value, real-time power consumption value, and real-time data processing rate value in the multi-dimensional state data into the parameter input layer of the operation simulation module; The parameter input layer of the operation simulation module matches the input real-time parameter value with the description of the device operation parameter change law stored in the digital twin model to determine the current stage of the device operation. The stage includes the temperature rising stage, the stable stage, or the falling stage, the peak stage or the valley stage of the data processing rate; Based on the device operation stage, call the corresponding change trend model in the description of the operation parameter change law. The change trend model includes a temperature change trend model, a correlation model between power consumption and data processing rate, and a periodic model of data processing rate; Deduce the temperature prediction value at a future time point through the temperature change trend model, deduce the power consumption prediction value at a future time point through the correlation model between power consumption and data processing rate, and deduce the data processing rate prediction value at a future time point through the periodic model of data processing rate; Merge the temperature prediction value, power consumption prediction value, and data processing rate prediction value to generate a simulation value of the device operation state.
8. The method for managing a computer room based on digital twin according to claim 1, wherein Analyze and process the abnormal characteristics of the dynamic operation representation to determine the abnormal associated devices and the influence range of the abnormal associated devices, including: Extract the simulation value time series data of each device and the color identification information in the three-dimensional visualization scene from the dynamic operation representation; Based on the parameter correlation relationship information in the digital twin model, set the normal change range of each device operation parameter and the normal display rule of the color identification; Compare the simulation value time series data with the normal change range to identify abnormal parameter values that exceed the normal change range; Analyze the matching situation between the color identification in the three-dimensional visualization scene and the normal display rule to identify the device nodes with abnormal color identification; Determine the device nodes with abnormal parameter values or abnormal color identification as the initial abnormal devices; According to the operation parameter association relationship information in the digital twin model, analyze the parameter change trigger conditions and transmission direction information of the initial abnormal devices, and deduce the set of associated devices affected by the parameter changes of the initial abnormal devices; Count the number and spatial distribution locations of the devices in the set of associated devices, and generate influence range description information including a list of associated device identifiers and a spatial distribution area; Merge the initial abnormal devices with the set of associated devices to determine the abnormally associated devices, and use the influence range description information as the influence range of the abnormally associated devices.
9. The digital twin-based computer room management method according to claim 8, characterized in that The step of, according to the operation parameter association relationship information in the digital twin model, analyzing the parameter change trigger conditions and transmission direction information of the initial abnormal devices, and deducing the set of associated devices affected by the parameter changes of the initial abnormal devices, includes: Extract the causal relationship description of the initial abnormal devices from the operation parameter association relationship information in the digital twin model; Analyze the parameter change trigger conditions in the causal relationship description to determine whether the abnormal parameter values of the initial abnormal devices meet the conditions for triggering parameter changes in other devices; If the trigger conditions are met, according to the parameter change transmission direction information in the causal relationship description, determine the first-level associated devices for parameter change transmission, where the first-level associated devices are the devices directly having a parameter association relationship with the initial abnormal devices; Perform the same analysis on the causal relationship description of the first-level associated devices to determine the second-level associated devices for parameter change transmission, where the second-level associated devices are the devices having a parameter association relationship with the first-level associated devices; Repeat the above analysis process until no new associated devices are identified or the preset transmission level limit is reached; Merge the initial abnormal devices, the first-level associated devices, and subsequent levels of associated devices to generate a set of associated devices affected by the parameter changes of the initial abnormal devices.
10. A computer room management device, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the memory to implement the digital twin-based computer room management method according to any one of claims 1-9 above.
Citation Information
Patent Citations
Digital twin system parameter control method and system
CN117572771A
Research method of intelligent electromechanical machining system based on digital twinning
CN118655773A
Manufacturing industry digital twinning three-dimensional visual management and control method and system
CN119417238A
Machine room environment management and control method, system and equipment based on digital twinning technology
CN119758764A
Building energy equipment energy consumption digital construction method and system based on digital twinning
CN119861571A
Cited By
Remote management and control method and system for data machine room
CN120892758A
Warehousing dynamic scheduling method and system based on digital twinning and deep learning, and storage medium
CN121212744A
Machine room equipment state monitoring method and system combined with three-dimensional visualization
CN121636300A