Fault simulation method, fault simulation device, equipment, storage medium and product
By pairing and injecting faults into the service nodes in the server cluster, synchronizing data to the relevant nodes, and judging their encryption capabilities, the problems of data loss and insufficient reliability judgment in the existing technology are solved, and the accuracy of fault simulation and the data processing capabilities of the cluster are improved.
Patent Information
- Application Number
- CN202510933069.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-07
AI Technical Summary
The existing technology easily causes service node data loss during fault simulation, and lacks an effective way to judge the reliability of the fault simulation.
By pairing multiple service nodes, selecting the target node for fault injection, and synchronizing the data to the relevant nodes, it is determined whether the relevant nodes can be encrypted normally and the fault simulation results are generated.
It effectively avoids data loss on faulty nodes, improves the accuracy and reliability of fault simulation, and ensures the data processing capability of the server cluster in the event of a fault.
Smart Images

Figure CN120455296B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and specifically to a fault simulation method, a fault simulation device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Fault injection simulation is a method of actively introducing errors or failures into a server cluster to test the cluster's response capabilities. It can demonstrate the resilience, fragility, and reliability of the server cluster by observing the ease with which the target system can be damaged and its self-recovery capabilities after being disturbed by other factors such as stress.
[0003] However, in the prior art, when performing fault simulation, actively injecting faults may cause data loss in the service node, and there is a lack of an effective way to judge the reliability in the fault simulation. Summary of the Invention
[0004] In view of the above problems, the present application provides a fault simulation method, a fault simulation device, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to the first aspect of the present application, a fault simulation method is provided, including: pairing a plurality of service nodes to obtain at least one node combination, wherein the node combination includes at least two service nodes, and the service nodes are used to encrypt data; selecting at least one target node from the at least one node combination based on the node status of the plurality of service nodes, and performing a fault injection operation on the target node so that the target node is in a faulty state; for any target node in a faulty state, synchronizing the data of the target node to a related node associated with the target node, wherein the node status of the related node is a normal state; judging whether the related node can encrypt the data normally, and generating a fault simulation result, wherein the fault simulation result represents whether the reliability of the fault injection is passed.
[0006] The second aspect of the present application provides a fault simulation device, including: a pairing module, used to pair multiple service nodes to obtain at least one node combination, wherein the node combination includes at least two service nodes, and the service node is used to encrypt data; an injection module, used to select at least one target node from at least one node combination based on the node status of multiple service nodes, and perform a fault injection operation on the target node to put the target node in a faulty state; a synchronization module, used to synchronize the data of the target node to the relevant nodes associated with the target node for any target node in a faulty state, wherein the node status of the relevant node is a normal state; a judgment module, used to judge whether the relevant node can encrypt the data normally, and generate a fault simulation result, wherein the fault simulation result represents whether the reliability of the fault injection is passed.
[0007] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0008] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instruction stored thereon, which implements the steps of the above-mentioned fault simulation method when the above-mentioned computer program or instruction is executed by a processor.
[0009] The fifth aspect of the present application further provides a computer program product, including a computer program or instructions, which implements the above-mentioned fault simulation method when executed by a processor.
[0010] According to an embodiment of the present application, a node combination is obtained by pairing multiple service nodes, and at least one target node is selected from at least one node combination based on the node status of the service node, and a fault injection operation is performed on the target node to put the target node in a faulty state. At the same time, the data of the target node is synchronized to the relevant nodes associated with the target node, and it is determined whether the relevant nodes can encrypt the data normally, thereby generating a fault simulation result regarding the reliability of the fault injection. Since the target node is selected based on the node status and the data in the target node performing the fault injection operation is synchronized to the relevant nodes, the loss of data of the service node that has failed is avoided. At the same time, the reliability of the fault injection is determined by combining whether the relevant nodes can encrypt the data, which effectively improves the accuracy of the fault simulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 An application scenario diagram of a fault simulation method according to an embodiment of the present application is shown;
[0012] Figure 2A flowchart of a fault simulation method according to an embodiment of the present application is shown;
[0013] Figure 3 A flowchart of a fault simulation method according to another embodiment of the present application is shown;
[0014] Figure 4 A structural block diagram of a fault simulation device according to an embodiment of the present application is shown;
[0015] Figure 5 A logical diagram showing the relationship between the fault simulation device and the server cluster according to an embodiment of the present application is shown;
[0016] Figure 6 A block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0017] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0018] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0020] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0021] The current server cluster can support the following types of failures: (1) There are four controllers in a server. For example, one controller, two controllers, or three controllers can fail. When there are multiple servers, any number of servers can fail. When facing multiple node failures, related technologies usually lose cache data, causing resource configurations such as storage pools and encrypted volumes to go offline, thereby reducing the encryption performance of the server cluster.
[0022] In view of this, the embodiments of the present application provide a fault simulation method, a fault simulation device, an electronic device, a computer-readable storage medium, and a computer program product, which can be applied to the field of data processing technology. The fault simulation method includes: pairing a plurality of service nodes to obtain at least one node combination, wherein the node combination includes at least two service nodes, and the service node is used to encrypt data; based on the node status of the plurality of service nodes, at least one target node is selected from the at least one node combination, and a fault injection operation is performed on the target node so that the target node is in a faulty state; for any target node in a faulty state, the data of the target node is synchronized to a related node associated with the target node, wherein the node status of the related node is a normal state; judging whether the related node can encrypt the data normally, and generating a fault simulation result, wherein the fault simulation result characterizes whether the reliability of the fault injection is passed.
[0023] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0024] Figure 1 An application scenario diagram of the fault simulation method according to an embodiment of the present application is shown.
[0025] like Figure 1 As shown, the server cluster includes a host 101, a controller 102, and a storage device 104. The host 101 can generate plaintext data, and the controller 102 can encrypt the plaintext data to obtain ciphertext data. The storage device 104 can store ciphertext data. The storage device 104 includes but is not limited to a hard disk. Multiple key servers 103 can be connected to the server cluster. The fault simulation method of the embodiment of the present application can be executed by the host 101 that controls the cluster system. Figure 1Although a single controller 102 and storage device 104 are used as an example for illustration, the embodiments of the present application are not limited thereto. A server cluster may include multiple controllers 102 and multiple storage devices 104. The generated encrypted data may be stored in a storage pool of the server cluster. The storage pool refers to the integration of multiple storage devices 104 into a single logical large-capacity storage unit.
[0026] According to embodiments of the present application, a server cluster may refer to a Multiple Controller System (MCS). A Multiple Controller System (MCS) is a distributed system architecture primarily characterized by centralized management and control of multiple service nodes through multiple controllers 102. A Multiple Controller System (MCS) is typically used to improve system reliability and flexibility and is suitable for scenarios requiring high concurrency processing and fault recovery capabilities.
[0027] Figure 2 A flowchart of a fault simulation method according to an embodiment of the present application is shown.
[0028] like Figure 2 As shown, the fault simulation method of this embodiment includes operations S210 to S240.
[0029] In operation S210, a plurality of service nodes are paired to obtain at least one node combination, wherein the node combination includes at least two service nodes, and the service nodes are used to encrypt data.
[0030] In operation S220 , at least one target node is selected from at least one node combination based on the node status of the plurality of service nodes, and a fault injection operation is performed on the target node, so that the target node is in a fault state.
[0031] In operation S230 , for any target node in a faulty state, data of the target node is synchronized to a related node associated with the target node, wherein the node state of the related node is a normal state.
[0032] In operation S240 , it is determined whether the relevant nodes can encrypt data normally, and a fault simulation result is generated, wherein the fault simulation result indicates whether the reliability of the fault injection is passed.
[0033] A service node can refer to a server or controller within a server that performs data encryption operations. A node's status can refer to whether the service node is available. Fault injection involves instructing a service node to be unavailable, such as causing the controller to cease operation or suspend receiving data to be encrypted.
[0034] During fault simulation, multiple service nodes in the server cluster are first paired to obtain at least one node combination. For example, if the server cluster has two servers, each with two controllers, the following node combinations can be obtained through pairing: (node1, node2), (node3, node4), and (server1, server2). Node1 and node2 are the two controllers in server 1, and node3 and node4 are the two controllers in server 2.
[0035] Based on the node status of each service node, at least one target node is selected from the at least one node combination as the node for fault injection. For example, if multiple service nodes are operating normally, node1 or both node1 and node3 can be selected as the target node. The target node can be selected randomly using a random algorithm or sequentially. When using a random algorithm for random selection, the random value of the random algorithm can be set as required, for example, the random value can be set to 0.4.
[0036] After selecting at least one target node, a fault injection operation is performed on each target node, so that the target node subjected to the fault injection operation is in a faulty state. At this time, the target node in the faulty state cannot continue the data encryption operation.
[0037] For each target node in a faulty state, since the target node cannot encrypt data and the service node still has data to be encrypted, the data within the target node, especially the data to be encrypted, can be synchronized with the relevant nodes associated with the target node, allowing the relevant nodes to encrypt the synchronized data. For example, after a fault injection, controller node1 can synchronize the data to be encrypted within controller node1 to controller node2. Or, if both controller node1 and controller node2 fail, that is, server server1 fails, the data to be encrypted within server server1 can be synchronized to server server2.
[0038] After synchronizing the data of the target node to the relevant nodes, it is determined whether the relevant nodes can encrypt the data (including the data of the relevant nodes themselves and the synchronized data). This can obtain the fault simulation results. Through the fault simulation results, it can be seen that the reliability of the system of the server cluster in this fault simulation when facing sudden failures.
[0039] According to an embodiment of the present application, a node combination is obtained by pairing multiple service nodes, and at least one target node is selected from at least one node combination based on the node status of the service node, and a fault injection operation is performed on the target node to put the target node in a faulty state. At the same time, the data of the target node is synchronized to the relevant nodes associated with the target node, and it is determined whether the relevant nodes can encrypt the data normally, thereby generating a fault simulation result regarding the reliability of the fault injection. Since the target node is selected based on the node status and the data in the target node performing the fault injection operation is synchronized to the relevant nodes, the loss of data of the service node that has failed is avoided. At the same time, the reliability of the fault injection is determined by combining whether the relevant nodes can encrypt the data, which effectively improves the accuracy of the fault simulation.
[0040] According to an embodiment of the present application, before performing a fault injection operation, the method further includes: obtaining node information of any service node; and logging into the service node using the node information to determine a target node from multiple logged-in service nodes.
[0041] Node information may include basic information such as user name, password, login port, IP, etc. In the embodiment of the present application, the user's authorization or consent is obtained before obtaining the user's personal information (such as user name and password).
[0042] Before executing the fault simulation method, each service node may be in a logged-in state, and may perform data encryption or fault simulation operations in the logged-in state.
[0043] The node information may be a code file compiled by the staff according to the basic information required for login. The code file is automatically executed to put the service node in a logged-in state, thereby selecting a target node from multiple logged-in service nodes.
[0044] According to an embodiment of the present application, at least one target node is selected from at least one node combination based on the node status of multiple service nodes, including: configuring fault type information, wherein the fault type information characterizes the operating parameters of the service node required to perform the fault injection operation for the fault simulation; generating a node status list based on the service nodes in the at least one node combination; and based on multiple node statuses, selecting at least one target node from the node status list according to the fault type information.
[0045] The fault type information can be specifically configured by the staff based on the fault simulation requirements. For example, the fault type information can be set to a single-node fault type, that is, only one target node needs to be selected for the fault injection operation.
[0046] When selecting the target node, you first need to configure the fault type information, and at the same time, count the node status of the service nodes in different node combinations into a node status list. The node status list can record the identifiers of multiple service nodes and the node status of each service node in sequence.
[0047] In a specific embodiment, if the fault type information is a single-node fault type, one can be randomly or sequentially selected from at least one service node in a normal state as a target node, thereby performing a fault injection operation on the target node.
[0048] According to an embodiment of the present application, a node status list is formed based on the node status of different service nodes, and the target node is selected from the node status list based on the node status, thereby improving the accuracy of selecting the target node and avoiding the occurrence of unsuccessful simulation caused by selecting a node in a fault state for fault simulation.
[0049] According to an embodiment of the present application, based on multiple node states, at least one target node is selected from the node state list according to fault type information, including: for any node combination, when the node combination includes m service nodes in normal state, n target nodes are selected from the m service nodes in normal state, where n is less than m and m≥2; when the node combination includes one service node in normal state, the node states of multiple service nodes are re-determined at preset time intervals until n target nodes are selected when the node combination includes m service nodes in normal state.
[0050] In the process of selecting specific target nodes, it is necessary to select according to the node status of different service nodes. For a certain node combination, if there are m service nodes in normal status in the service nodes in the node combination, then n target nodes can be selected from the m normal service nodes. For example, if there are 3 normal service nodes in the service node combination, 2 normal service nodes can be selected randomly or in sequence as target nodes.
[0051] If there is only one normal service node in the node combination, in order to avoid the node combination being unable to perform data encryption operations due to fault injection into the normal service node, a circular waiting phase can be entered, that is, waiting for a preset time and then re-selecting the target node from the node combination based on the node status of the service node.
[0052] It should be noted that the length of the preset time can be set according to actual needs, for example, it can be 10 seconds.
[0053] In order to further simulate the cluster reliability of the server cluster under extreme conditions, even if the node combination has only one normal service node, the service node can be used as the target node for fault injection, but it is necessary to ensure that at least one service node among the multiple service nodes in the multiple node combinations can receive data synchronized by all target nodes.
[0054] According to an embodiment of the present application, the fault type information includes a single-node fault type or a multi-node fault type.
[0055] According to an embodiment of the present application, a fault injection operation is performed on a target node so that the target node is in a fault state, including: in a case where the fault type information is a single-node fault type, a fault operation instruction is sent to the target node corresponding to the single-node fault type, so that the target node responds to the fault operation instruction and switches the state of the target node to a fault state; in a case where the fault type information is a multi-node fault type, a fault operation instruction is sent to multiple target nodes corresponding to the multi-node fault type respectively, so that any target node responds to the fault operation instruction and switches the state of the target node to a fault state.
[0056] A single-node fault type may refer to injecting a fault into only one service node in a combination of multiple nodes, and a multi-node fault type may refer to injecting a fault into at least two service nodes in a combination of multiple nodes.
[0057] If the fault type information configured by the staff is a single-node fault type, the target node is determined based on the target node selection method recorded above. At this time, relevant fault operation instructions can be sent to the target node, so that the target node responds to the fault operation instructions and adjusts its own state, so that the state of the target node is a fault state, thereby synchronizing the data.
[0058] If the fault type information configured by the staff is a multi-node fault type, multiple target nodes are determined based on the target node selection method recorded above. At this time, relevant fault operation instructions can be sent to each target node separately, so that the target node responds to the fault operation instruction and adjusts its own state, so that the state of the target node is a fault state, thereby synchronizing the data.
[0059] In a specific embodiment, if the node combinations are (node1, node2), (node2, node3), (node3, node4), and (node4, node1), totaling four service nodes, and the current fault type indicates a multi-node fault, controllers node1, node2, and node3 may all be selected as target nodes. Fault operation instructions are then sent to each of these controllers, adjusting their states to a faulty state. In this case, controllers node1, node2, and node3 can synchronize their data with controller node4 for encryption, thereby determining whether the fault injection has been reliable.
[0060] According to the embodiments of the present application, by configuring the fault type information of a single-node fault type or a multi-node fault type, the fault scenarios of the server cluster under different emergency situations can be simulated, thereby improving the accuracy of the fault simulation and effectively ensuring the performance of the server cluster in actual work.
[0061] According to an embodiment of the present application, whether the relevant nodes can encrypt data normally is judged, and a fault simulation result is generated, including: when the relevant nodes can encrypt data normally, a simulation pass result is generated, wherein the simulation pass result characterizes that the reliability of the fault injection is a pass state; when the relevant nodes cannot encrypt data, a simulation failure result is generated, wherein the simulation failure result characterizes that the reliability of the fault injection is a fail state, wherein the fault simulation result includes a simulation pass result or a simulation failure result.
[0062] The above embodiment is specifically described. If the controller node4 can normally encrypt its own data and synchronized data, it means that the reliability of this fault injection is passed, and thus a simulation pass result can be generated.
[0063] If the controller node4 fails to encrypt, it means that the entire server cluster cannot continue to encrypt data when multiple service nodes fail, indicating that the reliability of this fault injection fails, and a simulation failure result can be generated.
[0064] According to the embodiments of the present application, by judging whether the relevant nodes can work normally when the target node is in a fault state during the fault simulation process, it is determined whether the reliability of the fault injection is passed, thereby ensuring that the server cluster can continue to process data even if multiple service nodes are in a fault state in actual production, thereby indirectly improving the working performance of the server cluster.
[0065] According to an embodiment of the present application, the multiple service nodes include multiple controllers of any server, or multiple servers; the node combination includes at least two controllers of at least one server, or at least two servers.
[0066] In a specific embodiment, assume that there are two servers, such as the first server server1 and the second server server2, and each server includes two controllers, namely node1, node2, node3, and node4. At this time, the above multiple controllers and servers can be paired into the following five node combinations, namely cycle0 (node1, node2), cycle 1 (node2, node3), cycle 2 (node3, node4), cycle 3 (node4, node1), and cycle 4 (server1, server2), where cycle represents a node combination.
[0067] It should be noted that the above-mentioned server can not only be a server that performs data encryption, but also a key management server or a server with other functions independent of the data encryption server.
[0068] According to an embodiment of the present application, the fault type information includes a single-node fault type or a multi-node fault type.
[0069] According to an embodiment of the present application, when the fault type information is a single-node fault type, the data of the target node is synchronized to the relevant node associated with the target node, including: determining another service node in the node combination where the target node is located as the relevant node; synchronizing the data of the target node to the relevant node for data encryption.
[0070] If the fault type information indicates a single-node fault type, you can select a target node from the following five node combinations: cycle 0 (node1, node2), cycle 1 (node2, node3), cycle 2 (node3, node4), cycle 3 (node4, node1), and cycle 4 (server1, server2) for fault injection. For example, controller (i.e., service node) node2 is selected as the target node. At this time, the data in service node node2 can be synchronized to the relevant nodes for data encryption. For example, service node node1 or service node node3 can be selected as the relevant node.
[0071] In a specific embodiment, if the node status list shows that the node status of service node node1 is a fault state, service node node3 can be used as a related node. Similarly, if the node status list shows that the node status of service node node3 is a fault state, service node node1 can be used as a related node.
[0072] According to an embodiment of the present application, when the fault type information is a multi-node fault type, the data of the target node is synchronized to the relevant nodes associated with the target node, including: when another service node in the node combination where the target node is located is not in a faulty state, the other service node is determined as the relevant node; when another service node is in a faulty state, the service node that is not in a faulty state in the associated node combination related to the node combination is determined as the relevant node; and the data of the target node is synchronized to the relevant node for data encryption.
[0073] If the fault type information is a single-node fault type, you can select a target node from the five node combinations cycle0 (node1, node2), cycle 1 (node2, node3), cycle 2 (node3, node4), cycle 3 (node4, node1), and cycle 4 (server1, server2) for fault injection. For example, select service nodes node2 and node4 as target nodes. At this time, the data in service nodes node2 and node4 can be synchronized to related nodes for data encryption. For example, the data in service node node2 can be synchronized to service node node1 or service node node3, and the data in service node node4 can be synchronized to service node node1 or service node node3. At this time, service node node1 or service node node3 are both in normal status.
[0074] Assume that service node node3 is in a faulty state in node combination cycle 1 (node2, node3). At this time, the data in service node node2 can be synchronized to service node node1 in the associated node combination cycle 0 (node1, node2) related to service node node2, which is not in a faulty state.
[0075] If the status of service node node1 is faulty, cycle 2 (node3, node4) can be determined as the associated node combination, thereby synchronizing the data in service node node2 to service node node3 that is not in a faulty state, so that the synchronized data can be encrypted through service node node3.
[0076] According to an embodiment of the present application, the node combination includes a first node combination, and the service node in the first node combination includes at least two controllers of the server.
[0077] According to an embodiment of the present application, the fault simulation method also includes: generating a permanent key, an access management key and a data encryption key in response to a key generation request; encapsulating the permanent key and the access management key to obtain a key encryption key; encrypting the data encryption key using the key encryption key to obtain a data access key; storing the data access key in the controller so that the controller uses the data access key to encrypt the data to store the encrypted data in the encryption pool.
[0078] A Persistent Master Key (PMK) is a key with a long validity period and does not expire automatically by default. It is suitable for applications that require long-term stable operation. The Master Access Key (MAK) is a random number generated by the system. The MAK serves as the master key for encrypting the Data Encryption Key (DEK). The Data Encryption Key (DEK) is the key used by encryption and decryption algorithms to encrypt data. It is a string of characters used to alter data to make it appear random, similar to a physical key. Only those with the matching key can decrypt the data. The Data Access Key (DAK) is the key used to access data. The Data Access Key is generated by encrypting the Data Encryption Key. The Key Encryption Key (KEK) is the key used to encrypt other keys.
[0079] The encryption processor can send a key generation request to the key management module (KeyManager, KeyMgr) through an asynchronous message queue (Message Queue, MQ). The key management module (KeyManager, KeyMgr) responds to the key generation request to generate a permanent key, an access management key, and a data encryption key.
[0080] The encryption processor calls the encryption management module to first encapsulate the permanent key and access management key to obtain the key encryption key. The key management module calls the encryption management module (OpenSSL) to encrypt and encapsulate the data encryption key (DEK) based on the key encryption key to obtain the data access key (DAK). After obtaining the data access key, the encryption processor generates feedback information and transmits it to the encryption processor. After confirming the generation of the data access key, the encryption processor stores the data access key in the controller's secure memory. The controller can then use the data access key to encrypt data and store the encrypted data in the encryption pool.
[0081] The encryption management module can provide a unified encryption and decryption interface for the encryption call (PLIF_encryption) module to call, and perform actual data encryption and decryption operations through the integrated encryption algorithm. The encryption algorithm can be any type of encryption algorithm, such as SM (ShangMi) -3 and other cryptographic hash algorithms, SM (ShangMi) -4 and other block cipher algorithms or Advanced Encryption Standard (AES) algorithm.
[0082] According to an embodiment of the present application, by encapsulating the permanent key and the access management key into a key encryption key and using the key encryption key to perform secondary encryption on the data encryption key, the risk of data leakage can be reduced, thereby improving data security.
[0083] Figure 3 A flowchart of a fault simulation method according to another embodiment of the present application is shown.
[0084] According to an embodiment of the present application, the node combination further includes a second node combination, and the service nodes in the second node combination include at least two key servers.
[0085] According to an embodiment of the present application, the fault simulation method also includes: storing a permanent key in a key server in the second node combination; when the target node in the faulty state in the first node combination recovers to a normal state, the target node obtains a data access key from a related node and obtains a permanent key from any key server; generates a new data access key based on the permanent key; performs a consistency check on the data access key and the new data access key to obtain a check result; when the check result shows that the data access key is inconsistent with the new data access key, the target node in the normal state uses the new data access key to encrypt the data to store the encrypted data in the encryption pool.
[0086] See also Figure 3 Before the service node performs data encryption, it is first necessary to configure the key server in operation S301 so that the key server can store the permanent key, and then in operation S302, create an encryption pool and an encryption volume in the storage device, and then in operation S303, map the created encryption pool and encryption volume to the host, so that the host can send the data to be encrypted to the service node so that the service node can encrypt the data based on the data access key in operation S304.
[0087] In the process of storing the permanent key in the key server, the encryption processor sends a key write request to the key management module so that the key management module stores the permanent key in the key server in the second node combination and sends write feedback information to the encryption processor.
[0088] In another embodiment, the permanent key can be stored in a removable storage device (e.g., a portable hard drive, memory card, etc.). After the permanent key is stored in the removable storage device, the removable storage device can be disconnected to ensure the security of the permanent key. The purpose of storing the permanent key in the removable storage device is to provide an emergency key recovery method in the event of a failure of all key servers.
[0089] After this fault simulation, before the next round of fault simulation, the target node that performed the fault injection operation needs to be restored to its state. For example, a fault recovery instruction can be sent to the target node so that the target node responds to the fault recovery instruction and switches the state of the target node to normal.
[0090] In operation S305, it is determined whether the status of the target node has been restored to a normal state. For the target node that has been restored to a normal state, the target node can obtain the data access key from the relevant node in operation S306, and obtain the permanent key from the key server and the access management key and data access key from the server cluster in operation S307, thereby generating a new data access key DAK based on the permanent key PMK, the access management key MAK stored in the server cluster, and the data encryption key DEK.
[0091] If it is determined in operation S305 that the state of the target node has not been restored to a normal state, the state of the target node continues to be adjusted, thereby iteratively determining in operation S305 whether the state of the target node has been restored to a normal state.
[0092] In operation S308, the data access key DAK and the new data access key DAK are checked for consistency to obtain a verification result. For example, a consistency check can be performed through a similarity function (such as cosine similarity). If the data access key DAK is inconsistent with the new data access key DAK, it means that during the fault injection, the server cluster has replaced the data access key, and the relevant node has temporarily not used the latest data access key because the data has not been encrypted yet. At this time, in operation S309, the target node uses the new data access key DAK to encrypt the data.
[0093] If the data access key DAK is consistent with the new data access key DAK, data encryption is performed using the data access key DAK of the relevant node in operation S310, or the new data access key DAK may be used to encrypt data.
[0094] According to an embodiment of the present application, by restoring the target node to a normal state and performing consistency verification on the data access key in the relevant node and the new data access key generated based on the permanent key PMK in the key server, the use of expired data access keys for data encryption can be avoided, thereby improving data security.
[0095] According to an embodiment of the present application, a controller encrypts data using a data access key to store the encrypted data in an encryption pool, including: when the controller confirms that the encryption pool has a pool encryption tag, parsing the data access key to obtain a data encryption key encapsulated in the data access key, wherein the pool encryption tag indicates that the encryption pool is a storage pool in which data needs to be encrypted and stored; encrypting the data using the parsed data encryption key to obtain encrypted data; and storing the encrypted data in the encryption pool.
[0096] During data encryption, the data to be encrypted is first sent to the storage pool. The controller uses the pool management module (VG) to determine whether the storage pool has a pool encryption tag (dak-tag). If the pool encryption tag is confirmed, the controller uses the encryption call module (Plif encryption) to request a key from the encryption processor. The encryption processor retrieves the data access key (DAK) from the controller's secure memory and sends the DAK to the encryption call module. The encryption call module parses the DAK to obtain the data encryption key (DEK) encapsulated in the DAK. The encryption algorithm in the encryption management module uses the DEK to encrypt the data to be encrypted, obtaining the encrypted data. The encrypted target data is then stored in the storage pool (encryption pool). Similarly, the encryption management module can use the data encryption key encapsulated in the DAK to decrypt the encrypted data stored in the encryption pool.
[0097] If the pool management module determines that the storage pool does not have the pool encryption tag dak-tag, it means that the storage pool is a storage pool for storing unencrypted data. At this time, the data can be stored directly in the unencrypted storage pool without being encrypted.
[0098] The pool management module is a key component of the Logical Volume Manager (LVM). It combines multiple physical volumes (PVs) into a large logical storage pool for more flexible storage space management and allocation. A physical volume refers to a physical storage device or its partitions, such as a hard disk or solid-state drive.
[0099] According to an embodiment of the present application, by determining whether a pool encryption tag exists when data is encrypted, the encrypted and non-encrypted target data are stored in an encryption pool or a non-encrypted storage pool, thereby achieving the coexistence of encrypted and non-encrypted services, thereby improving the storage convenience during data storage.
[0100] Based on the above fault simulation method, this application also provides a fault simulation device. Figure 4 The device is described in detail.
[0101] Figure 4 A structural block diagram of a fault simulation device according to an embodiment of the present application is shown.
[0102] like Figure 4 As shown, the fault simulation device 400 of this embodiment includes a pairing module 410 , an injection module 420 , a synchronization module 430 , and a judgment module 440 .
[0103] The pairing module 410 is used to perform pairing processing on multiple service nodes to obtain at least one node combination, wherein the node combination includes at least two service nodes, and the service nodes are used to encrypt data.
[0104] The injection module 420 is configured to select at least one target node from at least one node combination based on the node status of multiple service nodes, and perform a fault injection operation on the target node to put the target node into a fault state.
[0105] The synchronization module 430 is configured to synchronize data of any target node in a faulty state to related nodes associated with the target node, wherein the node status of the related nodes is a normal state.
[0106] The judgment module 440 is used to judge whether the relevant nodes can encrypt data normally and generate a fault simulation result, wherein the fault simulation result represents whether the reliability of the fault injection is passed.
[0107] According to an embodiment of the present application, a node combination is obtained by pairing multiple service nodes, and at least one target node is selected from at least one node combination based on the node status of the service node, and a fault injection operation is performed on the target node to put the target node in a faulty state. At the same time, the data of the target node is synchronized to the relevant nodes associated with the target node, and it is determined whether the relevant nodes can encrypt the data normally, thereby generating a fault simulation result regarding the reliability of the fault injection. Since the target node is selected based on the node status and the data in the target node performing the fault injection operation is synchronized to the relevant nodes, the loss of data of the service node that has failed is avoided. At the same time, the reliability of the fault injection is determined by combining whether the relevant nodes can encrypt the data, which effectively improves the accuracy of the fault simulation.
[0108] According to an embodiment of the present application, the fault simulation device 400 further includes an acquisition module and a login module.
[0109] The acquisition module is used to obtain the node information of any service node.
[0110] The login module is used to log in to the service node using the node information to determine the target node from multiple logged-in service nodes.
[0111] According to an embodiment of the present application, the injection module 420 includes a configuration unit, a first generation unit, and a selection unit.
[0112] The configuration unit is used to configure fault type information, wherein the fault type information represents the operating parameters of the service node required to perform the fault injection operation for the fault simulation.
[0113] The first generating unit is configured to generate a node status list according to the service nodes in at least one node combination.
[0114] The selection unit is used to select at least one target node from the node status list based on multiple node states and according to fault type information.
[0115] According to an embodiment of the present application, the selection unit includes a first selection subunit and a second selection subunit.
[0116] The first selection subunit is configured to select n target nodes from m normal service nodes for any node combination, when the node combination includes m normal service nodes, where n is less than m and m≥2.
[0117] The second selection subunit is used to redetermine the node status of multiple service nodes at preset intervals when the node combination includes a service node in normal status, until n target nodes are selected when the node combination includes m service nodes in normal status.
[0118] According to an embodiment of the present application, the fault type information includes a single-node fault type or a multi-node fault type.
[0119] According to an embodiment of the present application, the injection module 420 further includes a first sending unit and a second sending unit.
[0120] The first sending unit is used to send a fault operation instruction to the target node corresponding to the single-node fault type when the fault type information is a single-node fault type, so that the target node responds to the fault operation instruction and switches the state of the target node to a fault state.
[0121] The second sending unit is used to send fault operation instructions to multiple target nodes corresponding to the multi-node fault type in sequence when the fault type information is a multi-node fault type, so that any target node responds to the fault operation instruction and switches the state of the target node to a fault state.
[0122] According to an embodiment of the present application, the judgment module 440 includes a second generation unit and a third generation unit.
[0123] The second generating unit is used to generate a simulation pass result when the relevant nodes can encrypt data normally, wherein the simulation pass result indicates that the reliability of the fault injection is in a pass state.
[0124] The third generating unit is used to generate a simulation failure result when the relevant node cannot encrypt the data, wherein the simulation pass result represents that the reliability of the fault injection is in a failed state, wherein the fault simulation result includes a simulation pass result or a simulation failure result.
[0125] According to an embodiment of the present application, the multiple service nodes include multiple controllers of any server, or multiple servers; the node combination includes at least two controllers of at least one server, or at least two servers.
[0126] According to an embodiment of the present application, the fault type information includes a single-node fault type or a multi-node fault type.
[0127] According to an embodiment of the present application, when the fault type information is a single-node fault type, the synchronization module 430 includes a first determination unit and a first synchronization unit.
[0128] The first determining unit is configured to determine another service node in the node combination where the target node is located as a related node.
[0129] The first synchronization unit is used to synchronize the data of the target node to the relevant nodes for data encryption.
[0130] According to an embodiment of the present application, when the fault type information is a multi-node fault type, the synchronization module 430 includes a second determination unit, a third determination unit, and a second synchronization unit.
[0131] The second determining unit is configured to determine another service node as a related node if another service node in the node combination where the target node is located is not in a fault state.
[0132] The third determining unit is configured to determine, when another service node is in a fault state, a service node that is not in a fault state in a combination of associated nodes related to the node combination as a related node.
[0133] The second synchronization unit is used to synchronize the data of the target node to the relevant nodes for data encryption.
[0134] According to an embodiment of the present application, the node combination includes a first node combination, and the service node in the first node combination includes at least two controllers of the server.
[0135] According to an embodiment of the present application, the fault simulation device 400 further includes a first generation module, a packaging module, an encryption module, and a storage encryption module.
[0136] The first generation module is used to generate a permanent key, an access management key and a data encryption key in response to a key generation request.
[0137] The encapsulation module is used to encapsulate the permanent key and the access management key to obtain the key encryption key.
[0138] The encryption module is used to encrypt the data encryption key using the key encryption key to obtain the data access key.
[0139] The storage encryption module is used to store the data access key in the controller, so that the controller encrypts the data using the data access key, and stores the encrypted data in the encryption pool.
[0140] According to an embodiment of the present application, the node combination further includes a second node combination, and the service nodes in the second node combination include at least two key servers.
[0141] According to an embodiment of the present application, the fault simulation device 400 further includes a storage module, a second acquisition module, a second generation module, a verification module, and a data encryption module.
[0142] The storage module is configured to store the permanent key in a key server in the second node combination.
[0143] The second acquisition module is used for, when the target node in the first node combination that is in a faulty state recovers to a normal state, enabling the target node to obtain a data access key from a related node and to obtain a permanent key from any key server.
[0144] The second generating module is used to generate a new data access key according to the permanent key.
[0145] The verification module is used to perform consistency verification on the data access key and the new data access key to obtain a verification result.
[0146] The data encryption module is used to encrypt data using the new data access key at the target node in a normal state when the verification result shows that the data access key is inconsistent with the new data access key, so as to store the encrypted data in the encryption pool.
[0147] According to an embodiment of the present application, the storage encryption module includes a key parsing unit, a data encryption unit, and a data storage unit.
[0148] The key parsing unit is used to parse the data access key to obtain the data encryption key encapsulated in the data access key when the controller confirms that the encryption pool has a pool encryption tag, wherein the pool encryption tag indicates that the encryption pool is a storage pool where data needs to be encrypted and stored.
[0149] The data encryption unit is used to encrypt data using the data encryption key obtained by parsing to obtain encrypted data.
[0150] The data storage unit is used to store encrypted data in the encryption pool.
[0151] According to embodiments of the present application, any multiple modules among pairing module 410, injection module 420, synchronization module 430, and judgment module 440 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of pairing module 410, injection module 420, synchronization module 430, and judgment module 440 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of pairing module 410, injection module 420, synchronization module 430, and judgment module 440 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0152] Figure 5 A logical diagram between a fault simulation device and a server cluster according to an embodiment of the present application is shown.
[0153] like Figure 5 As shown, after the user enters the encryption configuration instruction for encrypting the target data (i.e., the data to be encrypted) on the user interface of the user end, the compatibility support module (IC Compatibility Support Module, IC CSM) of the user end responds to the encryption configuration instruction by retrieving the license from the software authorization license (Encryption License) module and activating the encryption authorization license through the encryption module (IC_encryption_csm) of the license manager. After that, the compatibility support module parses the encryption configuration instruction and converts it into an internal processing format, and then sends the command to the encryption processing module for execution.
[0154] The compatibility support module is connected to the encryption processing module and the encryption management module through the pool management module (VG CSM), the management monitoring module (RAID CSM), and the protocol management module (VL CSM).
[0155] The encryption processing module receives encryption configuration commands transmitted from the license manager, sends the encryption configuration commands to the key management module through an asynchronous message queue (MQ), processes the returned results, records the command execution results to the server cluster, and feeds back to the user end.
[0156] The key management module receives and processes encryption configuration commands from the encryption processing module to generate keys (permanent keys, access management keys, data encryption keys, etc.) locally, and returns the results of key generation or management operations to the encryption processing module through the message queue.
[0157] According to an embodiment of the present application, the encryption management module provides a hardware-supported random number generation function to ensure the security of encryption operations.
[0158] According to the embodiments of the present application, encryption and decryption of target data involve an encryption call module (PLIF_encryption), an encryption processing module, and an encryption management module, ensuring the security and high performance of the data encryption service.
[0159] The encryption calling module inserts encryption and decryption operations in the target data IO, requests the data encryption key from the encryption processing module, and uses the obtained key to perform target data encryption and decryption operations by calling the encryption engine of the encryption management module.
[0160] The encryption processing module receives the key acquisition request from the encryption calling module, calls the encryption management module to acquire or generate the required key, and passes the acquired key to the encryption calling module.
[0161] The encryption management module provides a unified encryption and decryption interface for the encryption call module to call, and performs encryption and decryption operations on the target data through integrated encryption algorithms (such as AES, SM4, etc.). For example, the hardware-supported random number generation function provided by the Trusted Platform Module (TPM) ensures the security of encryption operations.
[0162] According to an embodiment of the present application, the key management unit in the key management module (KeyMgr) transmits the keys (such as data access keys and permanent keys, etc.) obtained or generated by the encryption processing module to the key server, mobile storage and secure memory of the controller based on the KMIP (Key Management Interoperability) protocol or the USB protocol.
[0163] According to an embodiment of the present application, the key management module calls the encryption management module to encapsulate the data encryption key based on the permanent key and the access management key, thereby obtaining the data access key. After obtaining the data access key, the encryption processing module generates feedback information and transmits it to the encryption processing module. After confirming the generation of the data access key, the encryption processing module stores the data access key in the secure memory of the controller. When performing data encryption and decryption, the pool management module (VG CSM) can be used to determine whether the storage pool has a pool encryption tag dak-tag. If the pool encryption tag dak-tag is present, the encryption calling module is used to request a key from the encryption processing module, so that the encryption processing module obtains the data access key DAK from the secure memory of the controller for data encryption and decryption.
[0164] Figure 6 A block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present application is shown.
[0165] like Figure 6 As shown, an electronic device 600 according to an embodiment of the present application includes a processor 601, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0166] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in one or more memories.
[0167] According to an embodiment of the present application, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.
[0168] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0169] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.
[0170] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided in the embodiments of the present application.
[0171] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the computer program is executed by the processor 601. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0172] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0173] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0174] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0176] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0177] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A fault simulation method, characterized in that: The fault simulation method comprises: Pairing the plurality of service nodes to obtain at least one node combination, the node combination including a first node combination, the first node combination including at least two controllers of a server as at least two service nodes, the service nodes being used to encrypt data; Selecting at least one target node from at least one of the node combinations based on the node states of the plurality of service nodes, and performing a fault injection operation on the target node so that the target node is in a fault state; For any target node in a faulty state, synchronize the data of the target node to a related node associated with the target node, where the node state of the related node is normal; Determine whether the relevant nodes can encrypt data normally, and generate a fault simulation result, wherein the fault simulation result indicates whether the reliability of the fault injection is passed; generating, in response to a key generation request, a permanent key, an access management key, and a data encryption key; Encapsulate the permanent key and access management key to obtain the key encryption key; The data encryption key is encrypted using the key encryption key to obtain a data access key; The data access key is stored in the controller, so that the controller encrypts the data using the data access key to store the encrypted data in the encryption pool.
2. The fault simulation method according to claim 1, characterized in that: Before performing the fault injection operation, it also includes: For any of the service nodes, obtaining node information of the service node; The service node is logged in using the node information to determine the target node from a plurality of logged-in service nodes.
3. The fault simulation method according to claim 1, characterized in that: Selecting at least one target node from at least one of the node combinations based on the node states of the plurality of service nodes includes: Configuring fault type information, wherein the fault type information represents operating parameters of a service node required to perform a fault injection operation for fault simulation; Generate a node status list according to at least one service node in the node combination; Based on multiple node states, at least one target node is selected from the node state list according to the fault type information.
4. The fault simulation method according to claim 3, characterized in that: Based on the plurality of node states, selecting at least one target node from the node state list according to the fault type information includes: For any of the node combinations, when the node combination includes m service nodes in normal state, n target nodes are selected from the m service nodes in normal state, where n is less than m and m≥2; When the node combination includes a service node in a normal state, the node states of the plurality of service nodes are re-determined at preset intervals until n target nodes are selected when the node combination includes m service nodes in a normal state.
5. The fault simulation method according to claim 3, characterized in that: The fault type information includes a single-node fault type or a multi-node fault type; Performing a fault injection operation on the target node so that the target node is in a fault state includes: When the fault type information is the single-node fault type, sending a fault operation instruction to a target node corresponding to the single-node fault type, so that the target node responds to the fault operation instruction and switches the state of the target node to the fault state; In the case where the fault type information is the multi-node fault type, fault operation instructions are sent in sequence to multiple target nodes corresponding to the multi-node fault type, so that any one of the target nodes responds to the fault operation instruction and switches the state of the target node to the fault state.
6. The fault simulation method according to claim 1, characterized in that: Determining whether the relevant nodes can encrypt data normally and generating a fault simulation result includes: If the relevant nodes are able to encrypt data normally, generating a simulation pass result, wherein the simulation pass result indicates that the reliability of the fault injection is in a pass state; In the case where the relevant node cannot encrypt the data, a simulation failure result is generated, wherein the simulation failure result indicates that the reliability of the fault injection is in a failed state, wherein the fault simulation result includes the simulation pass result or the simulation failure result.
7. The fault simulation method according to claim 1, characterized in that: The plurality of service nodes include a plurality of controllers of any server, or a plurality of the servers; the node combination includes at least two controllers of at least one of the servers, or at least two of the servers.
8. The fault simulation method according to claim 3, characterized in that: The fault type information includes a single-node fault type or a multi-node fault type; Wherein, when the fault type information is the single-node fault type, synchronizing the data of the target node to a related node associated with the target node includes: Determine another service node in the node combination where the target node is located as the relevant node; Synchronize the data of the target node to the relevant nodes for data encryption; Wherein, when the fault type information is the multi-node fault type, synchronizing the data of the target node to related nodes associated with the target node includes: If another service node in the node combination where the target node is located is not in a faulty state, determining the other service node as the relevant node; In the case that the another service node is in a fault state, determining a service node that is not in a fault state in an associated node combination related to the node combination as the related node; The data of the target node is synchronized to the related nodes for data encryption.
9. The fault simulation method according to claim 1, characterized in that: The node combination further includes a second node combination, wherein the service nodes in the second node combination include at least two key servers; Wherein, the fault simulation method further includes: a key server storing the permanent key in the second node combination; When the target node in the first node combination that is in a faulty state recovers to a normal state, the target node obtains the data access key from a related node and obtains the permanent key from any key server; generating a new data access key based on the permanent key; Performing a consistency check on the data access key and the new data access key to obtain a check result; When the verification result indicates that the data access key is inconsistent with the new data access key, the target node in a normal state encrypts the data using the new data access key to store the encrypted data in the encryption pool.
10. The fault simulation method according to claim 1, characterized in that: The controller encrypts the data using the data access key to store the encrypted data in the encryption pool, including: When the controller confirms that the encryption pool has a pool encryption tag, parsing the data access key to obtain a data encryption key encapsulated in the data access key, wherein the pool encryption tag indicates that the encryption pool is a storage pool in which data needs to be encrypted and stored; Encrypting the data using the data encryption key obtained by parsing to obtain encrypted data; The encrypted data is stored in the encryption pool.
11. A fault simulation device, characterized in that: The fault simulation device comprises: a pairing module, configured to perform pairing processing on the plurality of service nodes to obtain at least one node combination, wherein the node combination includes a first node combination, the first node combination includes at least two controllers of a server as at least two service nodes, and the service nodes are configured to encrypt data; an injection module, configured to select at least one target node from at least one of the node combinations based on the node states of the plurality of service nodes, and perform a fault injection operation on the target node, so that the target node is in a fault state; a synchronization module, configured to synchronize data of any target node in a faulty state to related nodes associated with the target node, wherein the node state of the related node is a normal state; a judgment module, configured to judge whether the relevant node is capable of encrypting data normally, and generate a fault simulation result, wherein the fault simulation result indicates whether the reliability of the fault injection is passed; A first generation module, configured to generate a permanent key, an access management key, and a data encryption key in response to a key generation request; The encapsulation module is used to encapsulate the permanent key and the access management key to obtain the key encryption key; An encryption module, configured to encrypt a data encryption key using a key encryption key to obtain a data access key; The storage encryption module is used to store the data access key in the controller, so that the controller encrypts the data using the data access key, and stores the encrypted data in the encryption pool.
12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the fault simulation method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the fault simulation method according to any one of claims 1 to 10 are implemented.
14. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program is used to implement the fault simulation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Fault injection method and device based on mirror image pair, equipment and storage medium
CN115328814A