Method, device and storage medium for secure migration control of a single floating drive in a dual-controller, dual-instance tape storage system
By acquiring the global state to divide the source instance and target instance, controlling the source instance to release the floating drive and generate a period identifier takeover record, and the target instance to take over the floating drive, the problem of inconsistent state caused by concurrent access in the dual-controller dual-instance tape storage system is solved, and the stability and security of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN JIETENG TECHNOLOGY CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-17
AI Technical Summary
In a dual-controller, dual-instance tape storage system, the source instance and the target instance may concurrently access the same drive, leading to inconsistent states and thus the risk of split-brain scheduling.
By obtaining the global state from the topology storage unit, the source instance and the target instance are divided. The source instance releases the floating drive and generates a takeover record with an era identifier. The target instance takes over the floating drive based on the takeover record with the era identifier, strictly limiting the access sequence and avoiding concurrent access.
It ensures state consistency and scheduling security during cross-instance migration of floating drives, and improves the operational stability and data security of dual-controller dual-instance tape storage systems.
Smart Images

Figure CN122086336B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage system control technology, and in particular to a method, device and storage medium for secure migration control of a single floating drive in a dual-controller dual-instance tape storage system. Background Technology
[0002] In current tape storage systems, a common deployment approach is to use a multi-instance shared drive resource pool, i.e., general resource pooling. However, in a dual-controller, dual-instance tape storage system, where only a single floating drive is migrated across instances, using a multi-instance shared drive resource pool can lead to the source and target instances concurrently accessing the same drive, resulting in inconsistent states and potentially causing split-brain scheduling risks. Summary of the Invention
[0003] The main objective of this application is to provide a method, device, and storage medium for secure migration control of a single floating drive in a dual-controller, dual-instance tape storage system. This aims to solve the technical problem that the source instance and the target instance may concurrently access the same drive, leading to inconsistent states and thus the risk of split-brain scheduling.
[0004] To achieve the above objectives, this application provides a secure migration control method for a single floating drive in a dual-controller, dual-instance tape storage system. The dual-controller, dual-instance tape storage system includes a first instance, a second instance, multiple fixed drives, and a single floating drive. The secure migration control method for the single floating drive in the dual-controller, dual-instance tape storage system includes:
[0005] The global state of the dual-controller dual-instance tape storage system is obtained from the topology storage unit, wherein the global state includes the first health state of the first instance, the second health state of the second instance, the ownership status and third health state of each fixed drive, and the current ownership instance of the floating drive.
[0006] Based on the global state and the preset strategy, the first instance and the second instance are divided into a source instance and a target instance;
[0007] Control the source instance to release the floating driver and generate a takeover record with an era identifier;
[0008] The target instance is controlled to take over the floating drive based on the takeover record with the period identifier.
[0009] In one embodiment, after obtaining the global state of the dual-controller dual-instance tape storage system from the topology storage unit, the process includes:
[0010] Perform a migration pre-check. If the migration pre-check passes, proceed with the step of dividing the first instance and the second instance into a source instance and a target instance based on the global state and a preset strategy.
[0011] The migration pre-check includes one of the following:
[0012] Set the system's automatic scheduling switch to "on";
[0013] Ensure compliance with security constraints, wherein the security constraints include at least one of the following: the floating driver is in a non-isolated state, a non-faulty state, or a non-migrating state;
[0014] It has been confirmed that there are no incomplete migration tasks.
[0015] In one embodiment, dividing the first instance and the second instance into a source instance and a target instance based on the global state and a preset strategy includes:
[0016] The current owner instance of the floating driver is determined as the source instance;
[0017] Based on the first health status, the second health status, the ownership status of each fixed driver, and the third health status, the target instance is calculated in conjunction with the preset strategy, wherein the preset strategy includes at least one of the following: fault compensation strategy, load balancing strategy, and time-sharing scheduling strategy.
[0018] If the calculated target instance is the same as the source instance, the migration process is terminated.
[0019] If the calculated target instance is different from the source instance, then the step of controlling the source instance to release the floating driver is executed.
[0020] In one embodiment, the step of calculating the target instance based on the first health state, the second health state, the ownership status of each of the fixed drivers, and the third health state, combined with the preset strategy, includes any one of the following:
[0021] When the preset strategy is a fault compensation strategy, if the number of fixed drivers in the source instance that are in a fault state reaches a preset fault threshold, and the number of fixed driver faults in the other instance between the first instance and the second instance is less than the number of fixed driver faults in the source instance, then the other instance is determined as the target instance.
[0022] When the preset strategy is a time-sharing scheduling strategy, the current time slice number is obtained, and the instance corresponding to the current time slice number is determined as the target instance according to the pre-configured mapping relationship between time slices and instances.
[0023] When the preset strategy is a load balancing strategy, the queue depth of the unprocessed tasks of the first instance and the second instance is obtained, and the instance with the smaller queue depth is determined as the target instance.
[0024] In one embodiment, controlling the source instance to release the floating driver includes:
[0025] Control the source instance to stop business input / output interactions with the floating driver;
[0026] If the critical tasks in transit of the floating drive are completed, control the source instance to perform cache flushing, metadata persistence, and instance ownership release of the floating drive;
[0027] Generate a release confirmation record based on the driver identifier, source instance identifier, release timestamp, and current master term number;
[0028] The release confirmation record is written to the target area of the topology storage unit to persist the release confirmation record.
[0029] In one embodiment, controlling the target instance to take over the floating drive based on the takeover record with a period identifier includes:
[0030] Verify the takeover record based on the aforementioned period identifier;
[0031] If the verification passes, the target instance is controlled to perform the takeover operation of the floating driver based on the takeover record, and the hardware status verification and metadata consistency verification of the floating driver are completed in sequence.
[0032] After confirming that the target instance has successfully taken over, an atomic update operation is performed on the topology of the floating drive, and the floating drive is included in the business service domain of the target instance.
[0033] In one embodiment, performing an atomic update operation on the home topology of the floating driver includes:
[0034] The current home topology record of the floating drive is read from the topology storage unit, and the current home topology record includes at least the current home instance identifier and the current status identifier of the floating drive;
[0035] Initiate a comparison and exchange operation to update the current home instance identifier to the identifier of the target instance, and update the current status identifier from the migrated status to the allocated status;
[0036] After the comparison and swap operation is successful, the takeover record is cleared from the topology storage unit;
[0037] If the comparison and swap operation fails, a migration failure handling process is triggered, and the floating driver is placed in an isolated state.
[0038] In one embodiment, after controlling the target instance to take over the floating drive based on the takeover record with the period identifier, the method further includes:
[0039] When a switch of the scheduling control module is detected, the new scheduling control module determines whether there are any unfinished migration task records in the topology storage unit.
[0040] If there are incomplete migration task records, query the current status of the source instance and the current status of the target instance corresponding to the migration task record;
[0041] Perform the corresponding recovery operation based on the current state of the source instance and the current state of the target instance;
[0042] The step of performing the corresponding recovery operation based on the current state of the source instance and the current state of the target instance includes at least one of the following:
[0043] If the source instance holds the floating drive and the target instance has not yet taken over, a rollback operation is performed to restore the floating drive to the source instance;
[0044] If the source instance has released the floating driver but the target instance has not yet taken over, then the floating driver is placed in an isolated state and new scheduling is blocked;
[0045] If the source instance has been released and the target instance has been taken over but the topology has not yet been committed, then submit a topology update or perform a reconciliation repair.
[0046] If the source instance has been released, the target instance has been taken over, and the topology has been committed, then the completion status is written and scheduling is resumed.
[0047] In addition, to achieve the above objectives, this application also provides a single floating drive secure migration control device in a dual-controller dual-instance tape storage system. The single floating drive secure migration control device in the dual-controller dual-instance tape storage system includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the single floating drive secure migration control method in the dual-controller dual-instance tape storage system as described above.
[0048] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, storing a program that implements a single floating drive secure migration control method in a dual-controller dual-instance tape storage system. The program that implements the single floating drive secure migration control method in a dual-controller dual-instance tape storage system is executed by a processor to implement the steps of the single floating drive secure migration control method in a dual-controller dual-instance tape storage system as described above.
[0049] This application provides a secure migration control method for a single floating drive in a dual-controller, dual-instance tape storage system. First, the global state of the dual-controller, dual-instance tape storage system is obtained from the topology storage unit. Based on the global state and a preset strategy, the first instance and the second instance are divided into a source instance and a target instance to clarify the cross-instance migration affiliation of a single floating drive. Next, the source instance is controlled to release the floating drive, cutting off its access rights. Simultaneously, a takeover record with a time stamp is generated to verify the validity of scheduling instructions and prevent split-brain scheduling behavior in a dual-controller environment. Finally, the target instance is controlled to take over the floating drive only based on the takeover record with the time stamp, strictly limiting the takeover basis of the target instance and preventing concurrent access to the same drive by the source instance and the target instance from the process.
[0050] In summary, this application overcomes the technical defect of split-brain scheduling risk caused by the concurrent access of the source instance and the target instance to the same drive during the migration of a single floating drive in a dual-controller dual-instance tape storage system. This is achieved by first releasing the source instance and then taking over the target instance according to the takeover record with a time stamp. This step-by-step constraint on the access permissions of the two instances to the floating drive overcomes the technical defect of the dual-controller dual-instance tape storage system. Attached Figure Description
[0051] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 A flowchart illustrating an embodiment of the secure migration control method for a single floating drive in a dual-controller, dual-instance tape storage system of this application;
[0054] Figure 2 This is a schematic diagram of the overall migration process provided in Embodiment 1 of the secure migration control method for a single floating drive in a dual-controller dual-instance tape storage system of this application.
[0055] Figure 3 This is a flowchart illustrating the process of determining the target instance in Embodiment 3 of the secure migration control method for a single floating drive in a dual-controller dual-instance tape storage system of this application.
[0056] Figure 4 This is a schematic diagram of the migration process provided in Embodiment 7 of the secure migration control method for a single floating drive in a dual-controller dual-instance tape storage system of this application.
[0057] Figure 5 This is a schematic diagram of the floating drive state change provided in Embodiment 7 of the secure migration control method for a single floating drive in a dual-controller dual-instance tape storage system of this application.
[0058] Figure 6 This is a schematic diagram of the recovery process for unfinished migration tasks after master switchover provided in Embodiment 8 of the secure migration control method for a single floating drive in a dual-controller dual-instance tape storage system of this application.
[0059] Figure 7 This is a schematic diagram of the architecture of the dual-controller dual-instance tape storage system in the embodiments of this application;
[0060] Figure 8 This is a schematic diagram of the hardware structure involved in the secure migration control of a single floating drive in the dual-controller dual-instance tape storage system of this application.
[0061] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0062] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0063] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0064] Currently, the common deployment method in tape storage systems is to use a multi-instance shared drive resource pool, i.e., general resource pooling. However, in a dual-controller, dual-instance tape storage system, where only a single floating drive is migrated across instances, using a multi-instance shared drive resource pool can lead to the source and target instances concurrently accessing the same drive, resulting in inconsistent states and potentially causing split-brain scheduling risks.
[0065] The main solution of this application is: to obtain the global state of the dual-controller dual-instance tape storage system from the topology storage unit, wherein the global state includes the first health state of the first instance, the second health state of the second instance, the ownership status and third health state of each fixed drive, and the current ownership instance of the floating drive; to divide the first instance and the second instance into source instance and target instance according to the global state and a preset strategy; to control the source instance to release the floating drive and generate a takeover record with a period identifier; and to control the target instance to take over the floating drive based on the takeover record with the period identifier.
[0066] This application overcomes the technical defect in dual-controller dual-instance tape storage systems where the source instance releases first, and the target instance takes over based on the takeover record with a time stamp. This is achieved by using an orderly control logic where the source instance releases first, and the target instance takes over later based on the takeover record with a time stamp. This step-by-step constraint on the access permissions of the two instances to the floating drive overcomes the technical defect that causes the source instance and target instance to access the same drive concurrently during the migration of a single floating drive, which leads to inconsistent states and the risk of split-brain scheduling. This ensures the consistency of the floating drive's state and the security of scheduling during cross-instance migration, and improves the operational stability of the dual-controller dual-instance tape storage system.
[0067] It should be noted that the executing entity in this embodiment can be a dual-controller dual-instance tape storage system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a single floating drive secure migration control device in a dual-controller dual-instance tape storage system capable of the above functions. This embodiment does not specifically limit it in this way. The following uses a single floating drive secure migration control device in a dual-controller dual-instance tape storage system as an example to describe this embodiment and the following embodiments.
[0068] Based on this, Embodiment 1 of this application proposes a secure migration control method for a single floating drive in a dual-controller, dual-instance magnetic tape storage system. Please refer to... Figure 1 The single floating drive secure migration control method in the dual-controller dual-instance tape storage system includes steps S100~S400:
[0069] Step S100: Obtain the global state of the dual-controller dual-instance tape storage system from the topology storage unit, wherein the global state includes the first health state of the first instance, the second health state of the second instance, the ownership status and third health state of each fixed drive, and the current ownership instance of the floating drive.
[0070] Optionally, the dual-controller dual-instance tape storage system includes a first instance, a second instance, multiple fixed drives, and a single floating drive.
[0071] In this embodiment, the topology storage unit is a hardware module in a dual-controller, dual-instance tape storage system used to persistently store system operation data. The global state is a structured data set characterizing the overall system operation. The first health state is data characterizing the normality of the hardware operation and service of the first instance. The second health state is data characterizing the normality of the hardware operation and service of the second instance. The ownership state is identification data identifying the binding relationship between the fixed drive and its corresponding instance. The third health state is data characterizing the normality of the fixed drive's own hardware operation. The currently owned instance is the instance currently providing service to the floating drive.
[0072] As an optional implementation, the instance status table, driver status table, and home topology table stored in the topology storage unit are obtained. During processing, all data entries within the topology storage unit are traversed and queried to extract the health data of the first and second instances, the binding relationship and hardware status data of fixed drivers, and the home instance data of floating drivers. After execution, the output or obtained result is complete system global status data.
[0073] As another optional implementation, real-time heartbeat data reported by each system component is acquired. During processing, the real-time heartbeat data is compared and fused with the persistent data in the topology storage unit item by item, and abnormal heartbeat data is removed before integrating the valid status information. After execution, the output or obtained result is the verified global status data of the system.
[0074] Step S200: Based on the global state and the preset strategy, the first instance and the second instance are divided into a source instance and a target instance.
[0075] In this embodiment, the preset strategy is the floating drive migration scheduling rule pre-configured by the system. The source instance is the instance that provides business services on the floating drive before migration. The target instance is the instance that provides business services on the floating drive after migration.
[0076] As an optional implementation, global state data and a fault compensation strategy are acquired. During processing, the number of faults in fixed drives in the global state is detected, the current owner instance of the floating drive is determined as the source instance, and the peer instance that meets the fault compensation conditions is determined as the target instance. After execution, the output or obtained result is the partitioning result of the source instance and the target instance.
[0077] As another optional implementation, global state data and time-sharing scheduling or load balancing strategies are obtained. During processing, the instance corresponding to the current time slice number is matched or the task queue depth of the dual instance is compared. The instance currently belonging to the floating driver is determined as the source instance, and the successfully matched instance is determined as the target instance. After execution, the output or obtained result is the partitioning result of the source instance and the target instance.
[0078] Step S300: Control the source instance to release the floating driver and generate a takeover record with an era identifier.
[0079] In this embodiment, release is the operation of the source instance unbinding itself from the service of the floating drive and terminating its access permissions. The period identifier is unique identifier data representing the effective term of the scheduling master. The takeover record is structured credential data used by the target instance to legally take over the floating drive.
[0080] As an optional implementation, the source instance identifier and the floating drive identifier are obtained. During the process, the source instance is controlled to stop business input / output interactions with the floating drive. After completing cache flushing and metadata persistence, ownership binding is released. A period identifier is generated based on the current master control term number. This period identifier is then combined with the source instance information, target instance information, and floating drive information to encapsulate and generate a takeover record. Upon completion, the output or obtained result is a source instance release completion signal and a takeover record with a period identifier.
[0081] As an alternative implementation, the running status of the source instance and the hardware status data of the floating driver are obtained. During processing, the preconditions for releasing the source instance are first verified. Once the conditions are met, the ownership release operation of the floating driver is executed. An era identifier is generated based on the system scheduling cycle, and the driver identifier and instance identifier are integrated to generate a takeover record. After execution, the output or obtained result is a source instance release completion signal and a takeover record with an era identifier.
[0082] Step S400: Control the target instance to take over the floating drive based on the takeover record with the period identifier.
[0083] In this embodiment, takeover is the operation by which the target instance establishes a business binding with the floating driver and obtains access permissions. Period identifier verification is the operation by which the target instance verifies the legality and validity of the takeover record.
[0084] As an optional implementation, a takeover record with a period identifier and a target instance identifier are obtained. During processing, the target instance verifies whether the period identifier is a valid identifier for the current system. If the verification is successful, the floating driver hardware connection is executed. After completing hardware status detection and metadata comparison, the takeover operation is performed. Upon completion, the output or obtained result is a target instance takeover completion signal.
[0085] As an alternative implementation, the takeover record and floating driver hardware parameters are obtained. During the process, the target instance first verifies the data integrity of the takeover record, then verifies the validity of the period identifier, and sequentially completes hardware status verification and metadata consistency verification before executing the takeover operation. After execution, the output or obtained result is the target instance takeover completion signal.
[0086] For example, in a dual-controller, dual-instance tape storage system under normal operation, the system scheduling module obtains health status data of the first and second instances, ownership relationships and hardware health status data of multiple fixed drives, and the current ownership instance data of the floating drives from the topology storage unit. The system detects a hardware failure in the fixed drive bound to the first instance, which cannot meet business operation requirements. Based on the fault compensation strategy, the first instance is designated as the source instance, and the second instance as the target instance. The scheduling module issues a release command to the source instance, which stops all business input / output interactions with the floating drives, completes cache data flushing and metadata persistent storage, and releases its ownership binding with the floating drives. The scheduling module generates a time period identifier based on the current valid master controller term number, encapsulates the source instance information, target instance information, and floating drive identifier to generate a takeover record with that time period identifier. After obtaining the takeover record, the target instance first verifies the validity of the time period identifier. If the verification passes, it sequentially completes the hardware status detection and metadata consistency comparison of the floating drives, successfully establishes a business binding relationship with the floating drives, completes the takeover operation of the floating drives, and the system synchronously updates the ownership topology data of the floating drives.
[0087] Optionally, combined Figure 2 The overall implementation process of this embodiment will be described. Step S10: Obtain the system status.
[0088] In this embodiment, the system status is the global operating status data of the dual-controller dual-instance tape storage system, including the first health status of the first instance, the second health status of the second instance, the ownership status and third health status of each fixed drive, and the current ownership instance of the floating drive.
[0089] As an optional implementation, the instance status table, driver status table, and home topology table stored in the topology storage unit are obtained. During processing, the scheduling control module traverses and queries all data entries within the topology storage unit, extracting the health data of the first and second instances, the binding relationship and hardware status data of fixed drivers, and the home instance data of floating drivers. After execution, the output or obtained result is complete system global status data.
[0090] As an alternative implementation, real-time heartbeat data reported by each system component is acquired. During processing, the scheduling and control module compares and merges the real-time heartbeat data with the persistent data in the topology storage unit item by item, eliminating abnormal heartbeat data and integrating valid status information. After execution, the output or obtained result is the verified global system status data.
[0091] Step S20: Determine the target instance.
[0092] In this embodiment, the target instance is the instance to be assigned after the floating driver completes the cross-instance migration operation, and the source instance is the instance to which the floating driver currently belongs.
[0093] As an optional implementation, system global status data and fault compensation strategy configuration parameters are obtained. During processing, the scheduling control module determines the current home instance of the floating driver as the source instance, counts the number of fixed drivers in fault state under dual-instance conditions, and determines the peer instance that meets the fault compensation conditions as the target instance. After execution, the output or obtained result is the partitioning result of the source instance and the target instance.
[0094] As another optional implementation, system global status data and time-sharing scheduling or load balancing strategy configuration parameters are obtained. During processing, the scheduling control module determines the currently owned instance of the floating driver as the source instance, matches the instance corresponding to the current time slice number or compares the task queue depth of the two instances, and determines the successfully matched instance as the target instance. After execution, the output or obtained result is the partitioning result of the source instance and the target instance.
[0095] After step S20 is completed, the system performs a migration determination operation. If the determination result is negative, the system directly terminates the current migration process and enters the end state. If the determination result is positive, the system performs the subsequent operation of freezing the source instance.
[0096] In this embodiment, the determination of whether to migrate is based on the identity of the target instance and the source instance, combined with the migration pre-check results.
[0097] As an optional implementation, source instance identifier information and target instance identifier information are obtained. During processing, a character equality comparison is performed between the source instance identifier and the target instance identifier, and the migration pre-check result is verified. If both conditions are met, it is determined as yes; otherwise, it is determined as no. After execution, the output or obtained result is the migration determination result.
[0098] As an alternative implementation, the source instance identity information, target instance identity information, and migration pre-check status data are obtained. During processing, the system comparison engine verifies the identity information of both instances, simultaneously verifies the compliance of the migration pre-check, and generates a migration determination result. After execution, the output or obtained result is a determination signal for migration execution or termination.
[0099] Step S30: Freeze the source instance.
[0100] In this embodiment, freezing the source instance is an operation that controls the source instance to stop all business input / output interactions with the floating driver and wait for the convergence of critical tasks in transit.
[0101] As an optional implementation, the source instance scheduling command and the floating driver identifier are obtained. During processing, the migration orchestration module issues an interaction termination command to the source instance, blocking all data read / write requests and control commands sent by the source instance to the floating driver, and continuously monitoring the execution progress of critical tasks in transit. Upon completion, the output or result is a signal indicating that the business input / output interaction between the source instance and the floating driver has completely terminated and the in-transit tasks have converged.
[0102] As an alternative implementation, the configuration information of the source instance's service interaction port is obtained. During processing, the migration orchestration module closes the service interaction ports corresponding to the source instance and the floating driver, blocking all data transmission channels between them, and waits for the critical tasks in transit to complete. After completion, the output or result is a signal indicating that the service interaction channels between the source instance and the floating driver are completely closed and the task convergence is complete.
[0103] Step S40: Release ownership of the source instance.
[0104] In this embodiment, releasing ownership of the source instance is the operation performed by the source instance to flush the floating drive cache to disk, persist metadata, and remove the business binding relationship.
[0105] As an optional implementation, the source instance identifier, floating drive identifier, and current master term number are obtained. During processing, the migration orchestration module controls the source instance to sequentially perform cache flushing, metadata persistence, and instance ownership release operations. Based on the drive identifier, source instance identifier, release timestamp, and current master term number, a release confirmation record is generated and written to the target area of the topology storage unit. After execution, the output or obtained result is a source instance ownership release completion signal.
[0106] As an alternative implementation, the source instance's runtime status data and the floating driver's hardware status data are obtained. During processing, the migration orchestration module first verifies the preconditions for releasing the source instance. Once the conditions are met, the floating driver ownership release operation is executed, generating a release confirmation record with a verification field and persisting it. After execution, the output or result is feedback information confirming the completion of the source instance ownership release.
[0107] After step S40 is completed, the system performs a determination operation to determine whether the release was successful. If the determination result is negative, the system performs the operation of setting the instance to an isolated state. If the determination result is positive, the system performs the subsequent takeover operation for the target instance.
[0108] In this embodiment, the determination of whether the release was successful is based on the persistence status of the release confirmation record and the release status of the source instance ownership.
[0109] As an optional implementation, persistent status data of the release confirmation record and ownership status data of the source instance are obtained. During processing, the scheduling control module reads the release confirmation record from the topology storage unit, verifies the data integrity and validity, and simultaneously detects the ownership status of the source instance. After execution, the output or obtained result is a determination signal indicating whether the release was successful or failed.
[0110] As an alternative implementation, source instance status feedback data is obtained, and topology storage unit write result data is processed. During the process, the scheduling control module synchronously verifies the source instance status feedback and the topology storage unit write result, generating a release success or failure determination result. After execution, the output or obtained result is a release status determination signal.
[0111] Step S50: The target instance takes over.
[0112] In this embodiment, target instance takeover is the operation in which the target instance completes the hardware connection, status verification and service binding of the floating driver based on the takeover record with period identifier.
[0113] As an optional implementation, a takeover record with a period identifier and a target instance identifier are obtained. During processing, the target instance verifies the validity of the period identifier in the takeover record. If the verification passes, the hardware connection of the floating driver is executed, and hardware status verification and metadata consistency verification are completed sequentially. After execution, the output or obtained result is the target instance takeover completion signal.
[0114] As an alternative implementation, complete takeover record data and floating driver hardware parameters are obtained. During processing, the target instance first verifies the data integrity of the takeover record, then verifies the validity of the period identifier. After completing the dual verification, the takeover operation of the floating driver is executed. Upon completion, the output or obtained result is the target instance takeover ready signal.
[0115] After step S50 is completed, the system performs a check to determine whether the takeover was successful. If the result is negative, the system sets the system to an isolated state. If the result is positive, the system performs subsequent atomic topology updates.
[0116] In this embodiment, the determination of whether the takeover was successful is based on the hardware status verification result and the metadata consistency verification result.
[0117] As an optional implementation, hardware status detection data and metadata comparison results of the floating driver are acquired. During processing, the scheduling control module reads the hardware status detection results and metadata comparison results. If both results are normal, the takeover is considered successful; otherwise, the takeover is considered a failure. After execution, the output or obtained result is a takeover success or failure determination signal.
[0118] As an alternative implementation, target instance takeover status feedback data is obtained. During processing, the scheduling control module receives status feedback from the target instance and generates a judgment result indicating whether the takeover was successful or failed. After execution, the output or obtained result is the takeover status judgment signal.
[0119] Step S60: Atomic update of the topology.
[0120] In this embodiment, atomic topology update is to perform an indivisible update operation on the home topology of the floating driver, thereby completing the switching of the home instance.
[0121] As an optional implementation, the floating drive's home topology record and target instance identifier are obtained. During processing, the scheduling control module reads the floating drive's current home topology record from the topology storage unit, initiates a comparison and exchange operation, updates the current home instance identifier to the target instance's identifier, updates the current status identifier from the migrated state to the allocated state, and clears the takeover record in the topology storage unit after successful operation. After execution, the output or obtained result is a home topology atomic update completion signal.
[0122] As another optional implementation, topology data from the topology storage unit is obtained, and configuration parameters are atomically updated. During processing, the scheduling control module performs an indivisible update of the home topology based on the configuration parameters, completes the synchronous replacement of instance identifiers and status identifiers, and clears the takeover record corresponding to this migration after a successful update. After execution, the output or obtained result is feedback information indicating that the home topology update is complete.
[0123] After step S60 is completed, the system determines that the migration is complete and enters the end state.
[0124] Step S70: Set to isolated state.
[0125] In this embodiment, setting it to isolated state is an operation that marks the floating driver as a protected state that prohibits it from participating in any automatic migration and service scheduling.
[0126] As an optional implementation, the migration failure trigger signal and the floating driver identifier are obtained. During processing, the status management module rewrites the floating driver's status identifier to an isolated state, issues a scheduling blocking command, and blocks all instances' service scheduling and automatic migration permissions for this driver. After execution, the output or obtained result is a protection signal indicating that the floating driver has entered the isolated state.
[0127] As an alternative implementation, migration failure status data and system isolation control rules are obtained. During processing, the status management module locks all scheduling interfaces of the floating drive according to the isolation control rules, marks them with a prohibited scheduling flag, and blocks all new scheduling requests. After execution, the output or result is a fault signal indicating that the floating drive's scheduling permissions are completely locked.
[0128] After step S70 is completed, the system enters the termination state.
[0129] For example, after the dual-controller, dual-instance tape storage system starts up, the scheduling control module obtains the global system status from the topology storage unit, including the health status of the first and second instances, the ownership and health status of fixed drives, and the current ownership instance of floating drives. The scheduling control module determines the target instance based on the fault compensation strategy, marking the first instance (current ownership instance of the floating drive) as the source instance and calculating the second instance as the target instance. The system determines that the target instance and the source instance are different instances and that the migration pre-check passes, then determines that migration is required and issues a migration command to the migration orchestration module. The migration orchestration module controls the first instance to stop all business input / output interactions with the floating drive, waits for the convergence of in-transit critical tasks, and then performs cache flushing, metadata persistence, and instance ownership release operations for the floating drive, generating a release confirmation record and writing it to the topology storage unit. After the system verifies the successful release of the source instance, it controls the second instance to perform a floating drive takeover operation based on the takeover record with a valid period identifier, sequentially completing hardware status verification and metadata consistency verification. After the system verifies the successful takeover of the second instance, it initiates a comparison and exchange operation to complete the atomic update of the floating drive's ownership topology, clears the takeover record, and the migration is complete, with the system entering the termination state. If the source instance fails to release or the target instance fails to take over, the system immediately puts the floating driver into an isolated state, blocking all new scheduling requests, and then enters the terminated state.
[0130] This embodiment establishes a complete control logic that includes acquiring system status, identifying the target instance, freezing the source instance, releasing ownership of the source instance, taking over the target instance, and atomically updating the topology. Combined with multi-node status determination and failure isolation mechanisms, it strictly constrains the access timing and status change process of floating drives during cross-instance migration. This completely avoids the inconsistency problem caused by the source instance and target instance concurrently accessing the same drive. It eliminates the risk of split-brain scheduling in a dual-controller environment through period identifier verification and atomic update operations. At the same time, it initiates isolation protection to prevent the spread of faults when migration fails, thereby improving the operational stability, data security, and fault tolerance of the dual-controller dual-instance tape storage system.
[0131] This embodiment first obtains the global state to divide the source instance and target instance, then controls the source instance to release the floating drive and generate a takeover record with a time stamp. Finally, the target instance completes the takeover of the floating drive based on the valid takeover record. This strictly limits the access sequence of the two instances to the floating drive, blocks the concurrent access behavior of the source instance and the target instance, avoids the inconsistency of state during the migration process, and ensures the validity of the scheduling instructions through time stamp verification. This eliminates the risk of split-brain scheduling in the dual-controller environment and improves the operational stability and data security of the dual-controller dual-instance tape storage system.
[0132] Based on any of the above embodiments, in Embodiment 2 of this application, after obtaining the global state of the dual-controller dual-instance tape storage system from the topology storage unit, the process includes:
[0133] Step A10: Perform a migration pre-check. If the migration pre-check passes, perform the step of dividing the first instance and the second instance into source instances and target instances according to the global state and the preset strategy.
[0134] Optionally, the pre-migration check includes at least one of the following: determining that the system's automatic scheduling switch is on, determining that security constraints are compliant, and determining that there are no incomplete migration tasks.
[0135] In this embodiment, the migration pre-check is a legality and security verification operation performed before the floating drive initiates a cross-instance migration. The system automatic scheduling switch is a status indicator that controls the system to enable or disable the automatic cross-instance migration function of floating drives. Security constraints are state restriction rules that determine whether a floating drive meets the migration prerequisites. Incomplete migration tasks are floating drive migration scheduling tasks that have not been completed or terminated within the system.
[0136] As an optional implementation, the system's automatic scheduling switch status data, floating driver operating status data, and migration task queue data are acquired. During processing, the status flags of the system's automatic scheduling switch are read sequentially, the floating driver is checked to see if it is in a non-isolated, non-faulty, and non-migration state, the migration task queue is traversed to check for any unfinished task entries, and Boolean AND logic judgments are performed on the three verification conditions. After execution, the output or obtained result is either a migration pre-check pass signal or a migration pre-check fail signal.
[0137] As an alternative implementation, the scheduling configuration data, driver status data, and migration task log data stored in the topology storage unit are acquired. During processing, relevant data in the topology storage unit are read in batches, and automatic scheduling switch status verification, security constraint compliance verification, and existence verification of incomplete migration tasks are performed simultaneously. Parallel verification logic is used to improve inspection efficiency. After execution, the output or result is either a migration pre-check pass signal that triggers the subsequent instance partitioning process, or a migration pre-check fail signal that terminates the current migration scheduling.
[0138] For example, after completing the global status acquisition, the dual-controller dual-instance tape storage system initiates the pre-verification process for floating drive migration. The system reads the on / off status of the automatic scheduling switch, verifies that the floating drive is in a normal, usable, non-isolated, non-faulty, and non-migrating state, and traverses the migration task queue to confirm that there are no incomplete migration tasks. After all three verification conditions are met, the migration pre-check passes, and the system performs the source instance and target instance partitioning operation for the first and second instances based on the global status and the preset fault compensation strategy.
[0139] This embodiment performs multi-dimensional migration pre-checks before migrating floating drives across instances, strictly limiting the legal conditions for initiating migrations. This prevents invalid migrations from being initiated when the scheduling function is disabled, the drive status is abnormal, or there are incomplete migration tasks. This ensures the security of the migration process and the rationality of its execution, providing a stable foundation for the accurate division of source and target instances in the future.
[0140] Based on any of the above embodiments, in Embodiment 3 of this application, the first instance and the second instance are divided into a source instance and a target instance according to the global state and a preset strategy, including:
[0141] Step S21: Determine the current home instance of the floating driver as the source instance.
[0142] In this embodiment, the currently owned instance is the instance that the floating driver has currently established a business binding relationship with and is hosting the business service. The source instance is the instance that the floating driver belonged to before performing the cross-instance migration operation.
[0143] As an optional implementation, the topology data of the floating drives stored in the topology storage unit is obtained. During processing, the owner instance identifier field corresponding to the floating drive is read, and the instance corresponding to this identifier field is directly marked as the source instance. After execution, the output or obtained result is the unique identifier information of the source instance.
[0144] As an alternative implementation, business binding log data of the first and second instances is obtained. During processing, the business binding logs of both instances are traversed, and instances with real-time business associations with the floating driver are selected and identified as the source instances. After execution, the output or obtained result is the instance identity information of the source instance.
[0145] Step S22: Calculate the target instance based on the first health status, the second health status, the ownership status of each fixed driver, and the third health status, in conjunction with the preset strategy. The preset strategy includes at least one of the following: fault compensation strategy, load balancing strategy, and time-sharing scheduling strategy.
[0146] In this embodiment, the target instance is the instance to be assigned to after the floating driver completes the cross-instance migration operation. The fault compensation strategy matches the target instance's scheduling rules based on the fixed driver's fault status. The load balancing strategy matches the target instance's scheduling rules based on the dual-instance service load status. The time-sharing scheduling strategy matches the target instance's scheduling rules based on a preset time period.
[0147] As an optional implementation, the following data are acquired: first health status data, second health status data, fixed driver ownership status data, fixed driver third health status data, and fault compensation strategy configuration parameters. During processing, the number of fixed drivers in fault status under dual-instance conditions is counted, and instances with a number of faulty drivers that meet the compensation condition are calculated as target instances. After execution, the output or obtained result is the unique identifier information of the target instance.
[0148] As an alternative implementation, the system obtains dual-instance service load data, system time slice data, and configuration parameters for load balancing and time-sharing scheduling strategies. During processing, the target instance is determined by prioritizing comparison of the dual-instance service load depth; if no load balancing is required, the target instance is matched based on the current time slice number. After execution, the output or obtained result is the instance identity information of the target instance.
[0149] Step S23: If the calculated target instance is the same as the source instance, then the migration process is terminated.
[0150] In this embodiment, the migration process is the complete execution flow of the floating driver switching ownership from the source instance to the target instance. Terminating the migration process means stopping all subsequent operations of this migration scheduling and restoring the system to its initial scheduling state.
[0151] As an optional implementation, source instance identifier information and target instance identifier information are obtained. During processing, a character-for-character equality comparison is performed between the source instance identifier and the target instance identifier. If the comparison results match, a migration termination command is issued. After execution, the output or obtained result is the migration process termination signal.
[0152] As an alternative implementation, the source instance identity information and the target instance identity information are obtained. During processing, the identity information of both instances is entered into the system's comparison engine for matching and verification. If the verification result shows that they are the same instance, the migration scheduling task is cleared. After execution, the output or obtained result is a migration scheduling task clear signal.
[0153] Step S24: If the calculated target instance is different from the source instance, then execute the step of controlling the source instance to release the floating driver.
[0154] In this embodiment, release is the operation of removing the source instance from the business binding relationship with the floating driver and terminating permission access.
[0155] As an optional implementation, source instance identifier information and target instance identifier information are obtained. During processing, the source instance identifier and target instance identifier are compared character by character. If the comparison results are inconsistent, the source instance release driver instruction is triggered. After execution, the output or obtained result is the source instance release operation start signal.
[0156] As an alternative implementation, the source instance identity information and the target instance identity information are obtained. During processing, the system comparison engine verifies the dual instance identity information. If the verification result indicates that they are different instances, a floating driver release execution command is issued. After execution, the output or obtained result is the floating driver release process start signal.
[0157] For example, after completing the migration pre-check, the dual-controller, dual-instance tape storage system first reads the current owner instance of the floating drive from the topology storage unit as the first instance, and determines the first instance as the source instance for this migration. The system retrieves the health status data of the first and second instances, counts the owner status and hardware health status of each fixed drive, calculates the target instance based on the fault compensation strategy, and detects that the fixed drive bound to the first instance has a fault, while all fixed drives of the second instance are in normal condition, so the second instance is calculated as the target instance. The system compares the identification information of the source instance and the target instance, determines that they are different instances, and then issues an execution command to initiate the subsequent operation of the source instance releasing the floating drive.
[0158] Optionally, refer to Figure 3This document describes a feasible method for determining the target instance. After system startup, three pre-emptive security checks are performed sequentially: checking whether security constraints are met, whether the automatic scheduling switch is enabled, and whether there are any incomplete migration tasks. If any check fails, the process terminates directly. After all three checks pass, the system first determines whether there is a fault compensation requirement: if so, the target instance is determined directly through the fault compensation strategy; if not, the system selects whether to perform target instance recalculation based on whether fast recalculation is triggered, and then checks the load balancing conditions. Based on the load balancing condition check results, the system matches the corresponding scheduling strategy: if satisfied, the load balancing strategy is used; otherwise, a strict time-sharing strategy is used. Finally, the system summarizes the calculation results of fault compensation, load balancing, or strict time-sharing strategies, generates the target instance, and then performs a secure migration of the floating driver, completing the target instance determination process.
[0159] This embodiment accurately completes the division of migration objects and the determination of migration triggers by determining the source instance step by step, calculating the target instance by combining multiple strategies, and performing instance identity verification and branch processing. This avoids the execution of invalid migration processes, improves the accuracy and rationality of floating driver migration scheduling, and provides accurate instance pointing basis for subsequent safe migration operations.
[0160] Based on any of the above embodiments, in Embodiment 4 of this application, the target instance is calculated according to the first health status, the second health status, the ownership status of each fixed driver, and the third health status, combined with the preset strategy, including any of the following:
[0161] Step S221: When the preset strategy is a fault compensation strategy, if the number of fixed drivers in the source instance that are in a fault state reaches a preset fault threshold, and the number of fixed driver faults in the other instance between the first instance and the second instance is less than the number of fixed driver faults in the source instance, then the other instance is determined as the target instance.
[0162] In this embodiment, the fault compensation strategy is a scheduling rule that matches the target home instance for the floating driver based on the fault status of the fixed driver. The preset fault threshold is a critical value representing the number of fixed driver faults that require floating driver fault compensation for the source instance. The fault status is a state where the fixed driver hardware is malfunctioning and unable to provide service.
[0163] As an optional implementation, the following methods are used: acquiring the number of fixed driver failures in the source instance, the number of fixed driver failures in another instance, and a preset failure threshold value. During processing, the number of failed drivers in the source instance is compared with the preset failure threshold, and the difference between the number of failed drivers in the source instance and the other instance is compared. If both comparison conditions are met, the other instance is marked as the target instance. After execution, the output or obtained result is the unique identifier information of the target instance.
[0164] As an alternative implementation, hardware status detection data of the fixed drivers in both instances and preset fault threshold configuration data are obtained. During processing, the actual number of faulty drivers in each of the two instances is counted. First, it is determined whether the source instance meets the fault compensation triggering condition. Then, the difference in the number of faulty drivers between the two instances is compared. If the condition is met, the other instance is identified as the target instance. After execution, the output or obtained result is the instance identity information of the target instance.
[0165] Step S222: When the preset strategy is a time-sharing scheduling strategy, obtain the current time slice number, and determine the instance corresponding to the current time slice number as the target instance according to the pre-configured mapping relationship between time slices and instances.
[0166] In this embodiment, the time-sharing scheduling strategy is a scheduling rule that allocates target-owned instances to floating drivers based on a preset time slice period. The time slice number is a unique sequence number representing the current runtime segment of the system. The mapping relationship between time slices and instances is a pre-configured binding relationship between time slice numbers and either a first instance or a second instance.
[0167] As an optional implementation, the system's current clock data, pre-configured time slices, and instance mapping table are obtained. During processing, the current time slice number is extracted based on the system clock, the mapping table is traversed to match the corresponding instance, and the successfully matched instance is determined as the target instance. After execution, the output or obtained result is the unique identifier information of the target instance.
[0168] As an alternative implementation, real-time time slice count data and instance-time slice binding configuration parameters are obtained. During processing, the current cumulative time slice number is read, and the corresponding belonging instance is locked according to a preset mapping rule, making that instance the target instance of the floating driver. After execution, the output or obtained result is the instance identity information of the target instance.
[0169] Step S223: When the preset strategy is a load balancing strategy, obtain the task queue depth of the first instance and the second instance, and determine the instance with the smaller queue depth as the target instance.
[0170] In this embodiment, the load balancing strategy matches the floating drive with the target home instance based on the dual-instance service load status. The pending task queue depth is the total number of tape storage service tasks waiting to be executed on the instance.
[0171] As an optional implementation, the task queue data for the first instance and the task queue data for the second instance are obtained. During processing, the total number of tasks in both instance task queues is counted, and the depth values of the two queues are compared. The instance with the smaller value is identified as the target instance. After execution, the output or obtained result is the unique identifier information of the target instance.
[0172] As another optional implementation, real-time reported data and queue depth statistics rules for dual-instance business tasks are obtained. During processing, the queue depth of tasks to be processed in both instances is calculated according to preset rules. The instance with lower load is selected by numerical comparison and identified as the target instance. After execution, the output or obtained result is the instance identity information of the target instance.
[0173] For example, a dual-controller, dual-instance tape storage system can switch between different scheduling strategies when calculating the target instance for floating drives, depending on the actual scenario. If the system detects that the number of faults in the fixed drives bound to the source instance has reached a preset fault threshold, while all fixed drives in the other instance are operating normally, a fault compensation strategy is used to identify the other instance as the target instance. When the system runs in hourly time slices, it obtains the current time slice number and identifies the corresponding instance as the target instance based on a pre-configured mapping relationship. During peak system traffic, the system collects the queue depths of pending tasks for both the first and second instances, identifying the lower-load instance with the smaller queue depth as the target instance, thus completing the accurate calculation of the target instance.
[0174] This embodiment uses three refined strategies—fault compensation, time-sharing scheduling, and load balancing—to calculate the target instance, adapting to different application scenarios of fault compensation, time-sharing sharing, and load balancing in dual-controller dual-instance tape storage systems. This improves the flexibility and adaptability of target instance determination, ensures that floating drive migration scheduling meets the actual operating needs of the system, and optimizes system resource allocation efficiency.
[0175] Based on any of the above embodiments, in Embodiment 5 of this application, controlling the source instance to release the floating driver includes:
[0176] Step S31: Control the source instance to stop business input / output interactions with the floating driver.
[0177] In this embodiment, business input / output interaction is a set of operations between the source instance and the floating driver, including business data reading and writing, control command transmission, and status information exchange.
[0178] As an optional implementation, the source instance scheduling command and the floating driver identifier are obtained. During processing, an interaction termination command is issued to the source instance, blocking all data read / write requests and control commands sent by the source instance to the floating driver. After execution, the output or result is a signal indicating that the service input / output interaction between the source instance and the floating driver has completely terminated.
[0179] As an alternative implementation, the service interaction port configuration information of the source instance is obtained. During the process, the service interaction ports corresponding to the source instance and the floating driver are closed, blocking all data transmission channels between them. After execution, the output or result is a signal indicating that the service interaction channels between the source instance and the floating driver are completely closed.
[0180] Step S32: If the in-transit critical task of the floating drive is completed, control the source instance to perform cache flushing, metadata persistence, and instance ownership release of the floating drive.
[0181] In this embodiment, "in-transit critical tasks" refers to core business processing tasks that have been deployed to the floating drive by the source instance but have not yet been completed. "Convergence complete" is the state where all in-transit critical tasks have been completed and there is no data pending processing. "Cache flushing" is the operation of writing temporary business data in the floating drive cache to the tape medium. "Metadata persistence" is the operation of writing metadata such as the floating drive's operating parameters, status information, and business configuration to the persistent storage medium. "Instance ownership release" is the operation of releasing the source instance's control permissions over the floating drive and its business binding relationship.
[0182] As an optional implementation, the execution status data of the critical tasks in transit for the floating drive is obtained. During processing, the execution progress of the critical tasks in transit is monitored in real time. After confirming that all tasks have been completed, the cache flushing operation, metadata writing operation, and ownership release operation are initiated sequentially. Upon completion, the output or obtained result is a signal indicating that the source instance has completed the ownership release of the floating drive.
[0183] As an alternative implementation, status feedback data from the source instance task management module is obtained. During processing, a convergence completion instruction sent by the task management module is received, and cache data synchronous writing, metadata solidification storage, and business binding relationship removal are executed sequentially. After execution, the output or obtained result is a status signal indicating that the floating driver has been removed from the source instance's control.
[0184] Step S33: Generate a release confirmation record based on the driver identifier, source instance identifier, release timestamp, and current master term number.
[0185] In this embodiment, the driver identifier is coded information that uniquely identifies the floating driver. The source instance identifier is coded information that uniquely identifies the source instance. The release timestamp is the time data that records the moment when the source instance completes the release of ownership. The current master term number is a unique number that represents the current effective scheduling master term of the system. The release confirmation record is structured credential data that records the source instance's completion of the floating driver release operation.
[0186] As an optional implementation, the driver identifier, source instance identifier, release timestamp, and current master term number are obtained. During processing, these four types of information are sequentially concatenated and encapsulated according to a preset data format to generate a standardized release confirmation record. After execution, the output or obtained result is the complete release confirmation record data.
[0187] As an alternative implementation, four types of core identifiers and time data are acquired, along with record encapsulation rules. During processing, each type of data is validated and encoded according to the encapsulation rules, generating a release confirmation record with a validation field. Upon completion, the output or obtained result is release confirmation record data with a validation mechanism.
[0188] Step S34: Write the release confirmation record to the target area of the topology storage unit to persist the release confirmation record.
[0189] In this embodiment, the target area of the topology storage unit is a pre-divided dedicated storage block within the topology storage unit for storing release confirmation records. Persistence is the operation of writing data to a non-volatile storage medium to achieve long-term storage.
[0190] As an optional implementation, the release confirmation record data and the target area address information of the topology storage unit are obtained. During processing, the release confirmation record is written to the designated storage block according to the target area address, completing the data persistence storage. After execution, the output or obtained result is a release confirmation record persistence completion signal.
[0191] As an alternative implementation, release confirmation record data and topology storage unit storage configuration information are obtained. During processing, the target area is located by matching the storage configuration information, a write operation is performed, and the storage result is fed back. After execution, the output or obtained result is feedback information indicating that the release confirmation record was successfully stored.
[0192] For example, after identifying the source instance and the target instance, the dual-controller, dual-instance tape storage system first controls the source instance to stop all business input / output interactions with the floating drive, blocking data transmission between them. The system continuously monitors the critical tasks en route to the floating drive. Once all critical tasks are completed (convergence achieved), the system controls the source instance to sequentially flush the temporary data in the floating drive's cache to the tape media, persistently storing the floating drive's metadata, and finally releasing the source instance's ownership of the floating drive. Based on the floating drive identifier, source instance identifier, timestamp of the release operation, and the current master controller's term number, the system generates a standardized release confirmation record and writes this record to the dedicated target area of the topology storage unit, completing the persistent storage of the release confirmation record and providing valid credentials for subsequent takeover by the target instance.
[0193] This embodiment completes the secure release operation of the floating drive by the source instance by terminating business interaction in stages, waiting for task convergence, executing data solidification and permission release, and generating and persisting the release confirmation record. This ensures that there is no data loss or state chaos during the release process, while retaining traceable release credentials. It avoids concurrent access to the floating drive by the source instance and the target instance from the source, and ensures the consistency of the state during the migration process.
[0194] Based on any of the above embodiments, in Embodiment Six of this application, controlling the target instance to take over the floating drive based on the takeover record with the period identifier includes:
[0195] Step S41: Verify the takeover record based on the period identifier.
[0196] In this embodiment, the period identifier is a unique coded data representing the current effective term of the system's scheduling master. The takeover record is the core credential data for a target instance to legally take over a floating drive. Verification is a judgment operation that checks the validity of the period identifier and the completeness of the takeover record.
[0197] As an optional implementation, the period identifier data carried in the takeover record and the system's current valid period identifier data are obtained. During processing, the period identifier in the takeover record and the system's current valid period identifier are compared character-for-character to determine if they are completely identical. After execution, the output or obtained result is a verification pass signal or a verification fail signal.
[0198] As an alternative implementation, complete takeover record data and period identifier validity determination rules are obtained. During processing, the data integrity of the takeover record is first verified, and then the period identifier is verified to be within its valid term according to the determination rules. Upon completion, the output or obtained result is either a verification pass signal or a verification fail signal.
[0199] Step S42: If the verification passes, control the target instance to perform the takeover operation of the floating driver based on the takeover record, and sequentially complete the hardware status verification and metadata consistency verification of the floating driver.
[0200] In this embodiment, the takeover operation is the process by which the target instance establishes a control connection with the floating drive and obtains service permissions. Hardware status verification is a detection operation to check whether the floating drive hardware is operating normally. Metadata consistency verification is a check operation to compare whether the metadata stored in the topology storage unit matches the metadata of the floating drive media.
[0201] As an optional implementation, the takeover record and target instance control instructions that have passed verification are obtained. During processing, the target instance establishes a communication connection with the floating drive based on the takeover record, first checks the hardware operating parameters of the floating drive to determine its hardware status, and then extracts the metadata of the topology storage unit and the floating drive for item-by-item comparison. After execution, the output or obtained result is a verification completion signal indicating that the hardware status is normal and the metadata is consistent.
[0202] As an alternative implementation, hardware detection data and metadata comparison rules for the floating driver are acquired. During processing, after the target instance completes the takeover connection, it calls the hardware detection module to verify the driver's operating status and performs metadata matching verification according to the comparison rules. After execution, the output or obtained result is a takeover ready signal indicating that both verifications have passed.
[0203] Step S43: After confirming that the target instance has successfully taken over, perform an atomic update operation on the topology of the floating drive and include the floating drive in the service domain of the target instance.
[0204] In this embodiment, the home topology is the topology data that records the binding relationship between the floating driver and the corresponding instance. An atomic update operation is an indivisible, one-time data update operation, ensuring that the update process has no intermediate states. The service domain is the service scope within which the target instance can issue service instructions and carry services.
[0205] As an optional implementation, the process involves acquiring a takeover success signal and the floating driver's home topology data. During processing, an atomic update command is initiated to replace the source instance identifier in the home topology with the target instance identifier. After the update is complete, a service access command is issued to the target instance. Upon completion, the output or result is a signal indicating that the home topology update is complete and the floating driver has been included in the service domain.
[0206] As an alternative implementation, topology data of the topology storage unit is obtained, and configuration parameters are atomically updated. During processing, an indivisible update of the topology is performed according to the configuration parameters. After a successful update, the target instance is granted service scheduling permissions to the floating drive. Upon completion, the output or result is a service availability signal indicating that the floating drive has officially belonged to the target instance.
[0207] For example, after a dual-controller, dual-instance tape storage system obtains a takeover record with a period identifier, it first compares and verifies the period identifier in the takeover record with the system's currently valid period identifier. If the verification passes, it controls the target instance to establish a connection with the floating drive based on the takeover record. The target instance first checks the hardware operating status of the floating drive. After confirming that the hardware is normal, it compares the metadata stored in the topology storage unit with the metadata in the floating drive's own media, confirming that the two are completely consistent. After the system confirms that the target instance has successfully taken over, it performs an atomic update on the floating drive's ownership topology, replacing the ownership instance from the source instance to the target instance. Then, it incorporates the floating drive into the target instance's business service domain, allowing the target instance to issue business input / output commands to the floating drive.
[0208] This embodiment verifies the validity of the takeover record by using a period identifier, then performs the takeover operation and dual state verification, and finally completes the atomic update of the home topology and the inclusion of the business domain. This strictly ensures the legality and security of the takeover of the target instance, avoids invalid takeover and state anomalies, and ensures that there is no intermediate state in the home topology through atomic updates, eliminates the problem of concurrent access to the floating driver by two instances, and eliminates the risk of split-brain scheduling.
[0209] Based on any of the above embodiments, in Embodiment 7 of this application, performing an atomic update operation on the home topology of the floating driver includes:
[0210] Step S431: Read the current home topology record of the floating drive from the topology storage unit. The current home topology record includes at least the current home instance identifier and the current status identifier of the floating drive.
[0211] In this embodiment, the current ownership topology record is structured data stored in the topology storage unit that represents the real-time ownership relationship and operating status of the floating drive. The current ownership instance identifier is coded information that uniquely identifies the instance to which the floating drive currently belongs. The current status identifier is coded information that identifies the real-time operating status of the floating drive.
[0212] As an optional implementation, the unique identifier of the floating drive and the storage address information of the topology storage unit are obtained. During processing, the corresponding storage address is located based on the unique identifier of the floating drive, the home topology record stored at that address is read, and the current home instance identifier and current status identifier are extracted. After execution, the output or obtained result is the complete data of the current home topology record.
[0213] As an alternative implementation, real-time synchronization data of the topology storage units and floating drive identity codes are obtained. During processing, the real-time synchronization data of the topology storage units is traversed, the floating drive identity codes are matched, the corresponding home topology record is retrieved, and the core identifier field is parsed. After execution, the output or obtained result is the parsed current home instance identifier and current status identifier data.
[0214] Step S432: Initiate a comparison and exchange operation, update the current home instance identifier to the identifier of the target instance, and update the current status identifier from the migrated status to the allocated status.
[0215] In this embodiment, the compare-and-swap operation is a hardware-level instruction operation used to achieve atomic data updates, ensuring the indivisibility of data updates. The migrated-out state is the state in which the floating drive is migrating from the source instance to the target instance. The allocated state is the state in which the floating drive has completed the switchover of its home instance and can normally support services.
[0216] As an optional implementation, the current home instance identifier, target instance identifier, current status identifier, and updated status identifier are obtained. During processing, a compare and exchange instruction is executed to compare the consistency between the current home instance identifier and the source instance identifier. If the comparison passes, the home instance identifier is replaced with the target instance identifier, and the status identifier is changed from the migrated state to the allocated state. After execution, the output or obtained result is the result of the compare and exchange operation.
[0217] As another optional implementation, the original data of the belonging topology record and the configuration information of the atomic update rules are obtained. During processing, a comparison and exchange operation is initiated according to the atomic update rules to complete the synchronous replacement of the instance identifier and the status identifier, ensuring that no intermediate state is generated during the update process. After execution, the output or obtained result is feedback information indicating that the identifier synchronization update is complete.
[0218] Step S433: After the comparison and exchange operation is successful, the takeover record is cleared from the topology storage unit.
[0219] In this embodiment, clearing the takeover record is an operation that deletes the takeover credential data corresponding to this migration stored in the topology storage unit, which is used to release storage resources and avoid repeated takeovers.
[0220] As an optional implementation, a successful operation signal and the takeover record storage address are acquired, compared, and exchanged. During processing, the takeover record storage address is located based on the successful signal, and a data deletion command is issued to clear all takeover record data at that address. After execution, the output or obtained result is a takeover record clearing completion signal.
[0221] As an alternative implementation, the topology storage unit data cleanup command and the unique identifier of the takeover record are obtained. During processing, the stored data is matched using the unique identifier of the takeover record, and a targeted cleanup operation is performed, while other valid data within the topology storage unit is preserved. After execution, the output or obtained result is feedback information indicating that the targeted cleanup of the takeover record has been completed.
[0222] Step S434: When the comparison and swap operation fails, a migration failure handling process is triggered to place the floating driver in an isolated state.
[0223] In this embodiment, the migration failure handling process is a fault protection process initiated when the atomic update of the home topology fails. The isolation state is a protection state in which the floating driver is prohibited from participating in any instance service scheduling and automatic migration.
[0224] As an optional implementation, the system acquires, compares, and exchanges operation failure signals and fault protection trigger rules. During processing, a migration failure handling procedure is triggered based on the failure signal, a status control command is issued, the status identifier of the floating drive is rewritten to the isolated state, and scheduling permissions for the drive by all instances are blocked. After execution, the output or obtained result is the protection signal for the floating drive to enter the isolated state.
[0225] As an alternative implementation, the system acquires atomic update failure feedback data and instructions from the system fault handling module. During processing, the fault handling module receives failure feedback data, activates the anomaly protection mechanism, locks the scheduling interface of the floating drive, and marks it as isolated. Upon completion, the output or result is a fault signal indicating that the floating drive's scheduling authority is completely locked.
[0226] For example, after a dual-controller, dual-instance tape storage system confirms that the target instance has successfully taken over the floating drive, it reads the current ownership topology record of the floating drive from the topology storage unit, extracts the current ownership instance identifier as the source instance identifier, and sets the current status identifier to "migration out". The system initiates a compare-and-swap operation, replacing the ownership instance identifier with the target instance identifier, and simultaneously updates the status identifier from "migration out" to "assigned". If the compare-and-swap operation is successful, the system automatically clears the takeover record corresponding to this migration from the topology storage unit. If the compare-and-swap operation fails, the system immediately triggers the migration failure handling procedure, placing the floating drive in an isolated state and prohibiting all instances from performing business scheduling and automatic migration operations on it.
[0227] Optionally, refer to Figure 4 The migration process starts from the initial state, first entering the pre-check state: if the pre-check passes, it proceeds to the freeze source instance state; if the pre-check fails, it directly enters the failure state. After freezing the source instance, it proceeds to the release source instance state; if freezing the source instance fails, it enters the failure state. After releasing the source instance, it proceeds to the takeover target instance state; if taking over the target instance fails, it enters the failure state. After taking over the target instance, it proceeds to the submit topology state; if taking over the target instance fails, it enters the failure state. If the submit topology state is successful, it proceeds to the completion state, indicating a successful migration task; if submitting the topology fails, it enters the failure state. Finally, both the completion and failure states flow to the process termination node, and the migration task ends. This flow logic strictly constrains the state changes in the migration process through the control rule of "successful preceding steps before proceeding to the next stage, and failure in any stage uniformly entering the failure state," completely avoiding intermediate abnormal states during the migration process and ensuring the security and consistency of the migration operation.
[0228] Optionally, refer to Figure 5During the migration process, the current status of the floating driver changes depending on the current process. The normal migration process involves the following state transitions: Initial Business State: When the floating driver is normally carrying out business, the status is "Assigned." At this time, the driver is statically bound to the source instance and only provides business services to the source instance. Migration Start-up Phase: When the migration process is initiated (triggered by the "Migration Start" event), the status changes from "Assigned" to "Migration Out." This corresponds to the stage in the migration process where the source instance is frozen and ownership is released. The driver is in a transitional state of migrating out of the source instance, and new business access is prohibited. Source Instance Release Phase: When the source instance completes the release of ownership of the floating driver (triggered by the "Release Complete" event), the status changes from "Migration Out" to "Migration In." This corresponds to the stage in the migration process where the target instance takes over. The driver is in a transitional state of migrating into the target instance, waiting for the target instance to complete the takeover verification. Target takeover phase: The target instance completes the takeover operation of the floating driver (triggers the "takeover complete" event), and the status flag changes from "migration in" back to "assigned". This corresponds to the stage in the migration process where the topology atomic update is completed. The driver is rebound to the target instance, normal business services are restored, and the status transition of this migration process is completed.
[0229] In case of anomalies / failures, the status indicator will change according to the following rules if a failure or anomaly occurs at any stage of the migration process: Failure Trigger: If a "failure" event is triggered in any of the "Assigned," "Migration Out," or "Migration In" states, the status indicator will directly change to "Failure," the driver will enter the failure state, and all business services will be suspended. Anomaly Trigger: If an "anomaly" event (such as release failure, takeover failure, or topology update failure) is triggered in the "Migration Out" or "Migration In" states, the status indicator will directly change to "Isolation," the driver will enter the isolation state, all automatic migration and business scheduling will be prohibited, and the failure will be prevented from spreading. Failure Recovery: If the failure is repaired in the failure state (triggered by the "Repair" event), the status indicator will change from "Failure" to "Recovery Observation," and the driver will enter the recovery observation phase; if the verification is completed in the recovery observation phase (triggered by the "Verification Passed" event), the status indicator will change back from "Recovery Observation" to "Assigned," and normal business services will be restored; if an anomaly occurs in the recovery observation phase, the status indicator will change to "Isolation." Restoration from isolation: The isolation state is the final protection state. The status identifier can only be changed back to "assigned" through the "manual recovery" event or back to "unassigned" through the "reconciliation recovery" event. Only after the abnormal state is manually repaired can the driver rejoin the system scheduling.
[0230] Supplementary status descriptions: Unassigned state: This is the initial / idle state of the floating drive. It can transition to the "Assigned" state via the "Assign" event, completing the initial business binding. The isolated state can also transition back to the "Unassigned" state via "Reconciliation Recovery," awaiting reallocation. Core control logic: Through strict state transition constraints, it ensures that the floating drive has only one valid state during the migration process, completely avoiding state inconsistencies caused by concurrent access from two instances, and guaranteeing the security and stability of the migration process.
[0231] This embodiment achieves atomic updates of the ownership topology through comparison and swap operations, ensuring the integrity and uniqueness of data updates. At the same time, it sets branch logic to clear the takeover record on success and trigger isolation protection on failure. This not only completes the effective switching of the ownership relationship of the floating driver, but also prevents the spread of faults when update is abnormal. It completely avoids the risks of state inconsistency and split-brain scheduling caused by concurrent access of dual instances, and improves the stability and security of system migration.
[0232] Based on any of the above embodiments, in Embodiment 8 of this application, after controlling the target instance to take over the floating drive based on the takeover record with the period identifier, the method further includes:
[0233] Step B10: When a switch of the scheduling control module is detected, the switched scheduling control module determines whether there are any unfinished migration task records in the topology storage unit.
[0234] In this embodiment, the scheduling control module switching is a runtime event in a dual-controller, dual-instance tape storage system where the scheduling master node changes. Incomplete migration task records are floating drive migration task data stored in the topology storage unit that have not yet reached the final completion state.
[0235] As an optional implementation, the switching signal of the scheduling control module and the migration task log data of the topology storage unit are acquired. During processing, the switched scheduling control module traverses the migration task logs within the topology storage unit, filtering for task entries with a status field indicating incompleteness. After execution, the output or obtained result is either a determination signal indicating the existence of incomplete migration task records or a determination signal indicating the absence of incomplete migration task records.
[0236] As an alternative implementation, the system scheduling state switching identifier and migration task storage block data are obtained. During processing, the dedicated storage block for migration tasks is locked based on the scheduling state switching identifier, and task status data is read in batches and validity is determined. After execution, the output or obtained result is the existence verification result of incomplete migration task records.
[0237] Step B11: If there are incomplete migration task records, query the current status of the source instance and the current status of the target instance corresponding to the migration task record.
[0238] In this embodiment, the current state of the source instance is the real-time running and permission-holding state of the original owner instance corresponding to the migration task. The current state of the target instance is the real-time running and takeover execution state of the target owner instance corresponding to the migration task.
[0239] As an optional implementation, the system obtains records of incomplete migration tasks and instance status query commands. During processing, the source instance identifier and target instance identifier are extracted from the records of incomplete migration tasks, a status query request is sent to the corresponding instance, and real-time status data is collected. After execution, the output or obtained result is the real-time status data of the source instance and the target instance.
[0240] As an alternative implementation, the topology storage unit instance status table and migration task-related instance information are obtained. During processing, the instance status table is traversed based on the instance information associated with the migration task, and the latest status field data of the corresponding instance is read. After execution, the output or obtained result is the status query result of the source instance and the target instance.
[0241] Step B12: Perform the corresponding recovery operation based on the current state of the source instance and the current state of the target instance.
[0242] In this embodiment, the recovery operation is a state repair, process rollback, or fault isolation operation performed on an incomplete migration task.
[0243] As an optional implementation, source instance status data, target instance status data, and recovery operation matching rules are obtained. During processing, the instance status data is compared item by item with the recovery operation matching rules to identify the matching recovery operation type. After execution, the output or obtained result is the start command for the corresponding recovery operation.
[0244] As an alternative implementation, the instance status determination result and system fault recovery configuration parameters are obtained. During processing, the corresponding configuration parameters are invoked based on the instance status determination result to generate a targeted recovery execution command. After execution, the output or obtained result is the execution signal for the targeted recovery operation.
[0245] Step B121: If the source instance holds the floating drive and the target instance has not yet taken over, a rollback operation is performed to restore the floating drive to the source instance.
[0246] In this embodiment, the rollback operation is the operation of restoring the ownership and running state of the floating driver to the initial state before the migration start.
[0247] As an optional implementation, the system obtains the source instance's permission data and the target instance's non-takeover status data. During processing, a rollback command is issued to unfreeze the migration and restore the source instance's service scheduling permissions to the floating drive. Upon completion, the output or result is a rollback completion signal indicating that the floating drive has been restored to the source instance.
[0248] As an alternative implementation, the initial state data of the migration task and the instance permission restoration rules are obtained. During the process, the floating drive ownership relationship is reset according to the initial state data, and the service control permissions of the source instance are restored according to the permission restoration rules. After execution, the output or result is a floating drive ownership relationship restoration completion signal.
[0249] Step B122: If the source instance has released the floating driver but the target instance has not yet taken over, then the floating driver is placed in an isolated state and new scheduling is blocked.
[0250] In this embodiment, blocking new scheduling is a control behavior that prevents the system from initiating any new automatic migration or service allocation operations on the floating driver.
[0251] As an optional implementation, the source instance release completion signal and the target instance non-takeover signal are obtained. During processing, the floating driver status field is rewritten to isolated status, and the automatic scheduling interface of the floating driver is disabled. After execution, the output or result is a floating driver isolation effective and a new scheduling blocking completion signal.
[0252] As an alternative implementation, instance permission status data and system isolation control rules are obtained. During processing, an isolation mechanism is triggered based on the permission status data, locking all scheduling channels of the floating driver and marking them with a no-scheduling flag. After execution, the output or result is a protection signal indicating that the floating driver scheduling channels are completely locked.
[0253] Step B123: If the source instance has been released and the target instance has been taken over but the topology has not yet been submitted, then submit the topology update or perform reconciliation repair.
[0254] In this embodiment, submitting a supplementary topology update is an operation to supplement and submit an incomplete ownership topology change. Reconciliation and repair is an operation to verify and correct the topology storage unit data against the actual state of the instance.
[0255] As an optional implementation, the source instance release data, target instance takeover data, and topology uncommitted status data are obtained. During processing, the topology atomic update interface is called to perform a supplementary commit, completing the replacement and update of the owner instance identifier. After execution, the output or obtained result is a topology supplementary commit completion signal.
[0256] As an alternative implementation, topology storage data, instance actual status data, and reconciliation and repair rules are obtained. During processing, the topology data and instance status data are checked item by item, and inconsistent topology entries are corrected to complete the repair. After execution, the output or obtained result is a topology data reconciliation and repair completion signal.
[0257] Step B124: If the source instance has been released, the target instance has been taken over, and the topology has been committed, then the completion status is written and scheduling is restored.
[0258] In this embodiment, "overwriting the completion status" is the operation of updating the status field of the migration task to the completion status. "Restoring scheduling" is the operation of removing control and allowing the floating driver to participate in normal service scheduling.
[0259] As an optional implementation, the migration task status data and topology submission completion signal are obtained. During processing, the migration task status field is rewritten to a completed state, and normal service scheduling permissions for the floating driver are granted. After execution, the output or result is a migration task status update completed and a scheduling recovery signal.
[0260] As an alternative implementation, task completion determination data and system scheduling recovery rules are acquired. During processing, the task status is updated based on the determination data, and the normal scheduling process of the floating drive is restarted according to the scheduling recovery rules. After execution, the output or obtained result is a running signal indicating that the floating drive has resumed normal business scheduling.
[0261] For example, in a dual-controller, dual-instance tape storage system, if the scheduling control module switches, the new scheduling control module immediately scans the topology storage units and detects an incomplete floating drive migration task record. The scheduling control module extracts the corresponding source and target instance information from this task record, finding that the source instance has completed the floating drive release operation, and the target instance has completed the takeover operation but the ownership topology has not yet been submitted. Based on this status combination, the system performs a reconciliation and repair operation, completing the supplementary submission update of the ownership topology, and then rewrites the migration task completion status, restoring normal business scheduling of the floating drive. If the system finds that the source instance still holds the floating drive and the target instance has not taken over, it directly performs a rollback operation, restoring the floating drive to the source instance and reverting to its initial business state. If the system finds that the source instance has released the floating drive and the target instance has not taken over, it places the floating drive in an isolated state, blocking all new scheduling requests and awaiting manual maintenance.
[0262] Optionally, refer to Figure 6This implementation corresponds to the recovery process of incomplete migration tasks after a master switch in a dual-controller, dual-instance tape storage system. Upon detecting a master switch, the system first loads the incomplete migration tasks, then determines if any migration tasks exist. If no migration tasks exist, normal scheduling is resumed directly. If migration tasks exist, the system queries the status of the source and target instances corresponding to the migration task. It first determines if the source instance holds the drive. If the source instance holds the drive, the system performs a rollback operation, then blocks new scheduling, and finally resumes normal scheduling. If the source instance does not hold the drive, the system further determines if the target instance has taken over. If the target instance has not taken over, the system sets the floating drive to an isolated state, then blocks new scheduling, and finally resumes normal scheduling. If the target instance has taken over, the system then determines if the topology has been committed. If the topology has been committed, the system writes the completion status, and then resumes normal scheduling. If the topology has not been committed, the system performs a supplementary commit or reconciliation repair operation, and then resumes normal scheduling.
[0263] This embodiment fully covers the migration task repair logic under the dual-controller system master switch scenario by detecting incomplete migration tasks after the scheduling control module switches, querying instance status, and performing multi-branch recovery operations. It performs precise recovery processing for different migration interruption states, which not only ensures the integrity and traceability of the migration process, but also enables isolation protection in abnormal states, avoiding state chaos and split-brain scheduling risks caused by master switch, and improving the fault recovery capability and operational stability of the dual-controller dual-instance tape storage system.
[0264] This application provides a dual-controller, dual-instance tape storage system, including multiple fixed drives, statically bound to a first instance or a second instance respectively; a floating drive, serving as the only drive in the system allowed for automatic cross-instance migration; a scheduling control module, used to uniformly acquire status, determine the target instance, and initiate migration actions; a migration orchestration module, used to execute actions such as freezing the source instance, releasing ownership, generating takeover records, target instance takeover, and topology atomic updates; a master control arbitration module, used to determine the current scheduling master through a master control lease mechanism; a state management module, used to maintain the drive state machine and migration task state machine; and a topology storage unit, used to persistently store release confirmation information, takeover records, owned topology, and period identifiers.
[0265] Optionally, refer to Figure 7 The processor schedules and processes the topology storage unit, scheduling control module, migration orchestration module, state management module, master control arbitration module, and isolation processing module. The scheduling control module is communicatively connected to the first instance, the second instance, the migration orchestration module, the master control arbitration module, the state management module, and the topology storage unit. The migration orchestration module is communicatively connected to the first instance and the second instance. The state management module is communicatively connected to the scheduling control plane and the migration orchestration module, and is used to record driver status, migration task status, and topology information.
[0266] Optionally, this embodiment discloses the overall architecture and full-process control method of the dual-controller dual-instance tape storage system. The system hardware and logic architecture, module interaction relationship and overall technical solution implementation steps are as follows.
[0267] The dual-controller, dual-instance tape storage system of this embodiment consists of a first instance, a second instance, multiple fixed drives, a floating drive, a scheduling control module, a migration orchestration module, a master control arbitration module, a status management module, and a topology storage unit. The first instance and the second instance are two independently operating tape storage service instances within the same dual-controller system, used to handle tape storage business processing and service scheduling. Multiple fixed drives establish static binding relationships with either the first or second instance, providing business services only to their corresponding bound instances and not participating in automatic cross-instance migration. The floating drive is the only tape drive in the system with automatic cross-instance migration permissions, and can switch ownership between the first and second instances under scheduling control. The scheduling control module establishes communication connections with the first instance, the second instance, the migration orchestration module, the master control arbitration module, the status management module, and the topology storage unit, respectively, to uniformly collect the system's global status, determine the target ownership instance of the floating drive, and initiate migration actions. The migration orchestration module establishes communication connections with both the first and second instances, to execute the entire process of freezing the source instance, releasing ownership, generating takeover records, taking over the target instance, and updating the topology atomically. The master control arbitration module establishes a communication connection with the scheduling control module and determines the currently effective scheduling master through the master control lease mechanism. The state management module establishes communication connections with both the scheduling control module and the migration orchestration module to maintain the driver state machine and migration task state machine, recording driver operating status, migration task execution status, and topology association information. The topology storage unit establishes a communication connection with the scheduling control module to persistently store release confirmation information, takeover records, ownership topology data, and period identifier data.
[0268] After system startup, it enters the initialization running state. The master control arbitration module completes the election of the scheduling master through the master control lease mechanism, determining the currently effective scheduling control module. The state management module loads the preset driver state machine and migration task state machine rules, completing the state machine initialization configuration. The topology storage unit loads the initial system topology data, recording the static binding relationship of fixed drivers and the initial home instance of floating drivers. The scheduling control module obtains the system global state from the topology storage unit through the communication link. The global state includes the first health state of the first instance, the second health state of the second instance, the home and third health states of each fixed driver, and the current home instance of the floating driver. The scheduling control module performs a migration pre-check, verifying the automatic scheduling switch status of the system, the compliance of floating driver security constraints, and the existence of incomplete migration tasks. After the pre-check passes, the target home instance determination process is initiated. The scheduling control module marks the current home instance of the floating driver as the source instance, and calculates the target instance by combining fault compensation strategy, time-sharing scheduling strategy, or load balancing strategy. If the calculated target instance is the same as the source instance, the system immediately terminates the current migration process. If the target instance and the source instance are different instances, the scheduling control module issues a migration execution instruction to the migration orchestration module.
[0269] After receiving the instruction, the migration orchestration module controls the source instance to cease all business input / output interactions with the floating drive and continuously monitors the execution status of the floating drive's in-transit critical tasks. Once the in-transit critical tasks have converged, the migration orchestration module controls the source instance to sequentially execute the floating drive's cache flushing, metadata persistence, and instance ownership release operations. The migration orchestration module generates a release confirmation record based on the drive identifier, source instance identifier, release timestamp, and current master control term number, and writes the release confirmation record to the target area of the topology storage unit for persistent storage. The scheduling control module generates a takeover record based on the current valid period identifier and stores the takeover record in the topology storage unit. The target instance retrieves the takeover record from the topology storage unit, performs a validity check based on the period identifier, and executes the floating drive takeover operation after successful verification, sequentially completing the floating drive's hardware status check and metadata consistency check.
[0270] After confirming successful takeover of the target instance, the scheduling control module reads the current ownership topology record of the floating drive from the topology storage unit. This record contains the current ownership instance identifier and current status identifier of the floating drive. The scheduling control module initiates a compare-and-swap operation, updating the current ownership instance identifier to the target instance's identifier and changing the current status identifier from migrated out to allocated. If the compare-and-swap operation succeeds, the scheduling control module clears the takeover record corresponding to this migration from the topology storage unit. If the compare-and-swap operation fails, the system immediately triggers the migration failure handling procedure, placing the floating drive in an isolated state and blocking automatic scheduling and service access to that drive from all instances.
[0271] The system monitors the operational status of the scheduling control module in real time. When a switch of scheduling control modules is detected, the new scheduling control module scans the topology storage units to determine if there are any incomplete migration task records. If an incomplete migration task record exists, the scheduling control module queries the current status of the source instance and the target instance corresponding to the migration task. If the source instance holds the floating drive and the target instance has not yet taken over, the system performs a rollback operation to restore the floating drive to the source instance. If the source instance has released the floating drive but the target instance has not yet taken over, the system puts the floating drive into an isolated state and blocks all new scheduling requests. If the source instance has released the floating drive, the target instance has taken over, but the home topology has not yet been submitted, the system performs a supplementary topology update or reconciliation repair operation. If the source instance has released the floating drive, the target instance has taken over, and the home topology has been submitted, the system writes the migration task completion status and restores normal business scheduling of the floating drive.
[0272] For example, after the dual-controller, dual-instance tape storage system completes its initial configuration, the master control arbitration module determines the primary scheduling control module through a lease mechanism, the state management module loads the drive and migration task state machines, and the topology storage unit stores the initial ownership topology data. The scheduling control module collects global status data and detects that the number of faults in the fixed drives bound to the first instance has reached a preset fault threshold. It then performs a migration pre-check to confirm that the automatic scheduling switch is on, the floating drives are in a normal and available state, and there are no incomplete migration tasks. The scheduling control module identifies the first instance as the source instance, calculates the second instance as the target instance based on the fault compensation strategy, and issues a migration command to the migration orchestration module. The migration orchestration module controls the first instance to stop its business interaction with the floating drives. After the in-transit critical tasks converge, it completes cache flushing, metadata persistence, and ownership release, generates a release confirmation record, and writes it to the topology storage unit. The scheduling control module generates a takeover record with a valid period identifier. After the second instance verifies the period identifier, it completes the takeover of the floating drives, hardware status verification, and metadata consistency verification. The scheduling control module initiates a comparison and exchange operation to complete the atomic update of the floating drive's ownership topology, clears the takeover record, and completes the entire migration process. If a scheduling control module switch occurs during the migration process, the new scheduling control module will detect that the migration task has not been completed, query and confirm that the source instance has been released, the target instance has been taken over, and the topology has not been submitted. Then, it will perform a topology update operation to update the migration task completion status and restore the system's normal business scheduling.
[0273] This embodiment establishes a complete system architecture with dual instances, fixed drives, and a single floating drive, along with scheduling control, migration orchestration, master controller arbitration, state management, and topology storage. It relies on standardized migration pre-checks, target instance determination, source release, target takeover, topology atomic updates, and master controller switching recovery processes. This strictly constrains the access timing and state change logic of floating drives during cross-instance migration, completely avoiding state inconsistencies caused by concurrent access to the same drive by the source and target instances. Period flags and master controller lease mechanisms eliminate the risk of split-brain scheduling in a dual-controller environment. Simultaneously, it adapts to various application scenarios such as fault compensation, time-sharing, and load balancing, improving the resource utilization, operational stability, and fault recovery capabilities of the tape storage system.
[0274] This application provides a single floating drive secure migration control device in a dual-controller dual-instance tape storage system. The single floating drive secure migration control device in the dual-controller dual-instance tape storage system includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the single floating drive secure migration control method in the dual-controller dual-instance tape storage system described in Embodiment 1 above.
[0275] The following is for reference. Figure 8 This document illustrates a structural schematic diagram of a single floating drive secure migration control device suitable for implementing a dual-controller dual-instance tape storage system according to embodiments of this application. The single floating drive secure migration control device in the dual-controller dual-instance tape storage system of this application can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablets, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The single floating drive secure migration control device in the dual-controller dual-instance tape storage system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0276] like Figure 8 As shown, the single floating drive secure migration control device in a dual-controller, dual-instance tape storage system may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the single floating drive secure migration control device in the dual-controller, dual-instance tape storage system. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the single floating drive secure migration control device in a dual-controller dual-instance tape storage system to wirelessly or wiredly communicate with other devices to exchange data. Although the figure shows a single floating drive secure migration control device in a dual-controller dual-instance tape storage system with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0277] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0278] The single-floating-drive secure migration control device in a dual-controller, dual-instance tape storage system provided in this application employs the single-floating-drive secure migration control method in the dual-controller, dual-instance tape storage system described in the above embodiments. This addresses the technical problem that concurrent access to the same drive by the source instance and target instance can lead to inconsistent states and consequently, the risk of split-brain scheduling. Compared to the prior art, the beneficial effects of the single-floating-drive secure migration control device in a dual-controller, dual-instance tape storage system provided in this application are the same as those of the single-floating-drive secure migration control device in the dual-controller, dual-instance tape storage system described in the above embodiments. Furthermore, other technical features of this single-floating-drive secure migration control device in a dual-controller, dual-instance tape storage system are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0279] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0280] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0281] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the single floating drive secure migration control method in the dual-controller dual-instance magnetic tape storage system described in the above embodiments.
[0282] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0283] The aforementioned computer-readable storage medium may be included in a single floating drive secure migration control device in a dual-controller dual-instance tape storage system; or it may exist independently and not be assembled into a single floating drive secure migration control device in a dual-controller dual-instance tape storage system.
[0284] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a single floating drive secure migration control device in a dual-controller dual-instance tape storage system, the single floating drive secure migration control device in the dual-controller dual-instance tape storage system: obtains the global state of the dual-controller dual-instance tape storage system from the topology storage unit, wherein the global state includes the first health state of the first instance, the second health state of the second instance, the ownership status and third health state of each of the fixed drives, and the current ownership instance of the floating drive; divides the first instance and the second instance into a source instance and a target instance according to the global state and a preset strategy; controls the source instance to release the floating drive and generates a takeover record with a time identifier; and controls the target instance to take over the floating drive based on the takeover record with the time identifier.
[0285] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0286] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation that may be implemented in systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0287] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0288] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the single floating drive secure migration control method in the aforementioned dual-controller dual-instance tape storage system. This addresses the technical problem that the source instance and target instance may concurrently access the same drive, leading to inconsistent states and thus the risk of split-brain scheduling. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the single floating drive secure migration control method in the dual-controller dual-instance tape storage system provided in the above embodiments, and will not be elaborated upon here.
[0289] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the single floating drive secure migration control method in a dual-controller dual-instance tape storage system as described above.
[0290] The computer program product provided in this application can solve the technical problem that source instances and target instances may concurrently access the same drive, leading to inconsistent states and thus the risk of split-brain scheduling. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the single floating drive secure migration control method in the dual-controller dual-instance tape storage system provided in the above embodiments, and will not be repeated here.
[0291] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for secure migration control of a single floating drive in a dual-controller, dual-instance magnetic tape storage system, characterized in that, The dual-controller dual-instance tape storage system includes a scheduling master, a first instance, a second instance, multiple fixed drives, and a single floating drive. The secure migration control method for the single floating drive in the dual-controller dual-instance tape storage system includes: The global state of the dual-controller dual-instance tape storage system is obtained from the topology storage unit, wherein the global state includes the first health state of the first instance, the second health state of the second instance, the ownership status and third health state of each fixed drive, and the current ownership instance of the floating drive. The current owner instance of the floating driver is determined as the source instance; Based on the first health status, the second health status, the ownership status of each fixed driver, and the third health status, a target instance is calculated in conjunction with a preset strategy, wherein the preset strategy includes at least one of a fault compensation strategy, a load balancing strategy, and a time-sharing scheduling strategy. If the calculated target instance is the same as the source instance, the migration process is terminated. If the floating drive does not belong to any instance, skip the drive removal process and add the drive to the target instance. If the calculated target instance is different from the source instance, then control the source instance to stop business input / output interactions with the floating driver; If the critical tasks in transit of the floating drive are completed, control the source instance to perform cache flushing, metadata persistence, and instance ownership release of the floating drive; Generate a release confirmation record based on the floating driver identifier, source instance identifier, release timestamp, and current scheduling master term number; The release confirmation record is written to the target area of the topology storage unit to persist the release confirmation record and generate a takeover record with a period identifier. The period identifier corresponds one-to-one with the current scheduling master term number generated by the master arbitration module. Only one unique period identifier is generated for each scheduling master term. When the scheduling master switches, the period identifier automatically expires and all takeover records based on the old period identifier are rejected. Verify the takeover record based on the aforementioned period identifier; If the verification passes, the target instance is controlled to perform the takeover operation of the floating driver based on the takeover record, and the hardware status verification and metadata consistency verification of the floating driver are completed in sequence. After confirming that the target instance has successfully taken over, an atomic update operation is performed on the topology of the floating drive, and the floating drive is included in the business service domain of the target instance.
2. The single floating drive secure migration control method in a dual-controller, dual-instance magnetic tape storage system as described in claim 1, characterized in that, After obtaining the global state of the dual-controller dual-instance tape storage system from the topology storage unit, the process includes: Perform a migration pre-check. If the migration pre-check passes, proceed with the steps of determining the source instance and calculating the target instance. The migration pre-check includes one of the following: Confirm that the system's automatic scheduling switch is enabled; Ensure compliance with security constraints, wherein the security constraints include at least one of the following: the floating driver is in a non-isolated state, a non-faulty state, or a non-migrating state; It has been confirmed that there are no incomplete migration tasks.
3. The single floating drive secure migration control method in a dual-controller, dual-instance magnetic tape storage system as described in claim 1, characterized in that, The step of calculating the target instance based on the first health status, the second health status, the ownership status of each fixed driver, and the third health status, combined with the preset strategy, includes any one of the following: When the preset strategy is a fault compensation strategy, if the number of fixed drivers in the source instance that are in a fault state reaches a preset fault threshold, and the number of fixed driver faults in the other instance between the first instance and the second instance is less than the number of fixed driver faults in the source instance, then the other instance is determined as the target instance. When the preset strategy is a time-sharing scheduling strategy, the current time slice number is obtained, and the instance corresponding to the current time slice number is determined as the target instance according to the pre-configured mapping relationship between time slices and instances. When the preset strategy is a load balancing strategy, the unprocessed task queue depths of the first instance and the second instance are obtained, and the instance with the smaller queue depth is determined as the target instance.
4. The single floating drive secure migration control method in a dual-controller, dual-instance magnetic tape storage system as described in claim 1, characterized in that, The atomic update operation on the home topology of the floating driver includes: The current home topology record of the floating drive is read from the topology storage unit, and the current home topology record includes at least the current home instance identifier and the current status identifier of the floating drive; Initiate a comparison and exchange operation to update the current home instance identifier to the identifier of the target instance, and update the current status identifier from the migrated status to the allocated status; After the comparison and swap operation is successful, the takeover record is cleared from the topology storage unit; If the comparison and swap operation fails, a migration failure handling process is triggered, and the floating driver is placed in an isolated state.
5. The single floating drive secure migration control method in a dual-controller, dual-instance magnetic tape storage system as described in claim 1, characterized in that, If the verification passes, after controlling the target instance to perform the takeover operation of the floating driver based on the takeover record, the method further includes: When a switch of scheduling master is detected, the new scheduling master determines whether there are any unfinished migration task records in the topology storage unit. If there are incomplete migration task records, query the current status of the source instance and the current status of the target instance corresponding to the migration task record; Perform the corresponding recovery operation based on the current state of the source instance and the current state of the target instance; The step of performing the corresponding recovery operation based on the current state of the source instance and the current state of the target instance includes at least one of the following: If the source instance holds the floating drive and the target instance has not yet taken over, a rollback operation is performed to restore the floating drive to the source instance; If the source instance has released the floating driver but the target instance has not yet taken over, then the floating driver is placed in an isolated state and new scheduling is blocked; If the source instance has been released and the target instance has been taken over but the topology has not yet been committed, then submit a topology update or perform a reconciliation repair. If the source instance has been released, the target instance has been taken over, and the topology has been committed, then the completion status is written and scheduling is resumed.
6. A single floating drive secure migration control device in a dual-controller, dual-instance magnetic tape storage system, characterized in that, The single floating drive secure migration control device in the dual-controller dual-instance tape storage system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the single floating drive secure migration control method in the dual-controller dual-instance tape storage system as described in any one of claims 1 to 5.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the single floating drive secure migration control method in a dual-controller dual-instance magnetic tape storage system as described in any one of claims 1 to 5.