Online expansion and contraction method, device, equipment and medium of database parallel file system

By redistribution of database files and automatic migration of data units during the scaling process of shared cluster databases, the problem of traditional methods requiring downtime operation is solved, load balancing and efficient data migration are realized, and the normal operation of the database is ensured.

CN118820177BActive Publication Date: 2025-05-06SHENZHEN INST OF COMPUTING SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410830078.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-05-06
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

When the existing technology expands the shared storage of a shared cluster database, it is necessary to shut down operations to adjust the I/O hotspot, resulting in complex operations and high risks, and affecting the normal read and write operations of the database.

Method used

By redistribution of database files when disk expansion and capacity of shared cluster databases, the target disk location of each data unit is obtained, and locked marks are used using data mapping tables, allowing read operations and prohibiting write operations, realizing automatic migration of data units.

Benefits of technology

Automatic load balancing of all disks after scaling is realized, avoiding frequent disk access, improving data migration efficiency, enhancing utilization and access efficiency of all disks after scaling is achieved, and ensuring data integrity and consistency of database read and write operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118820177B_ABST
    Figure CN118820177B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of database technology, and relates to a method, device, equipment and medium for online expansion and contraction of a cluster database parallel file system. When the shared cluster database parallel file system is expanded or contracted on disk, the method dynamically migrates the database files online, calculates the migrated data units and the target disk positions according to the data mapping table, creates a migration mapping table, and before the migration, locks the mapping relationship of the data units recorded in the data mapping table. During the migration, the data units are migrated according to the records in the migration mapping table. After the migration is completed, the mapping relationship of the data units recorded in the marked data mapping table is updated according to the target disk position, and the lock mark is removed. The method realizes online dynamic expansion and contraction of the shared cluster database parallel file system, automatic load balancing of all disk reads and writes, concurrency control during database reads and writes and online expansion and contraction, and automatic recovery of server abnormalities during dynamic expansion and contraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application is applicable to the field of database technology, and in particular relates to online scaling methods, devices, equipment, and media for a cluster database parallel file system. Background Art

[0002] A shared cluster database is a cluster system that allows multiple database instances to access and operate the same database simultaneously. In a shared cluster database, database files such as data files and control files are stored in shared storage. Multiple database instances can simultaneously access and modify data in the shared storage of the same database. As data volume increases and access requirements arise, the shared storage needs to be scaled up or down. However, after scaling up or down the shared storage of a shared cluster database, the shared storage may experience hotspots in I / O. This not only reduces storage resource utilization but can also significantly decrease access efficiency. Traditional solutions often require downtime, meaning that the database service must be manually stopped to adjust I / O hotspots. This approach is not only complex and risky, but also severely impacts normal read and write operations of the database. Therefore, improving the manageability and maintainability of online scaling of shared cluster databases has become a pressing issue. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a method, apparatus, device, and medium for online scaling of a database parallel file system to solve the problem of how to improve the manageability and maintainability of online scaling of a shared cluster database.

[0004] In a first aspect, an embodiment of the present application provides an online scaling method for a database parallel file system, the online scaling method comprising:

[0005] When performing disk expansion and contraction on a shared cluster database, redistributing database files in the shared cluster database to obtain target disk locations of all data units in each database file;

[0006] For any data unit, obtain a data mapping table of the database file corresponding to the data unit, wherein the data mapping table is used to record the mapping relationship between all data units belonging to the database file and the corresponding disk locations;

[0007] Lock-marking the mapping relationship of the data unit recorded in the data mapping table to obtain a marked data mapping table, wherein the lock mark in the marked data mapping table is used to allow a read operation and prohibit a write operation when migrating data in the data unit;

[0008] The data in the data unit is migrated according to the target disk position of the data unit. After the migration is completed, the mapping relationship of the data unit recorded in the marked data mapping table is updated according to the target disk position, and the lock mark is removed to obtain an updated data mapping table.

[0009] In a second aspect, an embodiment of the present application provides an online scaling device for a shared cluster database, the online scaling device comprising:

[0010] A redistribution calculation module is used to redistribute database files in the shared cluster database when the shared cluster database is expanded or reduced in disk capacity, and obtain the target disk location of all data units in each database file;

[0011] A mapping table acquisition module is used to acquire, for any data unit, a data mapping table of a database file corresponding to the data unit, wherein the data mapping table is used to record the mapping relationship between all data units belonging to the database file and corresponding disk locations;

[0012] a lock marking module, configured to lock-mark the mapping relationship of the data unit recorded in the data mapping table to obtain a marked data mapping table, wherein the lock mark in the marked data mapping table is used to allow a read operation and prohibit a write operation when migrating data in the data unit;

[0013] A migration module is used to migrate the data in the data unit according to the target disk position of the data unit. After the migration is completed, the mapping relationship of the data unit recorded in the marked data mapping table is updated according to the target disk position, and the lock mark is removed to obtain an updated data mapping table.

[0014] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the online scaling method for the database parallel file system as described in the first aspect is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the online scaling method of the database parallel file system as described in the first aspect is implemented.

[0016] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the present application obtains the target disk position of the data unit in each database file by redistributing the database files in the shared cluster database when the disk of the shared cluster database is expanded or reduced, obtains the data mapping table of the database file corresponding to the data unit for any data unit, locks the mapping relationship of the data unit recorded in the data mapping table, and obtains the marked data mapping table. The lock mark in the marked data mapping table is used to allow read operations and prohibit write operations when migrating data in the data unit. The data in the data unit is migrated according to the target disk position of the data unit. After the migration is completed, the mapping relationship of the data unit recorded in the marked data mapping table is updated according to the target disk position, and the lock mark is removed to obtain an updated data mapping table. Among them, when the shared cluster database performs disk expansion or contraction, the data unit is automatically migrated according to the calculation result of the target disk position of the data unit in the database file, and the data unit is evenly distributed to all the disks after expansion or contraction, thereby realizing automatic load balancing of all disks after expansion or contraction, avoiding the situation where the expanded or contracted disks are frequently accessed, thereby improving the efficiency of data migration, and improving the utilization rate and access efficiency of all disks after expansion or contraction, and before the data unit is migrated, the mapping relationship corresponding to the data unit is locked and marked, so that during the data unit migration process, the data unit is allowed to be read and the data unit is prohibited from being written, realizing the automatic migration of the data unit and the concurrent control of the database's read and write operations on the migrated data unit, ensuring the integrity and consistency of the data read and written by the database during the migration process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 This is a schematic diagram of an application environment of an online scaling method for a database parallel file system provided in Example 1 of the present application;

[0019] Figure 2 This is a flow chart of an online scaling method for a database parallel file system provided in Example 2 of the present application;

[0020] Figure 3 The second embodiment of the present application provides a schematic diagram of the mapping relationship between a database file, a data mapping table, and a shared storage disk group;

[0021] Figure 4 This is a flow chart of an online scaling method for a database parallel file system provided in Example 3 of the present application;

[0022] Figure 5 This is a flow chart of an online scaling method for a database parallel file system provided in Example 4 of the present application;

[0023] Figure 6 This is a flowchart of an online scaling method for a database parallel file system provided in Example 5 of the present application;

[0024] Figure 7 This is a flowchart of an online scaling method for a database parallel file system provided in Example 6 of the present application;

[0025] Figure 8 This is a flow chart of an online scaling method for a database parallel file system provided in Example 7 of the present application;

[0026] Figure 9 This is a flow chart of an online scaling method for a database parallel file system provided in Example 8 of the present application;

[0027] Figure 10 This is a schematic diagram of the structure of an online scaling device for a shared cluster database provided in Example 9 of the present application;

[0028] Figure 11 This is a structural diagram of a computer device provided in Example 10 of the present application. DETAILED DESCRIPTION

[0029] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0030] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0031] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0032] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0033] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0034] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0035] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0036] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0037] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0038] In order to illustrate the technical solution of the present application, specific embodiments are provided below.

[0039] The first embodiment of the present application provides an online scaling method for a database parallel file system, which can be applied in the following situations: Figure 1 In an application environment, a server communicates with a shared cluster database, providing online scaling services for the shared cluster database. The shared cluster database triggers online scaling tasks on the server. A shared cluster database can be a cluster consisting of multiple database nodes, each of which can be implemented as a physical server, virtual machine, or container. The computer device corresponding to the server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0040] See also Figure 2 , is a flow chart of an online scaling method for a database parallel file system provided in the second embodiment of the present application. The online scaling method for a database parallel file system is applied to Figure 1 The server in the server communicates with the shared cluster database to obtain the database files and data mapping tables sent by the shared cluster database. Figure 2 As shown, the online scaling method may include the following steps:

[0041] Step S201 : when performing disk expansion or contraction on a shared cluster database, database files in the shared cluster database are redistributed to obtain target disk locations of all data units in each database file.

[0042] In an embodiment of the present application, a shared cluster database may refer to a database cluster system that allows multiple database instances to access and operate the same shared storage at the same time. In a shared cluster database using a shared disk architecture, database files such as data files and control files are stored on the disk, and multiple database instances can access and modify the data on the disk simultaneously and in parallel. As the amount of data increases and access needs arise, it is necessary to perform disk expansion and contraction operations on the shared cluster database. The expansion and contraction of the disk may include: 1) adding or deleting disk capacity in any database node in the shared cluster database; 2) increasing or decreasing the number of disks in the shared cluster database due to an increase or decrease in the number of database nodes in the shared cluster database.

[0043] A data unit may refer to a data storage unit of a database file on disk. The data in each database file is stored on disk in units of data units. For example, the size of the data unit may be 4MB. The target disk location may refer to a disk storage location reallocated to the data unit in the database file when the shared cluster database performs disk expansion or contraction. When a data unit is migrated, the target disk location is the disk location to which the data unit needs to be migrated.

[0044] When scaling disks in a shared cluster database, in order to achieve load balancing across all disks after scaling, the data in the data units in the database files can be migrated. Before migrating the data units, a redistribution calculation can be performed on the storage location of the data units on the disk to determine the target disk location to which the data units need to be migrated. Specifically, before migration, a redistribution calculation can be performed on the disk storage location of the data units based on the original disk storage location of the data units in each database file and the storage status of all disks in the shared cluster database to determine the target disk location to which the data units need to be migrated.

[0045] Step S202: for any data unit, obtain a data mapping table of a database file corresponding to the data unit.

[0046] Step S203 , lock-marking the mapping relationship of the data units recorded in the data mapping table to obtain a marked data mapping table.

[0047] In an embodiment of the present application, a data mapping table can be used to record the mapping relationship between all data units of a database file and the corresponding disk locations. The lock mark can refer to a shared lock (Shared Lock) set for recording the mapping relationship of the data unit in the data mapping table. The marked mapping table can refer to a data mapping table after the mapping relationship of the recorded data unit is locked. The lock mark in the marked mapping table can be used to allow each database node to perform a read operation on the data unit with the locked mapping relationship when migrating data in the data unit with the locked mapping relationship, and prohibit each database node from performing a write operation on the data unit with the locked mapping relationship.

[0048] like Figure 3 As shown, embodiment 2 of the present application provides a schematic diagram of the mapping relationship between a database file, a data mapping table, and a shared storage disk group. The database file includes 6 data units. The data mapping table of the database file records the mapping relationship between all data units of the database file and the corresponding disk positions. The shared storage disk group includes 3 disks. The database file is evenly stored in the disks of the shared storage disk group in units of data units. For example, according to the data mapping table, it can be determined that data unit 1 in the database file is stored in disk 3, data unit 2 is stored in disk 2, data unit 2 is stored in disk 1, data unit 4 is stored in disk 2, data unit 5 is stored in disk 1, and data unit 6 is stored in disk 3.

[0049] Specifically, first, for any data unit, the data mapping table corresponding to the data unit is obtained according to the database file corresponding to the data unit. Then, the obtained data mapping table is traversed to determine the mapping row in the data mapping table that records the mapping relationship between the data unit and the corresponding disk location. Finally, by calling the lock mechanism interface of the shared cluster database, the mapping relationship of the data unit recorded in the data mapping table is locked and marked to obtain the marked mapping table.

[0050] Step S204, migrate the data in the data unit according to the target disk location of the data unit. After the migration is completed, update the mapping relationship of the data unit recorded in the marked data mapping table according to the target disk location, and remove the lock mark to obtain an updated data mapping table.

[0051] In the embodiment of the present application, the updated data mapping table may refer to updating the mapping relationship of the data units recorded in the marked data mapping table and removing the lock mark of the data mapping table.

[0052] Specifically, before migration, the original disk position corresponding to the data unit is determined according to the data mapping table corresponding to the data unit, and the data in the data unit is read from the original disk position corresponding to the data unit. During the migration process, the data in the read data unit is written into the target disk position of the data unit. After the migration is completed, the marked data mapping table corresponding to the data unit is obtained, and the mapping row in the marked data mapping table that records the mapping relationship between the data unit and the corresponding disk position is determined. The disk position corresponding to the data unit in the mapping row is updated to the target disk position of the data unit, and the shared lock set on the mapping row is released to obtain an updated data mapping table, and the read and write processes waiting on the data unit are awakened. According to the updated data mapping table, the read and write operations are re-executed.

[0053] For example, a shared cluster database uses a shared disk architecture, and the data in the database file is stored in disk 1, disk 2, and disk 3 as data units.

[0054] When a new disk 4 is added to a shared cluster database for disk expansion, the overall process may be: 1) based on the scale factor of the data unit group in the database file corresponding to the disk and the number of data units in the data unit group, determine the data units in the database file that need to be migrated to disk 4, and form a migration mapping table for the database file; 2) for any data unit that needs to be migrated to disk 4, before migration, broadcast the migration message of the data unit to each database node. If the response message received from each database node to the migration message does not contain a response result that does not allow lock marking, obtain the data mapping table corresponding to the data unit, and lock mark the mapping relationship of the data unit in the data mapping table; 3) migrate the data unit to disk 4. After the migration is completed, update the migration status of the data unit in the migration mapping table corresponding to the data unit to the migrated status, update the mapping relationship of the data unit in the data mapping table, remove the lock mark, and re-execute the read and write operations on the data unit.

[0055] When a shared cluster database deletes disk 3 to perform disk shrinkage, the overall process may be as follows: 1) based on the scale factor of disk 1 and the number of data units in disk 3, determine the data units in disk 3 that need to be migrated to disk 1. Correspondingly, based on the scale factor of disk 2 and the number of data units in disk 3, determine the data units in disk 3 that need to be migrated to disk 2, and form a migration mapping table for the database file corresponding to the data unit in disk 3; 2) for any data unit to be migrated in disk 3, before migration, broadcast the migration message of the data unit to each database node. If the response message received from each database node to the migration message does not contain a response result indicating that lock marking is not allowed, obtain the data mapping table corresponding to the data unit, and lock mark the mapping relationship of the data unit in the data mapping table; 4) migrate the data unit. After the migration is completed, update the migration status of the database file in the migration mapping table corresponding to the data unit to the migrated status, update the mapping relationship of the data unit in the data mapping table, remove the lock mark, and re-execute the read and write operations on the data unit.

[0056] Optionally, when a new disk 4 is added to the shared cluster database for disk expansion, or when a disk 3 is deleted from the shared cluster database for disk reduction, for any data unit being migrated, a record of the migration operation of the data unit is formed into a redo log, which is used to replay the migration operation of the data unit and restore the migration status of the data unit after the shared cluster database fails and restarts during the migration process. For the database files corresponding to the data units being migrated, the migration status of the database files is formed into an undo log, which is used to determine the database files that have not been completely migrated after the shared cluster database fails and restarts during the migration process, and continue to execute the migration operation of the unmigrated data units based on the status information of the data units in the migration mapping table of the database files that have not been completely migrated.

[0057] An embodiment of the present application automatically migrates data units based on the target disk position calculation results of the data units in the database file when scaling the disks in a shared cluster database, and evenly distributes the data units to all the disks after scaling, thereby achieving automatic load balancing of all the disks after scaling, avoiding the situation where the scaled disks are frequently accessed, and thereby improving the efficiency of data migration and the utilization and access efficiency of all the disks after scaling. Before the data unit is migrated, the mapping relationship corresponding to the data unit is locked, so that during the data unit migration process, the data unit is allowed to be read and the data unit is prohibited to be written, thereby achieving automatic migration of the data unit and concurrent control of the database's read and write operations on the migrated data units, and ensuring the integrity and consistency of the data read and written by the database during the migration process.

[0058] See also Figure 4 , is a flow chart of an online scaling method for a database parallel file system provided in Example 3 of this application. Figure 4 As shown, before redistributing the database files in the shared cluster database in step S201 to obtain the target disk locations of all data units in each database file, the following steps may also be included:

[0059] Step S401: Obtain the available storage capacity of all disks and the total available storage capacity of all disks in the shared cluster database.

[0060] Step S402 : For any disk, obtain the weight of the disk according to the available storage capacity and the total available storage capacity of the disk.

[0061] In an embodiment of the present application, the available storage capacity may refer to the unused storage capacity in the disk, the total available storage capacity may refer to the total unused storage capacity of all disks in the shared cluster database, and the weight of the disk may refer to the weight assigned to the disk based on the available storage capacity of the disk.

[0062] Specifically, for any disk, the ratio of the disk's available storage capacity to the total available storage capacity can be calculated, and the ratio of the disk's available storage capacity to the total available storage capacity can be directly used as the disk's weight. Alternatively, the disk's input and output speed, the disk's access frequency, and the disk's specific requirements, etc., can be comprehensively calculated with the ratio of the disk's available storage capacity to the total available storage capacity, and the comprehensive calculation result can be used as the disk's weight.

[0063] Step S403, obtaining the total weight of all disks according to the weights of all disks;

[0064] Step S404: Obtain the disk's scaling factor based on the disk's weight and the total weight.

[0065] In an embodiment of the present application, the total weight may refer to the sum of the weights of all disks in the shared cluster database, and the proportional factor may be used to determine the number of data units that need to be migrated in each disk. When a new disk is added for disk expansion, it may be used to determine the number of data units that need to be migrated from any disk except the newly added disk. When a disk is deleted for disk reduction, it may be used to determine the number of data units that need to be migrated to any disk except the disk to be deleted.

[0066] Specifically, first, the weights of all disks can be added up to obtain the total weight of all disks. Then, if the total weight of all disks is 1, the weight of the disk can be directly used as the disk's scaling factor, or the weight of the disk can be comprehensively calculated with factors such as the disk's performance and reliability, and the comprehensive calculation result can be used as the disk's scaling factor; if the total weight of all disks is not 1, the ratio of the disk's weight to the total weight of all disks can be calculated, and the ratio of the disk's weight to the total weight of all disks can be directly used as the disk's scaling factor. Correspondingly, factors such as the disk's performance and reliability can also be comprehensively calculated with the ratio of the disk's weight to the total weight of all disks, and the comprehensive calculation result can be used as the disk's scaling factor.

[0067] In an embodiment of the present application, a disk scaling factor is determined based on the actual available storage capacity of the disk, thereby providing a proportional basis for determining the data units that need to be migrated and redistributed in the database file after the disk is expanded or reduced in capacity, thereby ensuring load balancing of all disks after expansion or reduction.

[0068] See also Figure 5 , is a flow chart of an online scaling method for a database parallel file system provided in the fourth embodiment of the present application. Figure 5As shown, in the above step S201, when the shared cluster database is expanded or reduced in disk capacity, the database files in the shared cluster database are redistributed to obtain the target disk location of all data units in each database file, which may include the following steps:

[0069] Step S501 : when performing disk capacity expansion on a shared cluster database, all database files and data mapping tables of all database files in the shared cluster database are obtained.

[0070] Step S502 : for any database file, determine the disk locations of all data units in the database file according to the data mapping table of the database file.

[0071] Step S503 : According to the disk locations, the data units in the database file that are located on the same disk are determined as a data unit group.

[0072] In the embodiment of the present application, the disk location may refer to the original disk storage location of the data unit before migration, and the data unit group may refer to a combination of data units whose disk locations are the same disk.

[0073] Specifically, in the process of determining the data unit group, first, all database files in the shared cluster database and the data mapping tables corresponding to all database files are obtained. Then, for any database file, the original disk positions of all data units in the database file are determined based on the mapping relationship between all data units and corresponding disk positions recorded in the data mapping table of the database file. Finally, the data units in the database file whose original disk positions are the same disk are determined as data unit groups, thereby obtaining a collection of all data unit groups of the database file.

[0074] Step S504 : for any data unit group, determine the data units to be migrated in the data unit group and the corresponding target disk locations according to the number of all data units in the data unit group and the scale factor of the disk corresponding to the data unit group.

[0075] In an embodiment of the present application, when a data unit is migrated, the target disk location is the disk location to which the data unit needs to be migrated. When a disk is expanded, the target disk location is the newly added disk for expansion of the shared cluster database.

[0076] Specifically, first, for any data unit group, determine the number of all data units in the data unit group and the scale factor of the disk corresponding to the data unit group. Then, multiply the number of all data units in the data unit group and the scale factor of the disk corresponding to the data unit group to determine the number of data units that need to be migrated in the data unit group. According to the number of data units that need to be migrated, determine the data units that need to be migrated in the data unit group. Thus, according to the number of data units that need to be migrated in all data unit groups in the database file and the data units that need to be migrated, determine the total number of data units that need to be migrated in the database file and all data units that need to be migrated in the database file. Finally, divide the total number of data units that need to be migrated in the database file by the total number of newly added disks to determine the number of data units that need to be migrated to each newly added disk, thereby determining the target disk position of the data units that need to be migrated in the database file.

[0077] For example, if the total number of newly added disks for capacity expansion of a shared cluster database is 2, and the total number of data units that need to be migrated in the database file is 100, it can be determined that the number of data units that need to be migrated to each newly added disk is 50. Based on the number of data units that need to be migrated to each newly added disk, which is 50, target disk locations are allocated for all data units that need to be migrated in the database file.

[0078] In the embodiment of the present application, by determining the total number of data units that need to be migrated for each database file, the data units that need to be migrated, and the target disk location of the data units that need to be migrated based on the disk's scale factor between disk expansion and data migration, the decision-making time and errors in the migration process are reduced, and the efficiency of data migration is improved while ensuring that all disks are load balanced after the migration is completed.

[0079] See also Figure 6 , is a flowchart of an online scaling method for a database parallel file system provided in Example 5 of the present application, such as Figure 6 As shown, in the above step S201, when the shared cluster database is expanded or reduced in disk capacity, the database files in the shared cluster database are redistributed to obtain the target disk location of all data units in each database file, which may include the following steps:

[0080] Step S601 : when performing disk shrinking on a shared cluster database, all data units in the shrunk disk are obtained, and the number of all data units in the shrunk disk is determined.

[0081] Step S602 : for any disk in the shared cluster database except the disk being shrunk, determine the data unit at the target disk location in the disk being shrunk based on the number of all data units in the disk being shrunk and the scale factor of the disk.

[0082] In an embodiment of the present application, the disk to be shrunk may refer to a disk to be deleted from a shared cluster database. When a data unit is migrated, the target disk location is the disk location to which the data unit needs to be migrated. When a disk is shrunk, the target disk location is any disk in the shared cluster database except the disk to be shrunk.

[0083] Specifically, first, determine the total number of all data units in the disk to be deleted, and then, for any disk in the shared cluster database except the disk to be deleted, multiply the disk's scale factor by the total number of all data units in the disk to be deleted to determine the number of data units that need to be migrated to the disk. Based on this number, determine the data units in the disk to be deleted that need to be migrated to the disk, and thus determine the number of data units and data units that need to be migrated to each disk except the disk to be deleted, and obtain the number of data units to be migrated, the data units to be migrated and the corresponding target disk positions in each database file corresponding to all data units in the disk to be deleted.

[0084] For example, the total number of disks to be deleted for shrinking a shared cluster database is 2, and the total number of all data units in the disks to be deleted is 100. There are 4 disks in the shared cluster database excluding the disk to be deleted. If the number of data units to be migrated to each disk is determined to be 25 based on the proportional factor of each disk, then based on this number 25, the data units to be migrated to each disk in the two disks to be deleted are determined, thereby determining the number of data units to be migrated in each database file corresponding to all the data units in the two disks to be deleted, the data units to be migrated, and the corresponding target disk positions.

[0085] In an embodiment of the present application, after disk reduction and before data migration, the number of data units that need to be migrated, the data units that need to be migrated, and the target disk locations of the data units that need to be migrated in each database file corresponding to all data units in the disk being reduced in size are determined based on the disk's scale factor. This reduces decision-making time and errors in the migration process, improves the efficiency of data migration, and ensures that all disks are load balanced after the migration is completed.

[0086] See also Figure 7 , is a flow chart of an online scaling method for a database parallel file system provided in Example 6 of this application. Figure 7As shown, after obtaining the target disk locations of all data units in each database file in the above step S201, the following steps may also be included:

[0087] Step S701: for any database file, obtain all the data units to be migrated in the database file and the corresponding target disk locations.

[0088] Step S702: All the data units to be migrated and the corresponding target disk locations are combined to form a migration mapping table of the database file.

[0089] In an embodiment of the present application, the migration mapping table may refer to a file that records all data unit information to be migrated in the database file, wherein the migration mapping table may include all data unit information to be migrated in the corresponding database file and the corresponding target disk location information.

[0090] Specifically, when expanding the disk capacity of a shared cluster database, all data units that need to be migrated in the database file and the corresponding target disk locations can be determined based on the contents of steps S501 to S504 above. When shrinking the disk capacity of a shared cluster database, all data units that need to be migrated in the database file and the corresponding target disk locations can be determined based on the contents of steps S601 to S602 above. Based on all data units that need to be migrated in the database file and the corresponding target disk locations, a migration mapping table for the database file is formed.

[0091] Step S703, after migrating the data in the data unit according to the target disk location of the data unit, further includes: obtaining a migration mapping table of the database file corresponding to the data unit, and setting the state of the data unit in the migration mapping table to a migrated state.

[0092] In the embodiment of the present application, the state of the data unit may refer to the migration state of the data unit, wherein the state of the data unit may include an unmigrated state, a migrating state, and a migrated state.

[0093] Specifically, after the data in the data unit is migrated according to the target disk location of the data unit, a migration mapping table of the database file corresponding to the data unit is obtained, and the state of the data unit in the migration mapping table is updated to a migrated state.

[0094] In the embodiment of the present application, the migration mapping table provides clear migration data information and migration path information for data migration, reduces decision-making time and errors during the migration process, improves the efficiency of migration, and updates the status of the data unit to the migrated state after the migration is completed, reducing the risk of data loss or repeated migration during the migration process. At the same time, it provides the necessary information to support rollback and recovery operations for fault handling during the migration process, thereby ensuring the sustainability of the migration job.

[0095] See also Figure 8 , is a flow chart of an online scaling method for a database parallel file system provided in Example 7 of this application. Figure 8 As shown, the online scaling method may include the following steps:

[0096] Step S801 : for any data unit, when migrating the data unit, the migration operation of the data unit is recorded to form a redo log.

[0097] In embodiments of the present application, the redo log may refer to a REDO log, which can be used to replay the migration operation of a data unit according to the redo log after detecting a shared cluster database failure and restarting, restore the migration status of the data unit, and persist the data unit to the corresponding target disk location. Specifically, when a data unit is migrated, the migration operation of the data unit is recorded in the redo log.

[0098] Step S802: Generate an undo log of the migration status of the database file corresponding to the data unit.

[0099] In embodiments of the present application, the undo log may refer to an UNDO log. This log can be used to determine, based on the undo log, any database files in the shared cluster database that have not been migrated, after a shared cluster database failure restart is detected. The undo log can also be used to obtain a migration mapping table for the unmigrated database files. Based on the status information of the data units in the migration mapping table for the unmigrated database files, the unmigrated data units can be determined, and migration operations for the unmigrated data units can be continued. Specifically, when a data unit in a database file is migrated, the migration status of the database file is recorded in the undo log.

[0100] In the embodiment of the present application, by forming redo logs and undo logs, after the database fails and restarts during the migration process, it can quickly recover to the state before the failure and re-execute those migration operations that failed to be completed due to the failure, thereby ensuring the consistency and integrity of the data and providing guarantees for the sustainability of the migration operation.

[0101] See also Figure 9 , is a flow chart of an online scaling method for a database parallel file system provided in Example 8 of the present application. Figure 9As shown, before locking the mapping relationship of the data unit recorded in the data mapping table in the above step S203, the following steps may also be included:

[0102] Step S901: broadcasting a migration message of a data unit to a shared cluster database.

[0103] Step S902: Receive the response results of each database node in the shared cluster database to the migration message. If there is a response result that does not allow lock marking among all the received response results, continue to broadcast the migration message of the data unit to the shared cluster database until there is no response result that does not allow lock marking among all the received response results.

[0104] In an embodiment of the present application, the migration message may refer to a communication message that informs each database node in a shared cluster database that a data unit is about to be migrated, wherein the migration message may include identification information of the data unit to be migrated, the start time of the migration, the end time of the migration, and information on whether the data unit is allowed to be locked, etc. The response result may refer to a communication message that each database node in the shared cluster database feeds back the migration message.

[0105] Specifically, the migration message of the data unit to be migrated can be broadcast to each database node in the shared cluster database based on the User Datagram Protocol (UDP). When receiving the migration message, each database node forwards the migration message to the file management process to determine whether a write operation is being performed on the data unit corresponding to the migration message. If a write operation is being performed on the data unit, a response result indicating that lock marking is not allowed is fed back. If a write operation is not being performed on the data unit corresponding to the migration message, a response result indicating that lock marking is allowed is fed back. If a response result indicating that lock marking is not allowed exists among all the received response results, the migration message of the data unit broadcast by the shared cluster database is continuously executed until no response result indicating that lock marking is not allowed exists among all the received response results.

[0106] The embodiment of the present application ensures that each database node is aware of the current migration operation by broadcasting the migration message of the data unit to each database node in the shared cluster database and waiting for the response of each database node, thereby reducing system errors or failures caused by the lack of understanding of the database node, and ensuring that all database nodes are allowed to perform lock marking, avoiding the situation where the lock marking cannot be performed due to the database node being in the process of modifying the currently migrated data unit, resulting in an error in the lock marking.

[0107] Corresponding to the online expansion and contraction method of the database parallel file system in the above embodiment, Figure 10The structure diagram of the online expansion and contraction device of the shared cluster database provided by the ninth embodiment of the present application is shown. The online expansion and contraction device of the shared cluster database is applied to Figure 1 The server in the embodiment communicates with the shared cluster database to obtain the database file and data mapping table sent by the shared cluster database. For ease of description, only the parts related to the embodiment of the present application are shown.

[0108] See also Figure 10 , the online expansion and contraction device comprises:

[0109] The redistribution calculation module 1001 is used to redistribute the database files in the shared cluster database when the shared cluster database is expanded or reduced in disk capacity, and obtain the target disk location of all data units in each database file;

[0110] A mapping table acquisition module 1002 is configured to acquire, for any data unit, a data mapping table of a database file corresponding to the data unit, wherein the data mapping table is configured to record a mapping relationship between all data units belonging to the database file and corresponding disk locations;

[0111] a lock marking module 1003, configured to lock-mark the mapping relationship of the data unit recorded in the data mapping table to obtain a marked data mapping table, wherein the lock mark in the marked data mapping table is used to allow read operations and prohibit write operations when migrating data in the data unit;

[0112] The migration module 1004 is used to migrate the data in the data unit according to the target disk position of the data unit. After the migration is completed, the mapping relationship of the data unit recorded in the marked data mapping table is updated according to the target disk position, and the lock mark is removed to obtain an updated data mapping table.

[0113] Optionally, the online expansion and contraction device further includes:

[0114] A capacity acquisition module is used to obtain the available storage capacity of all disks in the shared cluster database and the total available storage capacity of all disks;

[0115] A weight determination module, configured to obtain, for any disk, a weight of the disk according to the available storage capacity of the disk and the total available storage capacity;

[0116] The total weight determination module is used to obtain the total weight of all disks based on the weights of all disks;

[0117] The scale factor determination module is configured to obtain the scale factor of the disk according to the weight of the disk.

[0118] Optionally, the redistribution calculation module 1001 includes:

[0119] The expansion acquisition module is used to obtain all database files and data mapping tables of all database files in the shared cluster database when the disk capacity of the shared cluster database is expanded;

[0120] A location determination module, configured to determine, for any database file, the disk locations of all data units in the database file according to a data mapping table of the database file;

[0121] A unit group determining module, configured to determine, according to the disk positions, data units in the database file whose disk positions are on the same disk as a data unit group;

[0122] The capacity expansion determination module is used to determine, for any data unit group, the data units to be migrated in the data unit group and the corresponding target disk locations based on the number of all data units in the data unit group and the scale factor of the disk corresponding to the data unit group.

[0123] Optionally, the redistribution calculation module 1001 includes:

[0124] A shrinking acquisition module is used to acquire all data units in the disk being shrunk when shrinking the shared cluster database, and determine the number of all data units in the disk being shrunk;

[0125] The shrinkage determination module is used to determine, for any disk in the shared cluster database except the disk to be shrunk, the data unit at the target disk position in the disk to be shrunk based on the number of all data units in the disk to be shrunk and the scale factor of the disk.

[0126] Optionally, the online expansion and contraction device further includes:

[0127] A location acquisition module is used to acquire, for any database file, all the data units to be migrated in the database file and the corresponding target disk locations;

[0128] A file forming module, configured to form a migration mapping table of the database file by combining all the data units to be migrated and the corresponding target disk locations;

[0129] The status setting module is used to, after migrating the data in the data unit according to the target disk location of the data unit, further include: obtaining the migration mapping table of the database file corresponding to the data unit, and setting the status of the data unit in the migration mapping table to the migrated state.

[0130] Optionally, the online expansion and contraction device further includes:

[0131] A redo log recording module is used to record the migration operation of any data unit to form a redo log when migrating the data unit. The redo log is used to replay the migration operation of the data unit according to the redo log after detecting the restart of the shared cluster database failure, restore the migration status of the data unit, and persist the data unit to the corresponding target disk location;

[0132] An undo log recording module is used to form an undo log with the migration status of the database file corresponding to the data unit. The undo log is used to determine the database files that have not been completely migrated in the shared cluster database according to the undo log after detecting the restart of the shared cluster database failure, and obtain the migration mapping table of the database files that have not been completely migrated. According to the migration mapping table of the database files that have not been completely migrated, the data units that have not been migrated are determined, and the migration operation of the data units that have not been migrated is continued.

[0133] Optionally, the online expansion and contraction device further includes:

[0134] A broadcast module, configured to broadcast the migration message of the data unit to the shared cluster database;

[0135] A receiving module is configured to receive response results fed back by each database node in the shared cluster database to the migration message, and if a response result indicating that lock marking is not allowed exists among all received response results, the migration message of broadcasting the data unit to the shared cluster database is continuously executed until no response result indicating that lock marking is not allowed exists among all received response results.

[0136] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0137] Figure 11 This is a schematic diagram of the structure of a computer device provided in Example 10 of this application. Figure 11 As shown, the computer device of this embodiment includes: at least one processor ( Figure 11 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in any of the above-mentioned embodiments of the online scaling method for a parallel file system of a database are implemented.

[0138] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 11 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0139] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.

[0140] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0141] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0142] The present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing it.

[0143] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0144] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0145] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.

[0146] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0147] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for online expansion and contraction of a database parallel file system, characterized in that: The online expansion and contraction method comprises: When the shared cluster database is expanded or reduced in disk capacity, the database files in the shared cluster database are redistributed to obtain the target disk location of all data units in each database file; For any data unit, obtain a data mapping table of a database file corresponding to the data unit, wherein the data mapping table is used to record a mapping relationship between all data units belonging to the database file and corresponding disk locations; Lock-marking the mapping relationship of the data unit recorded in the data mapping table to obtain a marked data mapping table, wherein the lock mark in the marked data mapping table is used to allow a read operation and prohibit a write operation when migrating data in the data unit; Migrating the data in the data unit according to the target disk position of the data unit, and after the migration is completed, updating the mapping relationship of the data unit recorded in the marked data mapping table according to the target disk position, and removing the lock mark to obtain an updated data mapping table; Before redistributing the database files in the shared cluster database to obtain the target disk locations of all data units in each database file, the method further includes: Obtaining the available storage capacity of all disks in the shared cluster database and the total available storage capacity of all disks; For any disk, obtaining a weight of the disk according to the available storage capacity of the disk and the total available storage capacity; According to the weights of all disks, get the total weight of all disks; Obtaining a scaling factor of the disk according to the weight of the disk and the total weight; When the shared cluster database is expanded or reduced in disk capacity, the database files in the shared cluster database are redistributed to obtain the target disk location of all data units in each database file, including: When the shared cluster database is expanded, all database files and data mapping tables of all database files in the shared cluster database are obtained; For any database file, determining the disk locations of all data units in the database file according to the data mapping table of the database file; According to the disk positions, the data units in the database file whose disk positions are the same disk are determined as a data unit group; For any data unit group, the number of all data units in the data unit group and the scale factor of the disk corresponding to the data unit group are multiplied to determine the number of data units in the data unit group that need to be migrated, and the data units in the data unit group that need to be migrated are determined according to the number of data units in the data unit group that need to be migrated; Determine the total number of data units that need to be migrated and all data units that need to be migrated in the database file according to the number of data units that need to be migrated in all data unit groups in the database file and the data units that need to be migrated; The total number of data units that need to be migrated in all database files in the shared cluster database is calculated by adding up the total number of data units that need to be migrated in the shared cluster database; Determine the total number of newly added disks when expanding the disk capacity of the shared cluster database, divide the total number of data units that need to be migrated in the shared cluster database by the total number of newly added disks, and determine the number of data units that need to be migrated to each newly added disk and the target disk position corresponding to the data units that need to be migrated.

2. The online expansion and contraction method of a database parallel file system according to claim 1, characterized in that: When the shared cluster database is expanded or reduced in disk capacity, the database files in the shared cluster database are redistributed to obtain the target disk location of all data units in each database file, including: When the shared cluster database is subjected to disk shrinkage, all data units in the disk being shrunk are obtained, and the number of all data units in the disk being shrunk is determined; For any disk in the shared cluster database except the disk to be shrunk, the data unit at the target disk position in the disk to be shrunk is determined according to the number of all data units in the disk to be shrunk and the scale factor of the disk.

3. The online expansion and contraction method of a database parallel file system according to claim 1, characterized in that: After obtaining the target disk locations of all data units in each database file, the method further includes: For any database file, obtain all the data units to be migrated in the database file and the corresponding target disk locations; Forming a migration mapping table of the database file with all the data units to be migrated and the corresponding target disk locations; After migrating the data in the data unit according to the target disk location of the data unit, the method further includes: A migration mapping table of a database file corresponding to the data unit is obtained, and a state of the data unit in the migration mapping table is set to a migrated state.

4. The online expansion and contraction method of a database parallel file system according to claim 3, characterized in that: The online expansion and contraction method further includes: For any data unit, when migrating the data unit, the migration operation record of the data unit is formed into a redo log. The redo log is used to replay the migration operation of the data unit according to the redo log after detecting the failure and restart of the shared cluster database, restore the migration status of the data unit, and persist the data unit to the corresponding target disk location; The migration status of the database file corresponding to the data unit is formed into an undo log. The undo log is used to determine the database files that have not been completely migrated in the shared cluster database according to the undo log after the shared cluster database failure is detected and restarted, and to obtain the migration mapping table of the database files that have not been completely migrated. According to the migration mapping table of the database files that have not been completely migrated, the data units that have not been migrated are determined, and the migration operation on the data units that have not been migrated is continued.

5. The online expansion and contraction method of a database parallel file system according to claim 1, characterized in that: Before locking and marking the mapping relationship of the data unit recorded in the data mapping table, the method further includes: broadcasting a migration message of the data unit to the shared cluster database; Receive response results fed back by each database node in the shared cluster database to the migration message, and if there is a response result that does not allow lock marking among all the received response results, continue to execute the migration message of broadcasting the data unit to the shared cluster database until there is no response result that does not allow lock marking among all the received response results.

6. An online expansion and contraction device for a shared cluster database, characterized in that: The online expansion and contraction device comprises: A redistribution calculation module is used to redistribute database files in the shared cluster database when the shared cluster database is expanded or reduced in disk capacity, so as to obtain the target disk location of all data units in each database file; A mapping table acquisition module, used for acquiring, for any data unit, a data mapping table of a database file corresponding to the data unit, wherein the data mapping table is used for recording a mapping relationship between all data units belonging to the database file and corresponding disk locations; A lock marking module, used for locking and marking the mapping relationship of the data unit recorded in the data mapping table to obtain a marked data mapping table, wherein the lock mark in the marked data mapping table is used to allow read operations and prohibit write operations when migrating data in the data unit; A migration module, configured to migrate the data in the data unit according to the target disk position of the data unit, and after the migration is completed, update the mapping relationship of the data unit recorded in the marked data mapping table according to the target disk position, and remove the lock mark to obtain an updated data mapping table; A capacity acquisition module, used to acquire the available storage capacity of all disks in the shared cluster database and the total available storage capacity of all disks; A weight determination module, used for obtaining the weight of any disk according to the available storage capacity of the disk and the total available storage capacity; A total weight determination module is used to obtain the total weight of all disks according to the weights of all disks; A scale factor determination module, used for obtaining the scale factor of the disk according to the weight of the disk; The redistribution calculation module comprises: The expansion acquisition module is used to acquire all database files and data mapping tables of all database files in the shared cluster database when the disk capacity of the shared cluster database is expanded; A location determination module, used for determining, for any database file, the disk locations of all data units in the database file according to the data mapping table of the database file; A unit group determination module, configured to determine, according to the disk positions, data units in the database file whose disk positions are the same disk as a data unit group; The capacity expansion determination module is used to determine the number of data units that need to be migrated in the data unit group by multiplying the number of all data units in the data unit group by the scale factor of the disk corresponding to the data unit group for any data unit group, and determine the data units that need to be migrated in the data unit group according to the number of data units that need to be migrated in the data unit group; Determine the total number of data units that need to be migrated and all data units that need to be migrated in the database file according to the number of data units that need to be migrated in all data unit groups in the database file and the data units that need to be migrated; The total number of data units that need to be migrated in all database files in the shared cluster database is calculated by adding up the total number of data units that need to be migrated in the shared cluster database; Determine the total number of newly added disks when expanding the disk capacity of the shared cluster database, divide the total number of data units that need to be migrated in the shared cluster database by the total number of newly added disks, and determine the number of data units that need to be migrated to each newly added disk and the target disk position corresponding to the data units that need to be migrated.

7. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the online scaling method for the database parallel file system as described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the online expansion and contraction method of the database parallel file system as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Distributed database system and data processing method in distributed database system

    CN106708968A

  • Data processing method and device

    CN113946542A

  • Data balancing method and device of data storage system and computer readable medium

    CN116501260A

  • Data migration method and device, computer equipment and storage medium

    CN117648056A