Mass data organization and storage method and device, electronic equipment and storage medium
Through partition offloading and management methods, the target data set segment is maintained locally from the overall data set, solving the problem of atomic application scenarios of high availability, easy-to-harness and hyper-transactions in traditional methods, and achieving efficient data management and consistency guarantees.
Patent Information
- Application Number
- CN202510934543.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Traditional massive data organization and storage methods are difficult to meet the needs of high availability, easy to control and super-transaction atomic application scenarios. Especially in natural resource data management, multi-level collaborative management and data failure isolation capabilities are insufficient, which affects the high availability of data and the super-transaction atomicity of concurrent access.
The partition offload method is used to separate the target data set segments in the target physical storage partition from the overall data set. Through partition offloading, management and mount methods, local maintenance and isolation of data are achieved, the availability of other partitions is ensured, and data consistency and visibility are ensured through segmentation rules and data structure verification.
It realizes high-availability data maintenance, improves data maintenance efficiency, meets the needs of ease of control, and ensures the consistency of data content through physical storage partitions, meets the atomic application scenarios of super-transactions, and solves the shortcomings of traditional methods.
Smart Images

Figure CN120429299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a method, device, electronic device and storage medium for organizing and storing massive data. Background Art
[0002] Through various efforts, the Natural Resources Department has accumulated thematic data containing tens of millions of elements. This data plays a crucial role in natural resource operations. However, due to its fundamental characteristics—massive, multi-source, heterogeneous, and dynamic—it requires continuous maintenance and dynamic updates. Currently, database partitioning technology has addressed the challenges of massive data storage and high-performance, high-concurrency applications. However, given the inherent characteristics (spatial and temporal) of natural resource data and application needs, the requirements for high data availability and manageability have not been fully met. Furthermore, the need for atomic, hyper-transactional applications of massive data under concurrent access has been overlooked.
[0003] Natural resource operations, which generate massive amounts of data, often feature collaborative operations at the national, provincial, municipal, and county levels. The organization and management of these data outputs must simultaneously meet the needs of all four levels, requiring multi-granularity organization and management. Currently, databases for these natural resource data at the municipal level and above are either simply organized based on the most basic unit of data production (e.g., county-level administrative districts) or randomly organize data for the entire region within a single storage space. These database construction techniques fail to fully reflect the business characteristics of multi-level division of labor and collaborative management, and lack data management capabilities independent of application logic and multi-granularity atomicity support. For example, databases built with a single storage space struggle to ensure transactional integrity and minimize fault isolation for data entry and updates managed at the county level. Provincial databases, simply organized at the county-level administrative district level, increase the complexity of backup and recovery operations at both the municipal and provincial levels. Both of these database construction approaches also pose the risk that failures in local data and its derivatives often impact overall database accessibility, significantly compromising the high availability requirements of ubiquitous sharing.
[0004] At present, the natural resources system applies database partitioning technology to solve the problems of high-performance and high-concurrency applications of massive data, but basically ignores the high availability, easy control and super-transaction atomic application scenario requirements of massive data. It often relies on high-cost, high-tech complexity and difficult-to-maintain cluster hardware and software deployment to meet the high availability requirements of massive databases, and relies on writing complex business codes to implement super-transaction atomic business scenario applications.
[0005] In summary, traditional methods of organizing and storing massive amounts of data are unable to fully meet the requirements of high-availability, easy-to-manage, and super-transactional atomic application scenarios. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method, device, electronic device and storage medium for organizing and storing massive data, so as to alleviate the problem that traditional methods for organizing and storing massive data are difficult to fully meet the requirements of high availability, easy control and super-transactional atomic application scenarios.
[0007] In a first aspect, an embodiment of the present invention provides a method for organizing and storing massive data, which is applied to a data management service interface at a partition granularity. The method includes: Unloading a target data set segment in a target physical storage partition from the overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition can independently perform data addition, deletion, and modification operations on the overall data set composed of data set segments in other physical storage partitions. Except for the application system that calls the partition unloading method, other application systems cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented. Managing the target data set segments in the target physical storage partition according to the target partition management function called by the application system to obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; Verifying the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall data set to verify whether there is a conflict, and verifying the data structure of the target data set segment in the managed target physical storage partition against the data structure of the overall data set to verify whether the data structures are consistent; After the verification is passed, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs according to the partition mounting method called by the application system for access by all application systems.
[0008] Furthermore, before unloading the target data set segment in the target physical storage partition from the overall data set according to the partition unloading method called by the application system, pre-preparation is performed, the method further comprising: When the application system calls the data management service interface, it determines a segmentation rule set for segmenting the entire data set, and segments the entire data set according to the segmentation rule set to obtain a plurality of data set segments; A plurality of physical storage partitions are created according to the segmentation rule set, and a mapping relationship is established between each of the data set segments and the corresponding physical storage partition, so that each of the data set segments is stored in the corresponding physical storage partition.
[0009] Furthermore, determining a segmentation rule set for segmenting the entire data set includes: A segmentation rule set for segmenting the entire data set is determined according to manually selected intrinsic features of the entire data set, wherein the intrinsic features include: time features and spatial features.
[0010] Furthermore, creating a plurality of physical storage partitions according to the segmentation rule set and establishing a mapping relationship between each of the data set segments and the corresponding physical storage partitions includes: A physical storage partition corresponding to each segmentation rule is created according to user customization, and then a mapping relationship between each data set segment and the corresponding physical storage partition is established according to each segmentation rule.
[0011] Furthermore, the partition unloading method carries data isolation boundary information, and unloads the target data set segment in the target physical storage partition from the overall data set according to the partition unloading method called by the application system, including: Determining a target physical storage partition corresponding to the data isolation boundary information; Unloading the target data set segment in the target physical storage partition from the overall data set.
[0012] Furthermore, the partition clearing includes logical clearing and physical clearing. The logical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, but not release the physical storage space occupied by the target physical storage partition. The physical clearing is to delete the target data set segment and the index partition corresponding to the target physical storage partition, and release the other physical storage space occupied by the target physical storage partition except for retaining the necessary physical storage space. Deleting the partition while deleting the target physical storage partition merges the segmentation rules of the deleted target physical storage partition into the data segmentation rule set of the overall data set; The migration partition is to copy the target data set segment in the target physical storage partition to be migrated to the first new physical storage partition, delete the migrated target physical storage partition, update the mapping relationship between the target data set segment and the corresponding first new physical storage partition, and rebuild the partition index corresponding to the target data set segment; The clone partition is to copy the target dataset segment in the cloned target physical storage partition to the second new physical storage partition, and create a corresponding partition index for the target dataset segment in the second new physical storage partition according to the index corresponding to the cloned target dataset segment and its index type and parameters; The partition splitting step comprises generating a plurality of sub-segmentation rules according to the segmentation rules corresponding to the split data set segments, assigning a respective physical storage partition to each of the sub-segmentation rules to obtain a plurality of target physical storage sub-partitions, migrating the data set segments from the split target physical storage partitions to the corresponding target physical storage sub-partitions according to the sub-segmentation rules, and creating corresponding partition indexes for the plurality of target data set sub-segments formed by the migration according to the indexes corresponding to the split target data set segments and their index types and parameters; The merging partitions is to generate a merging rule according to target segmentation rules corresponding to the multiple target dataset segments to be merged, allocate a corresponding physical storage partition to the merging rule, obtain a merged target physical storage partition, migrate the multiple target dataset segments to the merged target physical storage partition to form a merged dataset segment, release the physical storage of the merged target physical storage partition, and create a corresponding partition index on the merged dataset segment according to the index status of the overall dataset; The partition index maintenance is to rebuild and maintain the partition index on the target data set segment in the unloaded target physical storage partition.
[0013] Furthermore, the partition mounting method carries data boundary information, and according to the partition mounting method called by the application system, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs for access by all application systems, including: Determine a target physical storage partition corresponding to the data boundary information and having completed management; The target data set segments in the target physical storage partitions whose management has been completed are incorporated into the overall data set to which they belong for access by all application systems.
[0014] In a second aspect, an embodiment of the present invention further provides a device for organizing and storing massive data, which is applied to a data management service interface at a partition granularity, and the device includes: an unloading unit, configured to unload a target data set segment in a target physical storage partition from the overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition can independently perform data addition, deletion, and modification operations on the overall data set composed of data set segments in other physical storage partitions, and other application systems cannot access the target data set segment in the target physical storage partition except the application system that calls the partition unloading method, wherein when performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed, and the target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; a management unit, configured to manage the target data set segments in the target physical storage partition according to the target partition management function called by the application system, and obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; a verification unit, configured to verify the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall data set to verify whether there is a conflict, and to verify the data structure of the target data set segment in the managed target physical storage partition against the data structure of the overall data set to verify whether the data structures are consistent; The mounting unit is used to incorporate the target data set segment in the managed target physical storage partition into the overall data set to which it belongs for access by all application systems according to the partition mounting method called by the application system after verification.
[0015] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to execute any method described in the first aspect above.
[0017] In an embodiment of the present invention, a method for organizing and storing massive data is provided, which is applied to a data management business interface at a partition granularity. The method comprises: unloading a target data set segment in a target physical storage partition from an overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition performs data addition, deletion, and modification operations independently from the overall data set composed of data set segments in other physical storage partitions. Except for the application system that calls the partition unloading method, other application systems cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; according to the application The target partition management function called by the system manages the target data set segments in the target physical storage partition to obtain the target physical storage partition that has been managed, wherein the target partition management function includes at least one of the following: clearing partitions, deleting partitions, migrating partitions, cloning partitions, splitting partitions, merging partitions, and maintaining partition indexes; verifying the target segmentation rules corresponding to the target physical storage partition that has been managed against the segmentation rule set of the overall data set to verify whether there is a conflict, and verifying the data structure of the target data set segments in the target physical storage partition that has been managed against the data structure of the overall data set to verify whether the data structures are consistent; after the verification is passed, the target data set segments in the target physical storage partition that has been managed are incorporated into the overall data set to which they belong for access by all application systems according to the partition mounting method called by the application system.From the above description, it can be seen that in the method of organizing and storing massive data of the present invention, the partition unloading method is adopted to perform data maintenance on the target data set segment in the target physical storage partition independently of the overall data set composed of the data set segments in other physical storage partitions. That is to say, the data management and processing operations occurring on the target physical storage partition can not affect the availability of the data of other physical storage partitions on the same overall data set to the application system. That is, when there is a problem with the data of the target physical storage partition and data maintenance is required, the data in other physical storage partitions can still be used, meeting the high-availability data requirements. In addition, when data maintenance is required, it is only necessary to maintain the target physical storage partition corresponding to the data, and there is no need to perform overall maintenance on the overall data set. Only local data needs to be updated and maintained, and there is no need to maintain the entire data set, which greatly improves the efficiency of data maintenance and meets the demand for easy-to-manage data. In addition, the data set segments in the physical storage partition are used as atomic units to ensure the consistency of data content. Data in incomplete physical storage partitions during the data processing period is invisible to business applications (i.e., application systems). That is, for massive overall data sets, the visibility of local data to the business can be dynamically controlled, meeting the needs of atomic application scenarios of super transactions. In other words, the method of organizing and storing massive data of the present invention can meet the needs of high availability, easy management and atomic application scenarios of super transactions, alleviating the difficulty of traditional massive data organization and storage methods in fully meeting the needs of atomic application scenarios of high availability, easy management and super transactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flowchart of a method for organizing and storing massive data provided by an embodiment of the present invention; Figure 2 A schematic diagram of a device for organizing and storing massive amounts of data provided by an embodiment of the present invention; Figure 3 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following is a clear and complete description of the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] Traditional methods of organizing and storing massive amounts of data are unable to fully meet the requirements of high-availability, easy-to-manage, and super-transactional atomic application scenarios.
[0022] Based on this, in the method for organizing and storing massive data of the present invention, a partition unloading method is adopted to perform data maintenance on the target data set segment in the target physical storage partition independently of the overall data set composed of the data set segments in other physical storage partitions. That is to say, the data management and processing operations occurring on the target physical storage partition may not affect the availability of the data of other physical storage partitions on the same overall data set to the application system. That is, when there is a problem with the data of the target physical storage partition and data maintenance is required, the data in other physical storage partitions can still be used, meeting the high-availability data requirements. In addition, when data maintenance is required, only the target physical storage partition corresponding to the data needs to be maintained. That is, there is no need to perform overall maintenance on the entire data set, that is, only local data needs to be updated and maintained, and there is no need to maintain the entire data set, which greatly improves the efficiency of data maintenance and meets the demand for easy-to-manage data. In addition, the data set segments in the physical storage partitions are used as atomic units to ensure the consistency of data content. The data in the incomplete physical storage partitions during the data processing period is invisible to the business application (that is, the application system). That is, for the massive overall data set, the visibility of local data to the business can be dynamically controlled, which meets the needs of super-transactional atomic application scenarios. That is, the method of organizing and storing massive data of the present invention can meet the needs of high availability, easy control and super-transactional atomic application scenarios.
[0023] To facilitate understanding of this embodiment, a method for organizing and storing massive data disclosed in an embodiment of the present invention is first introduced in detail.
[0024] Example 1: According to an embodiment of the present invention, an embodiment of a method for organizing and storing massive data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0025] Figure 1FIG. 1 is a flow chart of a method for organizing and storing massive data according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: Step S102: Unloading the target dataset segment in the target physical storage partition from the overall dataset according to the partition unloading method called by the application system, so that the target dataset segment in the target physical storage partition can perform data addition, deletion, and modification operations independently from the overall dataset composed of the dataset segments in other physical storage partitions. Except for the application system that called the partition unloading method, other application systems cannot access the target dataset segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target dataset segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target dataset segment when the overall dataset is segmented. In an embodiment of the present invention, the aforementioned method for organizing and storing massive amounts of data can be applied to a partition-level data management service interface. When an application system desires to operate directly on a target dataset segment rather than the overall dataset, such as isolating defective data or optimizing and rebuilding a local index, the application system invokes a partition unloading method. When the partition unloading method is invoked, data isolation boundary information (e.g., logical isolation boundary and / or start time boundary) for a target dataset segment is independently managed and maintained within the overall dataset, and the target dataset segment in the target physical storage partition is then unloaded from the overall dataset.
[0026] The above-mentioned target physical storage partition (i.e., the physical storage partition to be uninstalled) can be a newly created empty physical storage partition, or it can be a physical storage partition that already has a target data set segment; although the uninstalled physical storage partition becomes another data set independent of the overall data set, the program logic of the interface records the data segmentation rules that the original physical storage partition (i.e., the physical storage partition before being uninstalled) should follow, and implements them when independently performing data addition, deletion, and modification operations on the target physical storage partition, so that the target physical storage partition can still follow the segmentation rules of the overall data set when it is mounted.
[0027] Taking the land classification patch dataset of Province X as an example, the unloaded subset of land classification patch data for County Y in Province X (i.e., the target dataset segment) is not allowed to be inserted into land classification patch data from other county-level administrative districts in Province X. Only land classification patch data from County Y in Province X can be inserted. In other words, data addition, deletion, and modification operations on the target dataset segment in the target physical storage partition follow the corresponding target segment rules. Because partition unloading only involves updating the metadata describing the entire dataset and its partitions, the impact on business applications (i.e., application systems) accessing the entire dataset (i.e., the entire logical partition dataset) during this operation is negligible. Except for applications maintaining the unloaded target dataset segment, after the target physical storage partition is unloaded, the data within it is invisible to business applications on the entire dataset (i.e., other application systems cannot access the target dataset segment in the target physical storage partition except the application that invoked the partition unloading method).
[0028] Step S104: managing the target data set segments in the target physical storage partition according to the target partition management function called by the application system to obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; Step S106: Verify the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall dataset to verify whether there is a conflict, and verify the data structure of the target dataset segment in the managed target physical storage partition against the data structure of the overall dataset to verify whether the data structures are consistent. Specifically, before mounting a partition, the interface (i.e., the partition-granular data management service interface) compares the target segmentation rules corresponding to the managed target physical storage partition with the segmentation rule set of the overall dataset (the individual segmentation rules corresponding to the overall dataset), and compares the data structure of the target dataset segments within the managed target physical storage partition with the data structure of the overall dataset. This verifies whether the unmounted target physical storage partition can be mounted onto the overall dataset. If the target segmentation rules corresponding to the managed target physical storage partition conflict with the individual segmentation rules corresponding to the overall dataset, or if the data structure of the target dataset segments within the managed target physical storage partition is inconsistent with the data structure of the overall dataset, subsequent mounting is not performed, and the process returns to step S104 above until verification is successful.
[0029] Step S108 , after verification is passed, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs according to the partition mounting method called by the application system for access by all application systems.
[0030] Specifically, in contrast to partition unmounting, when an application system determines that a partitioning operation requiring independence from the overall dataset has been completed and needs to incorporate the target dataset segment contained in the managed target physical storage partition into the overall dataset, it can perform a partition mount to disable the independent management and maintenance of the data boundary information (i.e., the data region boundaries and / or time window) for a target dataset segment within the overall dataset. Partition mount incorporates the target dataset segment contained in the target physical storage partition into the overall dataset (i.e., the overall logical partition dataset), making it accessible to business applications along with the existing data in the overall logical partition dataset. It is important to note that the target segmentation rules used to describe the target physical storage partition should not conflict with the segmentation rules of the overall dataset (i.e., the overall logical partition dataset) to which it is mounted. For example, taking the X Province National Land Survey Land Classification Map Dataset as an example, after completing data updates, index optimization, or storage migration, the unmounted land classification map data subset for County Y in X Province can be mounted back into the X Province National Land Survey Land Classification Map Dataset for access by business applications. Since the partition mount operation only involves updating the metadata describing the entire logical partitioned dataset and its partitioned data subset (i.e., the target dataset segment), the impact on business applications accessing the entire dataset during this operation is negligible.
[0031] In an embodiment of the present invention, a method for organizing and storing massive data is provided, which is applied to a data management business interface at a partition granularity. The method comprises: unloading a target data set segment in a target physical storage partition from an overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition performs data addition, deletion, and modification operations independently from the overall data set composed of data set segments in other physical storage partitions. Except for the application system that calls the partition unloading method, other application systems cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; according to the application The target partition management function called by the system manages the target data set segments in the target physical storage partition to obtain the target physical storage partition that has been managed, wherein the target partition management function includes at least one of the following: clearing partitions, deleting partitions, migrating partitions, cloning partitions, splitting partitions, merging partitions, and maintaining partition indexes; verifying the target segmentation rules corresponding to the target physical storage partition that has been managed against the segmentation rule set of the overall data set to verify whether there is a conflict, and verifying the data structure of the target data set segments in the target physical storage partition that has been managed against the data structure of the overall data set to verify whether the data structures are consistent; after the verification is passed, the target data set segments in the target physical storage partition that has been managed are incorporated into the overall data set to which they belong for access by all application systems according to the partition mounting method called by the application system.From the above description, it can be seen that in the method of organizing and storing massive data of the present invention, the partition unloading method is adopted to perform data maintenance on the target data set segment in the target physical storage partition independently of the overall data set composed of the data set segments in other physical storage partitions. That is to say, the data management and processing operations occurring on the target physical storage partition can not affect the availability of the data of other physical storage partitions on the same overall data set to the application system. That is, when there is a problem with the data of the target physical storage partition and data maintenance is required, the data in other physical storage partitions can still be used, meeting the high-availability data requirements. In addition, when data maintenance is required, it is only necessary to maintain the target physical storage partition corresponding to the data, and there is no need to perform overall maintenance on the overall data set. Only local data needs to be updated and maintained, and there is no need to maintain the entire data set, which greatly improves the efficiency of data maintenance and meets the demand for easy-to-manage data. In addition, the data set segments in the physical storage partition are used as atomic units to ensure the consistency of data content. Data in incomplete physical storage partitions during the data processing period is invisible to business applications (i.e., application systems). That is, for massive overall data sets, the visibility of local data to the business can be dynamically controlled, meeting the needs of atomic application scenarios of super transactions. In other words, the method of organizing and storing massive data of the present invention can meet the needs of high availability, easy management and atomic application scenarios of super transactions, alleviating the difficulty of traditional massive data organization and storage methods in fully meeting the needs of atomic application scenarios of high availability, easy management and super transactions.
[0032] The above content briefly introduces the method for organizing and storing massive data of the present invention. The specific contents involved are described in detail below.
[0033] In an optional embodiment of the present invention, a prerequisite for unloading a target data set segment in a target physical storage partition from an overall data set according to a partition unloading method called by an application system is that the overall data set is segmented and allocated physical storage partitions based on the segmentation configuration. The segmentation configuration and allocation of physical storage partitions are as follows: (1) When the application system calls the data management business interface, it determines a segmentation rule set for segmenting the entire data set, and segments the entire data set according to the segmentation rule set to obtain multiple data set segments; Specifically, the entire data set is segmented according to the manually selected intrinsic features of the entire data set, wherein the intrinsic features include: time features and spatial features.
[0034] During implementation, for an overall data set that may have more than 10 million records, the inherent characteristics of the overall data set are manually selected based on demand (e.g., continuous accumulation over time or regional coverage, selection by time period, administrative region, business classification, etc. If by time period, manual selection of characteristics by day, month, quarter, year; if by administrative region, selection by county, city, province, etc.), and then the interface of the present invention (data management business interface with partition granularity) determines the segmentation rule set for segmenting the overall data set. If by month, the segmentation rule set can be January 2024, February 2024, March 2024, etc. , April 2024, etc., and the entire dataset is segmented by month according to the segmentation rule set. As shown above, the data of January 2024 is used as a dataset segment, the data of February 2024 is used as a dataset segment, and the data of March 2024 is used as a dataset segment. In this way, multiple dataset segments are obtained, and each dataset segment corresponds to a segmentation rule. That is, based on the segmentation rule set, the data in the massive dataset (i.e., the entire dataset) can be aggregated into multiple dataset segments, and the records (i.e., data) in the same dataset segment have good proximity in time, spatial distribution, or other classification characteristics. Like ordinary data sets, logically partitioned data sets (i.e., data set segments) are known to related business applications, but business applications may not know the partitioning logic behind them (i.e., the specific segmentation rules). Through partition unloading (such as the content in step S102 above), business applications can avoid accessing local data that has failed, requires offline maintenance (such as migration to storage media with different performance, etc.), cannot meet the needs of super-transaction atomicity scenarios (such as statistical analysis during database warehousing), or has replaced content as a whole, thereby maintaining the sustainability of business application operations; for those business applications that need to avoid local data to run, since the volume of this local data is much smaller, the time required for fault recovery, offline maintenance, etc. on this local data is much shorter, which means that the business application interruption time caused by fault recovery or maintenance of local data is much shorter, and the recovery time is generally predictable. Taking the land classification patch dataset of Province X as an example, this dataset contains approximately 17 million land classification patches. The data can be segmented based on the county-level administrative districts where the patches are located, forming land classification patch data subsets (i.e., multiple dataset segments) corresponding one-to-one to the 103 county-level administrative districts in Province X. Each data subset contains only the land classification patch data for its respective administrative district.This data segmentation and aggregation strategy avoids random patch data storage layout in storage. In map application scenarios, spatially neighboring patches are visualized in a connected manner, which means that the data of these patches need to be read at the same time. The physical reading efficiency of data is greatly improved, and a large number of data screening operations are avoided. The user's direct experience is the improvement of patch map visualization performance; this data segmentation strategy also caters to multi-level authorization business application scenarios at the provincial, municipal, and county levels, that is, county-level applications only need to be authorized to access their own data set segments, municipal-level applications only need to be authorized to access several county-level data set segments under their jurisdiction, and provincial-level applications are only authorized to access all county-level data set segments; in addition, data set segments can be maintained independently of the overall data set after partitioning and unloading. In the process of maintaining one or more county-level land classification patch data (such as warehousing, spatial index creation, and migration from traditional disks to SSD high-performance storage), the operation of various natural resources business applications will generally not be interrupted, and the availability is good.
[0035] (2) Create multiple physical storage partitions according to the segmentation rule set, and establish a mapping relationship between each data set segment and the corresponding physical storage partition, so that each data set segment is stored in the corresponding physical storage partition; Specifically, a physical storage partition corresponding to each segmentation rule is created according to user customization, and then a mapping relationship between each data set segment and the corresponding physical storage partition is established according to each segmentation rule.
[0036] During implementation, users can customize the storage location for each dataset segment based on the popularity of each dataset segment, fault tolerance, etc., and then create a physical storage partition corresponding to each dataset segment in the corresponding storage location based on the user's customization. The created physical storage partition also corresponds to the corresponding segmentation rule. Then, based on the segmentation rule, a mapping relationship between each dataset segment and the corresponding physical storage partition can be established. After the mapping relationship is established, each dataset segment is stored in the corresponding physical storage partition according to the database's own data storage positioning function.
[0037] When creating each physical storage partition, physical storage (e.g., with different technologies and disk redundancy) is allocated with the minimum capacity (e.g., a database segment) that matches the data storage requirements of the physical storage partition. The newly created physical storage partition becomes part of the entire logical partition dataset as an empty data subset. Because only the minimum capacity physical storage partition is allocated, the impact of creating a new physical storage partition on business applications' simultaneous access to the entire dataset is negligible.
[0038] In an optional embodiment of the present invention, a partition unloading method carries data isolation boundary information and unloads a target data set segment in a target physical storage partition from the overall data set according to the partition unloading method called by the application system, specifically comprising the following steps: 1) Determine the target physical storage partition corresponding to the data isolation boundary information; 2) Unload the target dataset segment in the target physical storage partition from the overall dataset.
[0039] In an optional embodiment of the present invention, clearing a partition includes logical clearing and physical clearing. Logical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, but does not release the physical storage space occupied by the target physical storage partition. Physical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, and release the other physical storage space occupied by the target physical storage partition except for retaining the necessary physical storage space (such as a database block or page). The above-mentioned index partition is a physical storage area that stores only the index information of the target data set segment, which is a sub-area of the overall data set index storage area. The index of the data set segment is referred to as the partition index. Deleting a partition not only deletes the target physical storage partition, but also integrates the segmentation rules of the deleted target physical storage partition into the data segmentation rule set of the entire data set. Migrating a partition involves copying a target data set segment in a target physical storage partition to be migrated to a first new physical storage partition, deleting the target physical storage partition to be migrated, updating a mapping relationship between the target data set segment and the corresponding first new physical storage partition, and rebuilding a partition index corresponding to the target data set segment. The cloning partition is to copy the target dataset segment in the cloned target physical storage partition to the second new physical storage partition, and create a corresponding partition index for the target dataset segment in the second new physical storage partition according to the index corresponding to the cloned target dataset segment and its index type and parameters; Splitting the partitions involves generating multiple sub-segmentation rules based on the segmentation rules corresponding to the split data set segments, allocating a respective physical storage partition to each sub-segmentation rule, and obtaining multiple target physical storage sub-partitions. The data set segments in the split target physical storage partitions are migrated to the corresponding target physical storage sub-partitions according to the sub-segmentation rules, and creating corresponding partition indexes for the multiple target data set sub-segments formed by the migration according to the indexes corresponding to the split target data set segments and their index types and parameters. Merging partitions involves generating a merge rule based on the target segmentation rules corresponding to the multiple target dataset segments being merged, allocating a corresponding physical storage partition to the merge rule, obtaining a merged target physical storage partition, migrating the multiple target dataset segments being merged to the merged target physical storage partition to form a merged dataset segment, releasing the physical storage of the merged target physical storage partition, and creating a corresponding partition index on the merged dataset segment based on the index status of the overall dataset. Maintaining partition indexes rebuilds and maintains partition indexes on target dataset segments in the unloaded target physical storage partition.
[0040] The following is a detailed introduction to each of the above management functions: Clearing a partition includes: logical clearing and physical clearing. Logical clearing is to delete the data content of the target physical storage partition and its derivative data such as indexes without releasing the physical storage space occupied by the target physical storage partition. Logical clearing is achieved by simply marking the data and indexes of the target physical storage partition as invalid by the interface, without calling the operating system to reclaim the storage space occupied by the target physical storage partition. Physical clearing is to delete the data content of the target physical storage partition and its derivative data such as indexes, and call the operating system to release other physical storage space occupied by the target physical storage partition except for retaining the necessary physical storage space (such as a database block or page).
[0041] To delete a partition, after calling the physical partition clearing operation, the interface program determines whether to merge the segmentation rules of the target dataset segment to be deleted with the relevant segmentation rules in the overall dataset by adjusting the segmentation rules of the overall dataset based on the parameters provided by the application system when calling the partition delete operation. For example, when deleting the January 2024 partition, you can specify whether to modify the segmentation rules of the existing February 2024 partition in the overall dataset's segmentation rules set to accommodate the data records belonging to January 2024. Alternatively, you can directly delete the segmentation rules that determine the January 2024 dataset segment from the overall dataset's segmentation rules set.
[0042] Migrate partition. This interface creates a new physical storage partition in the target storage area, updates the mapping relationship between the target dataset segment and the corresponding new physical storage partition, copies the data content of the migrated target physical storage partition to the new physical storage partition, and deletes the migrated target physical storage partition. The storage areas before and after migration can be different areas of the same physical storage or belong to different physical storages. When migrating a partition, the partition index can be rebuilt synchronously or asynchronously.
[0043] Clone partition: While retaining the data content of the cloned target physical storage partition, this interface creates a new physical storage partition in the target storage area. Its data content is copied from the cloned target physical storage partition. You can also choose whether to create corresponding indexes for the new physical storage partition according to the index types and index columns of the entire dataset.
[0044] Split partitions to form multiple sub-segmentation rules based on the target segmentation rules of the target physical storage partition being split. The multiple target dataset sub-segments determined by these sub-segmentation rules cannot overlap with each other, and the union of the multiple target dataset sub-segments stipulated by these sub-segmentation rules is equivalent to the target dataset segment of the target physical storage partition being split. Create a target physical storage sub-partition in an unloaded state according to each sub-segmentation rule, and migrate its data content from the split target physical storage partition to the corresponding target physical storage sub-partition according to the sub-segmentation rules. You can also choose whether to create corresponding partition indexes for the multiple target dataset sub-segments formed by the migration according to the index type and index columns on the overall dataset. For massive datasets that are partitioned based on administrative districts, the split partition operation can be used to cope with the scenario of newly established administrative districts. This interface keeps the multiple split partitions belonging to the overall dataset by updating the segmentation rule content of the overall dataset.
[0045] Merging partitions is the opposite of splitting partitions. For multiple target physical storage partitions that are already in the unloaded state, the target segmentation rules of these partitions are used to form a merge rule. The data set determined by the merge rule is the union of the data sets of the multiple target physical storage partitions being merged. Based on the new merge rule, a merged target physical storage partition in the unloaded state is created. Data is migrated from each merged target physical storage partition to the merged target physical storage partition. It is also possible to choose whether to create corresponding partition indexes for the merged dataset segments according to the index types and index columns of the overall dataset. This interface updates the segmentation rule content of the overall dataset so that the merged target dataset segments formed by the merge belong to the overall dataset. For massive datasets that are partitioned based on administrative districts, the merge partition operation can be used to handle scenarios where administrative districts are merged.
[0046] Maintain partition indexes. For the unloaded target physical storage partition, you can rebuild the partition index based on the data content of the target physical storage partition in accordance with the index type and index column specifications on the overall data set for the purpose of index optimization. You can also asynchronously create partition indexes for new partitions formed by partition management and maintenance operations in accordance with the indexes on the overall data set.
[0047] In an optional embodiment of the present invention, a partition mounting method carries data boundary information. According to the partition mounting method called by the application system, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs for access by all application systems. Specifically, the method includes the following steps: (1) Determine the target physical storage partition corresponding to the data boundary information; (2) The target data set segments in the managed target physical storage partitions are incorporated into the overall data set to which they belong for access by all application systems.
[0048] Based on existing database partition management technology and large table partitioning operations, this invention develops a partition-granular data management interface, providing a highly available and manageable method for organizing and storing massive data in a detailed manner based on the data's inherent characteristics. This method involves segmenting massive data sets into multiple segments based on the data's time period, spatial distribution, and primary access and utilization patterns. Based on the differences in access performance, security strength, and management requirements resulting from the time period and spatial distribution of the data segments, a storage layout is implemented that matches the storage medium and reliability requirements. By separating overall massive data business applications from localized data management workloads, this method achieves high availability, manageability, and super-transactional atomicity based on the inherent characteristics of massive data while maintaining high performance and high concurrency.
[0049] It should be noted that in addition to the above interface methods, the present invention also provides auxiliary interfaces such as GetDatasetPartitionInfo, SetDatasetPartitionInfo, GetPartitionInfo, SetPartitionInfo, and ValidatePartitonToDataset. Among them, GetDatasetPartitionInfo and SetDatasetPartitionInfo are used to query or configure the segmentation rules of the overall logical partitioned dataset, the partitions it contains, and the information of each partition in the uninstalled state; GetPartitionInfo and SetPartitionInfo are used to query or configure the data segmentation rules, partition status, and related index information of a specific partition.
[0050] The method for organizing and storing massive amounts of data of the present invention has the following advantages: (1) Data management and processing operations on the target physical storage partition must not affect the availability of data in other physical storage partitions on the same overall logical partition data set (i.e., the overall data set) to the business; (2) Taking the physical storage partition as the atomic unit, the consistency of data content is guaranteed. The data of the incomplete physical storage partition during the data processing period is invisible to the business application; (3) For massive data sets, the visibility of local data to the business can be dynamically controlled; (4) Maintain indexes at the partition level. Index maintenance on a specific physical storage partition will not affect the validity of indexes on other partitions. (5) Quickly isolate the physical storage partition of the problem data without affecting the availability of the data content of other physical storage partitions to the business; (6) Implement differentiated storage allocation and on-demand adjustments for massive amounts of data, such as placing data in some physical storage partitions that require high-profile access on high-performance storage; (7) Split the long-term, continuous data management and maintenance work that affects the entire system into multiple small units of local, short-term, independent and concurrent data management and maintenance work; (8) Achieve data maintenance in narrow time windows with predictable duration; (9) Provide configurable data segmentation rules that reflect the inherent characteristics of natural resource industry data; (10) It provides a mechanism for users to customize data organization, layout and storage based on the intrinsic characteristics of data objects. Users can implement differentiated management and maintenance of local data of massive scale sets at different life cycles. (11) It improves the availability of massive data sets in non-cluster configurations and has certain advantages in low cost and easy maintenance compared to cluster deployment.
[0051] Taking the storage of land survey data and the application of land statistics in Province X as an example, in accordance with national regulations on conducting land surveys at the county level, Province X built a provincial land survey database based on the land survey data of its 103 counties and cities. The core of this database is approximately 17 million land type map data, a large amount of data. Relying on the province's established and operational land space basic information platform, the land survey data supports land statistics, farmland protection, land approval, and other businesses. In order to ensure that the land survey data database construction process in Province X does not interrupt business operations as much as possible and that the stored data can serve the above-mentioned applications as soon as possible, the present invention can be used to carry out its construction work: 1. The provincial land class map data is segmented by county-level administrative district codes. At this point, the application functions on it can run, but there is no output because the overall land class map dataset is empty. 2. Now we need to store the land classification map data of County Y, Province X in the database. To do this, we create a physical storage partition. All land classification map data with the administrative district code of County Y will be placed in this physical storage partition. 3. To ensure that land statistics can only be performed when the land classification map data of County Y is complete, the land classification map data segments of County Y are unloaded from the overall data set of the provincial land classification map, so that empty or incomplete land classification maps of County Y cannot be accessed by land statistics. 4. Use standard SQL to load data into the target physical storage partition of the Y County land classification map in the unloaded state. The data entry may be controlled by multiple database transactions, and the specific number often depends on the considerations of the technician who calls this interface; 5. After the storage is completed, the target physical storage partition of the Y County land class map data segment will be mounted to the overall data set of the provincial land class map. The land statistics business can calculate the land area according to multi-level land classes based on the complete Y County land class map data.
[0052] Without the partition unloading and mounting strategy, land statistics services could access incomplete land classification data for County Y during the storage process. The resulting output would often be meaningless and misleading to statisticians. Using a traditional database transaction to address this issue would be limited by the large data volume and database log configuration. By adopting the partition unloading and mounting strategy, we can achieve data storage atomicity beyond the database transaction level, as demonstrated in this example.
[0053] Example 2: An embodiment of the present invention further provides a device for organizing and storing massive data. The device for organizing and storing massive data is mainly used to execute the method for organizing and storing massive data provided in the first embodiment of the present invention. The following is a detailed introduction to the device for organizing and storing massive data provided in the embodiment of the present invention.
[0054] Figure 2 FIG. 1 is a schematic diagram of a device for organizing and storing massive data according to an embodiment of the present invention. Figure 2 As shown, the data management service interface applied to the partition granularity mainly includes: an unloading unit 10, a management unit 20, a verification unit 30, and a mounting unit 40, wherein: an unloading unit, configured to unload a target data set segment in a target physical storage partition from the overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition can independently perform data addition, deletion, and modification operations on the overall data set composed of data set segments in other physical storage partitions, and other application systems other than the application system that called the partition unloading method cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed, which is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; a management unit, configured to manage target data set segments in a target physical storage partition according to a target partition management function called by an application system, and obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; a verification unit, configured to verify the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall data set to verify whether there is a conflict, and to verify the data structure of the target data set segment in the managed target physical storage partition against the data structure of the overall data set to verify whether the data structures are consistent; The mounting unit is used to incorporate the target data set segment in the managed target physical storage partition into the overall data set to which it belongs for access by all application systems according to the partition mounting method called by the application system after verification.
[0055] In an embodiment of the present invention, a device for organizing and storing massive data is provided, which is applied to a data management business interface at a partition granularity. The device comprises: unloading a target data set segment in a target physical storage partition from an overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition performs data addition, deletion, and modification operations independently of the overall data set composed of data set segments in other physical storage partitions. Except for the application system that calls the partition unloading method, other application systems cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; according to the application The target partition management function called by the system manages the target data set segments in the target physical storage partition to obtain the target physical storage partition that has been managed, wherein the target partition management function includes at least one of the following: clearing partitions, deleting partitions, migrating partitions, cloning partitions, splitting partitions, merging partitions, and maintaining partition indexes; verifying the target segmentation rules corresponding to the target physical storage partition that has been managed against the segmentation rule set of the overall data set to verify whether there is a conflict, and verifying the data structure of the target data set segments in the target physical storage partition that has been managed against the data structure of the overall data set to verify whether the data structures are consistent; after the verification is passed, the target data set segments in the target physical storage partition that has been managed are incorporated into the overall data set to which they belong for access by all application systems according to the partition mounting method called by the application system.From the above description, it can be seen that in the device for organizing and storing massive data of the present invention, a partition unloading method is adopted to perform data maintenance on the target data set segment in the target physical storage partition independently of the overall data set composed of the data set segments in other physical storage partitions. That is to say, the data management and processing operations occurring on the target physical storage partition may not affect the availability of the data of other physical storage partitions on the same overall data set to the application system. That is, when there is a problem with the data of the target physical storage partition and data maintenance is required, the data in other physical storage partitions can still be used, meeting the high-availability data requirements. In addition, when data maintenance is required, it is only necessary to maintain the target physical storage partition corresponding to the data, and there is no need to perform overall maintenance on the overall data set. Only local data needs to be updated and maintained, and there is no need to maintain the entire data set, which greatly improves the efficiency of data maintenance and meets the demand for easy-to-manage data. In addition, the data set segments in the physical storage partition are used as atomic units to ensure the consistency of data content. Data in incomplete physical storage partitions during the data processing period is invisible to business applications (i.e., application systems). That is, for massive overall data sets, the visibility of local data to the business can be dynamically controlled, meeting the needs of atomic application scenarios of super transactions. In other words, the method of organizing and storing massive data of the present invention can meet the needs of high availability, easy management and atomic application scenarios of super transactions, alleviating the difficulty of traditional massive data organization and storage methods in fully meeting the needs of atomic application scenarios of high availability, easy management and super transactions.
[0056] Optionally, the device is also used to: when the application system calls the data management business interface, determine a segmentation rule set for segmenting the entire data set, and segment the entire data set according to the segmentation rule set to obtain multiple data set segments; create multiple physical storage partitions according to the segmentation rule set, and establish a mapping relationship between each data set segment and the corresponding physical storage partition, so that each data set segment is stored in the corresponding physical storage partition.
[0057] Optionally, the device is further configured to determine a segmentation rule set for segmenting the entire data set based on manually selected intrinsic features of the entire data set, wherein the intrinsic features include: time features and spatial features.
[0058] Optionally, the device is further configured to: create a physical storage partition corresponding to each segmentation rule according to user customization, and then establish a mapping relationship between each data set segment and the corresponding physical storage partition according to each segmentation rule.
[0059] Optionally, the partition unloading method carries data isolation boundary information, and the unloading unit is further used to: determine a target physical storage partition corresponding to the data isolation boundary information; and unload the target data set segment in the target physical storage partition from the overall data set.
[0060] Optionally, clearing a partition includes: logical clearing and physical clearing. Logical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, but not release the physical storage space occupied by the target physical storage partition; physical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, and release the other physical storage space occupied by the target physical storage partition except for retaining the necessary physical storage space; deleting a partition is to merge the segmentation rules of the deleted target physical storage partition into the data segmentation rule set of the overall data set while deleting the target physical storage partition; migrating a partition is to copy the target data set segment in the migrated target physical storage partition to the first new physical storage partition, delete the migrated target physical storage partition, update the mapping relationship between the target data set segment and the corresponding first new physical storage partition, and rebuild the partition index corresponding to the target data set segment; cloning a partition is to copy the target data set segment in the cloned target physical storage partition to the second new physical storage partition, and map the second new physical storage partition to the index corresponding to the cloned target data set segment and its index type and parameters. Create corresponding partition indexes for target data set segments in physical storage partitions; split partitions to generate multiple sub-segmentation rules according to the segmentation rules corresponding to the split data set segments, assign respective physical storage partitions to each sub-segmentation rule, obtain multiple target physical storage sub-partitions, migrate the data set segments from the split target physical storage partitions to the corresponding target physical storage sub-partitions according to the sub-segmentation rules, and create corresponding partition indexes for the multiple target data set sub-segments formed by the migration according to the indexes corresponding to the split target data set segments and their index types and parameters; merge partitions to generate merge rules according to the target segmentation rules corresponding to the multiple target data set segments to be merged, assign corresponding physical storage partitions to the merge rules, obtain merged target physical storage partitions, form merged data set segments after migrating the multiple target data set segments to be merged, release the physical storage of the merged target physical storage partitions, and create corresponding partition indexes on the merged data set segments according to the index situation of the overall data set; maintain partition indexes to rebuild and maintain the partition indexes on the target data set segments in the unloaded target physical storage partitions.
[0061] Optionally, the partition mounting method carries data boundary information, and the mounting unit is further used to: determine the target physical storage partition that has been managed corresponding to the data boundary information; and incorporate the target data set segment in the target physical storage partition that has been managed into the overall data set to which it belongs for access by all application systems.
[0062] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.
[0063] like Figure 3 As shown, an electronic device 600 provided in an embodiment of the present application includes: a processor 601, a memory 602 and a bus, wherein the memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device is running, the processor 601 communicates with the memory 602 through the bus, and the processor 601 executes the machine-readable instructions to perform the steps of the method for organizing and storing massive data as described above.
[0064] Specifically, the above-mentioned memory 602 and processor 601 can be general-purpose memories and processors, which are not specifically limited here. When the processor 601 runs the computer program stored in the memory 602, it can execute the above-mentioned method of organizing and storing massive data.
[0065] The processor 601 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 601 or by instructions in the form of software. The above-mentioned processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 602, and processor 601 reads the information in memory 602 and performs the steps of the above method in conjunction with its hardware.
[0066] Corresponding to the above-mentioned method for organizing and storing massive data, an embodiment of the present application also provides a computer-readable storage medium, which stores machine-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the steps of the above-mentioned method for organizing and storing massive data.
[0067] The device for organizing and storing massive data provided in the embodiments of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiments of the present application, its implementation principle and the technical effect produced are the same as those in the aforementioned method embodiments. For the sake of brief description, where the device embodiment is not mentioned, reference can be made to the corresponding content in the aforementioned method embodiments. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0068] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0069] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0070] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0071] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0072] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method for organizing and storing massive data described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0073] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0074] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. However, these modifications, changes, or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for organizing and storing massive data, characterized in that: The method for applying a data management service interface at a partition granularity includes: Unloading a target data set segment in a target physical storage partition from the overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition can independently perform data addition, deletion, and modification operations on the overall data set composed of data set segments in other physical storage partitions. Except for the application system that calls the partition unloading method, other application systems cannot access the target data set segment in the target physical storage partition. When performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed. The target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented. Managing the target data set segments in the target physical storage partition according to the target partition management function called by the application system to obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; Verifying the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall data set to verify whether there is a conflict, and verifying the data structure of the target data set segment in the managed target physical storage partition against the data structure of the overall data set to verify whether the data structures are consistent; After the verification is passed, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs according to the partition mounting method called by the application system for access by all application systems.
2. The method according to claim 1, characterized in that Before unloading the target data set segment in the target physical storage partition from the overall data set according to the partition unloading method called by the application system, pre-preparation is performed, the method further comprising: When the application system calls the data management service interface, it determines a segmentation rule set for segmenting the entire data set, and segments the entire data set according to the segmentation rule set to obtain a plurality of data set segments; A plurality of physical storage partitions are created according to the segmentation rule set, and a mapping relationship is established between each of the data set segments and the corresponding physical storage partition, so that each of the data set segments is stored in the corresponding physical storage partition.
3. The method according to claim 2, characterized in that Determining a segmentation rule set for segmenting the entire data set, including: A segmentation rule set for segmenting the entire data set is determined according to manually selected intrinsic features of the entire data set, wherein the intrinsic features include: time features and spatial features.
4. The method according to claim 2, characterized in that Creating a plurality of physical storage partitions according to the segmentation rule set and establishing a mapping relationship between each of the data set segments and the corresponding physical storage partition, including: A physical storage partition corresponding to each segmentation rule is created according to user customization, and then a mapping relationship between each data set segment and the corresponding physical storage partition is established according to each segmentation rule.
5. The method according to claim 1, characterized in that The partition unloading method carries data isolation boundary information and unloads the target data set segment in the target physical storage partition from the overall data set according to the partition unloading method called by the application system, including: Determining a target physical storage partition corresponding to the data isolation boundary information; Unloading the target data set segment in the target physical storage partition from the overall data set.
6. The method according to claim 1, wherein: The partition clearing includes logical clearing and physical clearing. The logical clearing is to delete the target data set segment and the index partition corresponding to the target data set segment in the target physical storage partition, but not release the physical storage space occupied by the target physical storage partition. The physical clearing is to delete the target data set segment and the index partition corresponding to the target physical storage partition, and release the other physical storage space occupied by the target physical storage partition except for retaining the necessary physical storage space. Deleting the partition while deleting the target physical storage partition merges the segmentation rules of the deleted target physical storage partition into the data segmentation rule set of the overall data set; The migration partition is to copy the target data set segment in the target physical storage partition to be migrated to the first new physical storage partition, delete the migrated target physical storage partition, update the mapping relationship between the target data set segment and the corresponding first new physical storage partition, and rebuild the partition index corresponding to the target data set segment; The clone partition is to copy the target dataset segment in the cloned target physical storage partition to the second new physical storage partition, and create a corresponding partition index for the target dataset segment in the second new physical storage partition according to the index corresponding to the cloned target dataset segment and its index type and parameters; The partition splitting step comprises generating a plurality of sub-segmentation rules according to the segmentation rules corresponding to the split data set segments, assigning a respective physical storage partition to each of the sub-segmentation rules to obtain a plurality of target physical storage sub-partitions, migrating the data set segments from the split target physical storage partitions to the corresponding target physical storage sub-partitions according to the sub-segmentation rules, and creating corresponding partition indexes for the plurality of target data set sub-segments formed by the migration according to the indexes corresponding to the split target data set segments and their index types and parameters; The merging partitions is to generate a merging rule according to target segmentation rules corresponding to the multiple target dataset segments to be merged, allocate a corresponding physical storage partition to the merging rule, obtain a merged target physical storage partition, migrate the multiple target dataset segments to the merged target physical storage partition to form a merged dataset segment, release the physical storage of the merged target physical storage partition, and create a corresponding partition index on the merged dataset segment according to the index status of the overall dataset; The partition index maintenance is to rebuild and maintain the partition index on the target data set segment in the unloaded target physical storage partition.
7. The method according to claim 1, characterized in that The partition mounting method carries data boundary information, and according to the partition mounting method called by the application system, the target data set segment in the managed target physical storage partition is incorporated into the overall data set to which it belongs for access by all application systems, including: Determine a target physical storage partition corresponding to the data boundary information and having completed management; The target data set segments in the target physical storage partitions whose management has been completed are incorporated into the overall data set to which they belong for access by all application systems.
8. A device for organizing and storing massive amounts of data, characterized in that: A data management service interface applied to partition granularity, the device comprising: an unloading unit, configured to unload a target data set segment in a target physical storage partition from the overall data set according to a partition unloading method called by an application system, so that the target data set segment in the target physical storage partition can independently perform data addition, deletion, and modification operations on the overall data set composed of data set segments in other physical storage partitions, and other application systems cannot access the target data set segment in the target physical storage partition except the application system that calls the partition unloading method, wherein when performing data addition, deletion, and modification operations on the target data set segment in the target physical storage partition, the corresponding target segmentation rule is followed, and the target segmentation rule is the segmentation rule corresponding to the target data set segment when the overall data set is segmented; a management unit, configured to manage the target data set segments in the target physical storage partition according to the target partition management function called by the application system, and obtain a managed target physical storage partition, wherein the target partition management function includes at least one of the following: clearing a partition, deleting a partition, migrating a partition, cloning a partition, splitting a partition, merging partitions, and maintaining a partition index; a verification unit, configured to verify the target segmentation rules corresponding to the managed target physical storage partition against the segmentation rule set of the overall data set to verify whether there is a conflict, and to verify the data structure of the target data set segment in the managed target physical storage partition against the data structure of the overall data set to verify whether the data structures are consistent; The mounting unit is used to incorporate the target data set segment in the managed target physical storage partition into the overall data set to which it belongs for access by all application systems according to the partition mounting method called by the application system after verification.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Hard disk data processing method and device and electronic equipment
CN115016739A
Data validation of data migrated from a source database to a target database
US10963435B1