Storage space management method, device and equipment and computer readable storage medium

By automatically calculating and adjusting the number of storage units in the data storage space, the problem of data write latency in real-time big data processing scenarios is solved, achieving efficient data storage space management and system automation, and improving the stability and business continuity of the data processing system.

CN121541829APending Publication Date: 2026-02-17CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511692582.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In real-time big data processing scenarios, existing technical architectures suffer from data write delays when data volumes surge suddenly, and the lack of automated processing methods affects the efficiency of real-time data processing for businesses.

Method used

By collecting operational metrics from the data processing system, the target number of storage units in the data storage space is automatically calculated, and dynamic adjustments are made when preset conditions are met, including data processing rate and write latency ratio judgment. The state machine and routing service are used to achieve storage unit number configuration and data reallocation without manual intervention.

Benefits of technology

It achieves dynamic matching between data processing system load and storage capacity, improves the efficiency and automation of data storage space management, reduces manual intervention, and ensures the stability and business continuity of the data processing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541829A_ABST
    Figure CN121541829A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed computing, in particular to a storage space management method, device and equipment and a computer readable storage medium. Whether the number of the storage units in the current data storage space needs to be adjusted or not is judged based on the operation indexes of the data processing system, the target number is automatically calculated and adjustment is rapidly completed when adjustment is needed, manual intervention is not needed in the whole process, dynamic matching of the load and the storage capacity of the data processing system is achieved, and the data processing efficiency is improved. And the management efficiency and the automation degree of the data storage space are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of distributed computing technology, and in particular to a storage space management method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] In current big data real-time processing scenarios, the mainstream technical architecture typically uses the Apache Flink engine to read source data from Kakfa and write the processed data into the Paimon data lake. In practical applications, this architecture often faces sudden surges in data volume, leading to significant latency issues when writing data to the Paimon data lake, severely impacting the business's needs for real-time data processing and usage.

[0003] The industry currently uses manual intervention to address the aforementioned write latency issue, which involves manually adjusting parameters in the architecture after stopping data processing tasks. This approach suffers from low automation and inefficiency. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a storage space management method, apparatus, device, and computer-readable storage medium to improve the efficiency and automation of data storage space management.

[0005] In a first aspect, embodiments of this disclosure provide a storage space management method, including: The operational metrics of the data acquisition and processing system include the data processing rate of the data source in the data processing system and the write latency of the data storage space. When the operating indicators meet the preset conditions, calculate the target number of storage units in the data storage space; Submit the target number to the data storage space so that the data storage space configures the number of storage units to the target number.

[0006] In some embodiments, when the operating indicators meet preset conditions, calculating the target number of storage units in the data storage space includes: When the ratio of the data processing rate to the preset processing rate is greater than a first preset ratio, the target number of storage units in the data storage space is calculated; or, When the ratio of write latency to preset latency is greater than the second preset ratio, the target number of storage units in the data storage space is calculated.

[0007] In some embodiments, calculating the target number of storage units in the data storage space includes: Get the current number of storage units; Double the current number is determined as the first candidate number; The number of storage units corresponding to the available resource information of the data processing system is the second candidate number; Obtain the maximum preset number of storage units to get the third candidate number; The minimum value among the first, second, and third candidate numbers is determined as the target number.

[0008] In some embodiments, submitting a target quantity to a data storage space to cause the data storage space to configure the number of storage units to the target quantity includes: Create an information record in a preset information storage location and configure the state machine corresponding to the information record to the first state; wherein, the storage key value of the information record is the timestamp when the information record is created, and the storage value of the information record is the metadata to be submitted, the metadata to be submitted includes the number of storage units, and the number of storage units is the target number; In response to the state machine being in the first state, a first statement is created based on the information record. The first state indicates that there is an information record to be processed in the preset information storage location. The first statement is submitted to the data storage space through the routing service, so that the data storage space can obtain the metadata to be submitted based on the first statement and change the number of storage units to the target number.

[0009] In some embodiments, the method further includes: In response to the state machine being configured to the second state, and the metadata to be submitted and the metadata in the data storage space being verified, the data to be allocated in the data storage space is allocated to the target number of storage units. The second state indicates that the number of storage units has been successfully changed to the target number. Configure the state machine to the third state, which indicates that the data in the data storage space has been redistributed.

[0010] In some embodiments, the information record includes the task path of the data processing task in the data processing system at the current moment; the method further includes: In response to the state machine being configured to the third state, the data processing task is resumed based on the task path.

[0011] In some embodiments, the method further includes: In response to a successful data processing task recovery, configure the state machine to the fourth state; or, In response to a data processing task recovery failure, the version of the metadata in the data storage space is rolled back to the previous version, and the state machine is configured to the fifth state.

[0012] Secondly, embodiments of this disclosure provide a storage space management device, comprising: The acquisition module is used to collect the operating indicators of the data processing system, including the data processing rate of the data source in the data processing system and the write latency of the data storage space. The calculation module is used to calculate the target number of storage units in the data storage space when the operation indicators meet the preset conditions. The submission module is used to submit the target quantity to the data storage space, so that the data storage space configures the number of storage units to the target quantity.

[0013] Thirdly, embodiments of this disclosure provide an electronic device, including: Memory; Processor; and Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.

[0014] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described in the first aspect.

[0015] Fifthly, embodiments of this disclosure also provide a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the storage space management method described above.

[0016] The storage space management method, apparatus, device, and computer-readable storage medium provided in this disclosure determine whether the number of storage units in the current data storage space needs to be adjusted based on the operating indicators of the data processing system, and automatically calculate the target number and quickly complete the adjustment when adjustment is required. The entire process does not require manual intervention, realizes dynamic matching between the load of the data processing system and the storage capacity, and improves the efficiency and automation of data storage space management. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1A flowchart of a storage space management method provided in this embodiment of the disclosure; Figure 2 This is a schematic diagram of the state changes of a state machine provided in an embodiment of this disclosure; Figure 3 A flowchart of a storage space management method provided in another embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of a storage space management device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0020] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0021] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0022] This disclosure provides a storage space management method, which will be described below with reference to specific embodiments.

[0023] Figure 1 A flowchart illustrating a storage space management method provided in an embodiment of this disclosure. Figure 1 As shown, the method includes the following steps: S101, Operating indicators of the data acquisition and processing system.

[0024] Among these, the operational metrics include the data processing rate of the data source in the data processing system and the write latency of the data storage space.

[0025] A data processing system refers to a distributed system composed of components such as data acquisition, computation, storage, and scheduling. In this embodiment, it specifically includes the Flink engine, the data source Kakfa, and the data storage space Paimon Data Lake.

[0026] Operational metrics refer to quantitative data that reflect the system's operational status and performance level. Data processing rate refers to the speed at which the data source outputs data to the computing system, specifically the number of records consumed per second by a Kafka consumer group, reflecting the real-time flow of input data to the data processing system. Data storage space write latency is specifically the time difference between when the Paimon data lake receives a write request and when the write is confirmed to be complete; this metric reflects the real-time performance status of the storage layer in processing data.

[0027] S102. When the operating indicators meet the preset conditions, calculate the target number of storage units in the data storage space.

[0028] Preset conditions refer to the threshold standards that the system sets in advance to determine whether the number of storage units needs to be adjusted. They are used to trigger the adjustment of the number of storage units when the storage performance cannot meet the data processing requirements.

[0029] A storage unit refers to the smallest logical unit in a data storage space used for data partitioning and storage. Specifically, it refers to a bucket in the Paimon data lake. The number of buckets determines the data storage performance of the Paimon data lake. The more buckets there are, the more write tasks can be processed simultaneously, and the lower the write latency.

[0030] The target number refers to the optimal number of storage units that should be adjusted to.

[0031] S103. Submit the target quantity to the data storage space so that the data storage space configures the number of storage units to the target quantity.

[0032] After the calculated target quantity is passed to the data storage space, the data storage space updates the metadata according to the target quantity and synchronously adjusts the data sharding strategy based on the storage units (buckets) based on the target quantity.

[0033] This embodiment of the disclosure determines whether the number of storage units in the current data storage space needs to be adjusted based on the operating indicators of the data processing system, and automatically calculates the target number and quickly completes the adjustment when adjustment is needed. The entire process does not require manual intervention, realizes dynamic matching between the load of the data processing system and the storage capacity, and improves the efficiency and automation of data storage space management.

[0034] Based on the above embodiments, when the operating indicators meet the preset conditions, the target number of storage units in the data storage space is calculated, including: when the ratio of the data processing rate to the preset processing rate is greater than a first preset ratio, the target number of storage units in the data storage space is calculated; or, when the ratio of the write latency to the preset latency is greater than a second preset ratio, the target number of storage units in the data storage space is calculated.

[0035] The preset processing rate refers to the critical value of Kafka's data processing rate when the data processing system is resource-saturated. The preset latency refers to the critical value of write latency for the data storage space (Paimon data lake) when the data processing system is resource-saturated.

[0036] When the ratio of the real-time data processing rate to the preset processing rate is greater than the first preset ratio, or when the ratio of the real-time write latency to the preset latency is greater than the second preset ratio, it indicates that the current data processing system's operating indicators have exceeded the normal range, and the number of storage units needs to be adjusted in a timely manner.

[0037] Optionally, the first preset ratio is 1.8, and the second preset ratio is 2.0.

[0038] Accordingly, calculating the target number of storage units in the data storage space includes: obtaining the number of storage units at the current moment to obtain the current number of storage units; determining twice the current number as the first candidate number; determining the number of storage units corresponding to the available resource information of the data processing system as the second candidate number; obtaining the maximum preset number of storage units to obtain the third candidate number; and determining the minimum value among the first, second, and third candidate numbers as the target number.

[0039] Access the Paimon data lake's metadata storage node via the metadata query interface, read the currently active storage unit configuration information, determine the current number of storage units, and record it as the current number. Since there is a need to expand storage units to alleviate load pressure, double the current number is selected as the first candidate number.

[0040] Optionally, the number of storage units in the original state is consistent with the number of Kafka partitions to ensure a one-to-one mapping between consumption parallelism and storage buckets, avoid data skew, and ensure optimal task processing performance under limited resources.

[0041] Since writing data units requires computational resources such as memory and storage bandwidth from the data processing system, the number of expanded storage units needs to be controlled within the range supported by the available resources of the data processing system. Based on the available resource information of the data processing system, the maximum number of storage units that can be supported at present is determined, resulting in a second candidate number.

[0042] The maximum preset number of storage units is a parameter configured by maintenance personnel or set by the system default, used to limit the upper limit of the number of storage units. The maximum preset number of storage units is read from the configuration information of the data processing system to obtain the third candidate number.

[0043] Furthermore, the minimum value among the first candidate quantity, the second candidate quantity, and the third candidate quantity is selected as the target number of storage units.

[0044] This embodiment of the disclosure determines a first candidate quantity based on the current number of storage units and load requirements, a second candidate quantity based on the system resource status, and a third candidate quantity based on the maximum preset quantity. This enables dynamic calculation of the target quantity based on resource and system constraints, achieves dynamic balance between load and resources, and improves the flexibility of storage space management.

[0045] In some embodiments, submitting a target quantity to a data storage space so that the data storage space configures the number of storage units to the target quantity includes: creating an information record in a preset information storage location and configuring the state machine corresponding to the information record to a first state; wherein the storage key value of the information record is a timestamp when the information record is created, the storage value of the information record is metadata to be submitted, the metadata to be submitted includes the number of storage units, and the number of storage units is the target quantity; in response to the state machine being in the first state, creating a first statement based on the information record, the first state indicating that there is an information record to be processed in the preset information storage location; and submitting the first statement to the data storage space through a routing service so that the data storage space obtains the metadata to be submitted based on the first statement and changes the number of storage units to the target quantity.

[0046] The default information storage location is used for real-time data and status, such as a remote dictionary service (Redis). An information record refers to structured data containing metadata to be submitted, using the timestamp when the information record was created as the storage key and the metadata to be submitted as the storage value.

[0047] Simultaneously, process control is based on a state machine. The information record includes a state field, corresponding to the state of the state machine.

[0048] For example, the specific format of the information record is as follows: { "jobId": "flink-job-{hash}", "savepointPath": "hdfs: / / ns1 / flink / savepoints / ${jobId} / ${timestamp}", "targetBuckets": 10, "version": "v2", "status": "pending", "checksum": "sha256(metadata)", } Among them, jobId refers to the identifier of the current data processing task, savepointPath refers to the task path of the data processing task, targetBuckets is the target number of storage units, version is the version number of the metadata, status is the state of the state machine, and checksum represents data integrity verification.

[0049] Retrieve information records with the status field set to the first state from the preset information storage location and acquire a distributed lock. Create the first statement based on the information record.

[0050] The first statement is an idempotent Data Definition Language (DDL) statement containing a version number and a timestamp, including the version number, timestamp, and target number of storage units for the metadata.

[0051] For example, the first statement is as follows: ALTER TABLE `${table}` SET( 'bucket'='${targetBuckets}', 'rescale.version'='${version}', 'rescale.timestamp'='${timestamp}' ) WITH ( 'atomic'='true', 'async'='true' ) The first statement is sent to the metadata storage node of the data storage space via the router service. The metadata storage node is responsible for maintaining the configuration information of the data storage space. It obtains metadata such as the target number of storage units by parsing the parameters in the first statement, and makes configuration changes so that the number of storage units in the data storage space is configured to the target number.

[0052] This embodiment of the disclosure centrally manages information records by pre-setting information storage locations, and prevents concurrent processing conflicts by using state machines in conjunction with distributed locks, ensuring the reliability and stability of metadata submission. Furthermore, metadata submission does not require manual intervention, reducing operation and maintenance costs.

[0053] Figure 2 This is a schematic diagram illustrating the state changes of a state machine provided in an embodiment of this disclosure. Figure 2As shown, the first state (pending) indicates the initial state of the information record, meaning that the information record is waiting to be processed. The second state (committed) indicates that the number of storage units has been successfully changed to the target number. The third state (rolled) indicates that the data in the data storage space has been redistributed. The fourth state (compelted) indicates that the data processing task has been successfully resumed. The fifth state (failed) indicates that an error occurred in any of the states.

[0054] Based on the above embodiments, in response to the state machine being configured to the second state and the metadata to be submitted and the metadata in the data storage space being verified, the data to be allocated in the data storage space is allocated to the target number of storage units; the state machine is then configured to the third state.

[0055] After the metadata in the data storage space is updated, the state machine's state is written back, configuring it to the second state. The state machine being configured to the second state means that the status field of the information record in Redis is configured as committed.

[0056] The metadata to be submitted corresponding to the information record is verified against the metadata in the data storage space. If the current version of the metadata in the data storage space is consistent with the version of the metadata to be submitted, and the version number of the metadata in the data storage space is incremented by 1 relative to the version number of the previous version of the metadata, and the timestamp is in an incrementing state, then it is determined that the metadata to be submitted and the metadata in the data storage space have passed the verification.

[0057] When the above conditions are met simultaneously, a batch task is triggered to allocate newly added or incompletely partitioned data to the newly added storage units. Specifically, a batch processing job with a parallelism of the target number is submitted.

[0058] Among them, the data to be allocated refers to data whose timestamp corresponds to a time later than the timestamp contained in the first statement.

[0059] After the batch task is completed, configure the status field of the information record to be rolled.

[0060] Furthermore, in response to the state machine being configured to the third state, the data processing task is resumed based on the task path.

[0061] Because data processing tasks in the business flow need to be paused when metadata is updated and storage units are reallocated, the flow task recovery process is executed after the batch task is completed to resume normal Flink flow task processing.

[0062] The task path, or savepointPath field in the information record, contains the task state, operator state, and data sharding information of the data processing task (i.e., the Flink streaming task) before it was paused. Based on the information recorded in the task path, the parallelism of the Flink streaming task is set to the target number, and the processing progress of the data processing task is restored to its state before the pause.

[0063] Accordingly, in response to a successful data processing task recovery, the state machine is configured to the fourth state; or, in response to a failed data processing task recovery, the version of the metadata in the data storage space is rolled back to the previous version, and the state machine is configured to the fifth state.

[0064] If no anomalies occur after the stream task is restored, the state machine is configured to the fourth state, and the status field of the information record is configured to completed, indicating that the storage space expansion process is complete and the data processing system has returned to normal operation.

[0065] If an anomaly occurs during the stream task recovery process, such as the data processing task failing to start or an anomaly occurring after the data processing task has started, the metadata in the data storage space will be atomically restored, rolled back to the previous metadata version, and the state machine will be configured to the fifth state, providing a basis for subsequent manual troubleshooting and retries.

[0066] This embodiment of the disclosure improves the system's automation level by using a state machine for process control, automatically triggering actions such as data redistribution and data processing task recovery. Simultaneously, it ensures the continuity and accuracy of business flows by recovering data processing tasks through task paths.

[0067] Figure 3 A flowchart illustrating a storage space management method according to another embodiment of this disclosure. Figure 3 As shown, the method includes the following steps: S301, Operating indicators of the data acquisition and processing system.

[0068] S302. When the operating indicators meet the preset conditions, calculate the target number of storage units in the data storage space.

[0069] S303. Create an information record in the preset information storage location.

[0070] Specifically, the Flink task ID, HDFS path (savepointPath), number of targets, metadata version number, and status are stored as the value in a pre-defined information storage location (e.g., Redis), with the timestamp used as the key. The status includes one of the following: pending, committed, rolled, completed, or failed.

[0071] S304, hot update of metadata.

[0072] Specifically, a first statement is created based on the information record in the first state (pending). The first statement is submitted to the data storage space through the routing service, so that the data storage space can obtain the metadata to be submitted based on the first statement and change the number of storage units to the target number.

[0073] S305, batch and stream collaborative switching.

[0074] In response to the state machine being configured to the second state (committed) and the metadata to be committed being verified with the metadata in the data storage space, the data to be allocated in the data storage space is allocated to the target number of storage units, and the state machine is configured to the third state (rolled).

[0075] In response to the state machine being configured to the third state, the data processing task is resumed based on the task path.

[0076] In response to a successful data processing task recovery, the state machine is configured to the fourth state (completed); or, in response to a failed data processing task recovery, the version of the metadata in the data storage space is rolled back to the previous version, and the state machine is configured to the fifth state.

[0077] This embodiment of the disclosure determines whether the number of storage units in the current data storage space needs to be adjusted based on the operating indicators of the data processing system. When adjustment is required, it automatically calculates the target number and quickly completes the adjustment. Furthermore, it achieves closed-loop automation of the entire process based on a state machine, simplifying manual operation to autonomous system execution, saving time and effort.

[0078] Figure 4 This is a schematic diagram of a storage space management device provided in an embodiment of the present disclosure. The storage space management device provided in this embodiment can execute the processing flow provided in the storage space management method embodiment, such as… Figure 4As shown, the storage space management device 40 includes: a data acquisition module 41, a calculation module 42, and a submission module 43; the data acquisition module 41 is used to acquire the operating indicators of the data processing system, including the data processing rate of the data source in the data processing system and the write latency of the data storage space; the calculation module 42 is used to calculate the target number of storage units in the data storage space when the operating indicators meet preset conditions; the submission module 43 is used to submit the target number to the data storage space so that the data storage space configures the number of storage units to the target number.

[0079] Optionally, the calculation module 42 is specifically used to calculate the target number of storage units in the data storage space when the ratio of the data processing rate to the preset processing rate is greater than a first preset ratio; or, when the ratio of the write latency to the preset latency is greater than a second preset ratio, calculate the target number of storage units in the data storage space.

[0080] Optionally, the calculation module 42 is specifically used to obtain the number of storage units at the current time, and obtain the current number of storage units; determine twice the current number as the first candidate number; determine the number of storage units corresponding to the available resource information of the data processing system as the second candidate number; obtain the maximum preset number of storage units, and obtain the third candidate number; and determine the minimum value among the first candidate number, the second candidate number, and the third candidate number as the target number.

[0081] Optionally, the submission module 43 is used to create an information record in a preset information storage location and configure the state machine corresponding to the information record to a first state; wherein, the storage key value of the information record is the timestamp when the information record is created, and the storage value of the information record is the metadata to be submitted, the metadata to be submitted includes the number of storage units, and the number of storage units is the target number; in response to the state machine being in the first state, a first statement is created based on the information record, the first state indicating that there is an information record to be processed in the preset information storage location; the first statement is submitted to the data storage space through the routing service, so that the data storage space obtains the metadata to be submitted based on the first statement and changes the number of storage units to the target number.

[0082] Optionally, the storage space management device 40 further includes a data allocation module, which, in response to the state machine being configured to a second state and the metadata to be submitted being verified with the metadata in the data storage space, allocates the data to be allocated in the data storage space to a target number of storage units, the second state indicating that the number of storage units has been successfully changed to the target number; and configures the state machine to a third state, the third state indicating that the data in the data storage space has been redistributed.

[0083] Optionally, the information record includes the task path of the data processing task in the data processing system at the current moment; the storage space management device 40 also includes a task recovery module, which is used to recover the data processing task based on the task path in response to the state machine being configured to the third state.

[0084] Optionally, the task recovery module is also used to configure the state machine to the fourth state in response to a successful data processing task recovery; or, in response to a data processing task recovery failure, to roll back the version of the metadata in the data storage space to the previous version and configure the state machine to the fifth state.

[0085] Figure 4 The storage space management device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0086] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device provided in an embodiment of this disclosure can execute the processing flow provided in the storage space management method embodiment, such as... Figure 5 As shown, the electronic device 50 includes: a memory 51, a processor 52, a computer program, and a communication interface 53; wherein the computer program is stored in the memory 51 and is configured to be executed by the processor 52 using the memory space management method described above.

[0087] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the storage space management method described in the above embodiments.

[0088] Furthermore, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the storage space management method described above.

[0089] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0090] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A storage space management method characterized by comprising: The method comprises: collecting an operation index of a data processing system, the operation index comprising a data processing rate of a data source in the data processing system and a write delay of a data storage space; when the operation index meets a preset condition, calculating a target number of storage units in the data storage space; submitting the target number to the data storage space, so that the data storage space configures the number of storage units as the target number.

2. The method of claim 1, wherein, The method further comprises: when the operation index meets a preset condition, calculating a target number of storage units in the data storage space; when the ratio of the data processing rate to a preset processing rate is greater than a first preset ratio, calculating the target number of storage units in the data storage space; or 3. The method of claim 1, wherein, when the ratio of the write delay to a preset delay is greater than a second preset ratio, calculating the target number of storage units in the data storage space. The method further comprises: obtaining the number of storage units at the current time to obtain a current number of storage units; determining that twice the current number is a first candidate number; determining that the number of storage units corresponding to the available resource information of the data processing system is a second candidate number; obtaining a maximum preset number of storage units to obtain a third candidate number; 4. The method of claim 1, wherein, determining that the minimum value of the first candidate number, the second candidate number, and the third candidate number is the target number. The method further comprises: creating an information record in a preset information storage location, and configuring a state machine corresponding to the information record to a first state; wherein the storage key value of the information record is a timestamp when the information record is created, and the storage value of the information record is to-be-submitted metadata, the to-be-submitted metadata comprising the number of storage units, and the number of storage units being the target number; in response to the state machine being in the first state, creating a first statement based on the information record, the first state indicating that there is a to-be-processed information record in the preset information storage location; 5. The method of claim 4, wherein, submitting the first statement to the data storage space through a routing service, so that the data storage space obtains the to-be-submitted metadata based on the first statement, and changes the number of storage units to the target number. The method further comprises: in response to the state machine being configured to a second state and the to-be-submitted metadata being verified with metadata in the data storage space, allocating to-be-allocated data in the data storage space to the target number of storage units, the second state indicating that the number of storage units is successfully changed to the target number; 6. The method of claim 5, wherein, configuring the state machine to a third state, the third state indicating that data re-allocation in the data storage space is complete. The information record comprises a task path of a data processing task in the data processing system at the current time; the method further comprises: In response to the state machine being configured as the third state, the data processing task is resumed based on the task path.

7. The method of claim 6, wherein, The method further includes: In response to the data processing task being resumed successfully, the state machine is configured as a fourth state; or, In response to the data processing task being resumed unsuccessfully, a version of metadata in the data storage space is rolled back to a previous version, and the state machine is configured as a fifth state.

8. A storage space management apparatus characterized by comprising: Comprise: a collection module, configured to collect a running index of a data processing system, the running index comprising a data processing rate of a data source in the data processing system and a write delay of a data storage space; a calculation module, configured to calculate a target number of storage units in the data storage space when the running index meets a preset condition; a submission module, configured to submit the target number to the data storage space, so that the data storage space configures the number of storage units as the target number.

9. An electronic device, comprising: Comprise: a memory; a processor; and a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1-7. The computer program is executed by the processor to implement the method of any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Data storage method, device and equipment based on distributed system and storage medium

    CN115481295A

  • Bucket dividing method and device for data lake

    CN120020750A

  • Bucket dividing regulation and control method and device for financial data warehouse and electronic equipment

    CN120873092A

  • Machine learning-based interactive conversation system

    US20240021196A1

  • Data storage method and apparatus, electronic device, and storage medium

    US20240272814A1