Method, device, equipment and product for determining data storage condition in database
By establishing threads based on the concurrent number configuration parameters of the database cluster and assigning data storage statistics tasks to these threads in parallel, the problems of inefficient statistics and inability to meet real-time requirements in the existing technology are solved, and efficient and flexible data table storage status statistics are achieved.
Patent Information
- Application Number
- CN202510481246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the data table storage status in the prior art, when the data table is stored in a statistical database, it is impossible to flexibly set threads according to the situation of the database cluster, resulting in wasted computing resources, inefficient statistics, and unable to meet the real-time requirements.
According to the concurrency configuration parameters of the database cluster, a corresponding number of threads is established, and the data storage status statistics task is assigned to each thread to execute in parallel, and the storage status of a single data table in the database is calculated and generated based on the execution results of all threads.
It improves statistical efficiency, meets real-time requirements, and can flexibly set up statistical tasks in data tables, which facilitates database administrators to timely understand and handle data table storage status issues and improves database operation performance.
Smart Images

Figure CN120011195A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of databases, and in particular relates to a method, device, equipment and product for determining data storage conditions in a database. Background Art
[0002] With the advent of the big data era, the amount of data stored in the database is becoming increasingly large and complex. In the daily management and maintenance of the database, database administrators need to understand the storage status of each data table in the database in a timely manner. Under the existing technical conditions, when performing storage status statistics on data tables, statistics can only be performed in a unified and fixed manner. The threads used for statistics are relatively single and cannot be flexibly set according to the situation of the database cluster, wasting a lot of computing resources. Especially when facing large-scale databases, the statistical efficiency is low and cannot meet the real-time requirements. At the same time, the form of statistical tasks is relatively single and cannot be flexibly set, which makes it difficult for database administrators to promptly discover and deal with problems in the storage status of data tables, thereby affecting the query performance and overall operation efficiency of the database. Summary of the invention
[0003] In view of this, the present invention aims to overcome the defects in the prior art and proposes a method, device, equipment and product for determining the data storage status in a database.
[0004] To achieve the above object, the technical solution of the present invention is achieved as follows: In a first aspect, the present invention discloses a method for determining data storage conditions in a database, comprising: Create a corresponding number of threads based on the concurrent configuration parameters of the database cluster; Assign the data storage situation statistics task to each thread for parallel execution. The data storage situation statistics task is used to count the storage status of a single data table in the database. Based on the execution results of all threads, the storage status of a single data table in the database is calculated and generated.
[0005] In another embodiment of the present invention, the storage status includes the disk space occupied by the storage of the data table; based on the execution results of all threads, the storage status of a single data table in the database is calculated and generated, including: determining the disk space occupied by the storage of the single data table based on the system table and data dictionary of the database.
[0006] In another embodiment of the present invention, the storage status includes the degree of data skew of the data table; based on the execution results of all threads, the storage status of a single data table in the database is calculated and generated, including: based on the storage occupied disk space of a single data table shard stored on all nodes of the database cluster, the data skew coefficient of the single data table is calculated, and the data skew coefficient is used to characterize the degree of data skew.
[0007] In another embodiment of the present invention, a corresponding number of threads are established according to a concurrency configuration parameter of a database cluster, including: reading a configuration file of the database cluster to determine the concurrency configuration parameter.
[0008] In another embodiment of the present invention, the data storage situation statistics task is assigned to each thread for parallel execution, including: determining the scope of the data tables involved in the data storage situation statistics task according to a preset file, the scope of the data tables including: the entire database data table and the specified data table.
[0009] In another embodiment of the present invention, allocating data storage situation statistics tasks to various threads for parallel execution includes: evenly allocating data storage situation statistics tasks to various threads.
[0010] In another embodiment of the present invention, before allocating the data storage situation statistics task to each thread for parallel execution, the method further includes: determining the data storage situation statistics task type according to preset parameters.
[0011] In a second aspect, the present invention discloses a device for determining data storage conditions in a database, the device comprising: Establish a thread module to create a corresponding number of threads according to the concurrent configuration parameters of the database cluster; A task execution module is used to assign data storage situation statistics tasks to various threads for parallel execution. The data storage situation statistics tasks are used to count the storage status of a single data table in the database. The result generation module is used to calculate and generate the storage status of a single data table in the database according to the execution results of all threads.
[0012] In a third aspect, the present invention discloses an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above method.
[0013] In a fourth aspect, the present invention discloses a computer program product, including a computer program, which implements the above method when executed by a processor.
[0014] Compared with the prior art, the present invention has the following advantages: The present invention discloses a method, device, equipment and product for determining the data storage situation in a database, including establishing a corresponding number of threads according to the concurrent number configuration parameters of the database cluster; assigning data storage situation statistics tasks to each thread for parallel execution; and calculating and generating the storage status of a single data table in the database according to the execution results of all threads. The present invention discloses a method, device, equipment and product for determining the data storage situation in a database, which can establish a corresponding number of threads according to the characteristics of the database cluster, fully mobilize computing resources, and execute data storage situation statistics tasks in parallel, effectively improving the statistical efficiency and meeting the real-time requirements. At the same time, the statistical tasks of the data table can be flexibly set to realize the statistics of the storage status of different types of data tables, so that the database administrator can timely understand the problems existing in the storage status of the data table in the database from different angles, timely maintain the database, and improve the operation performance of the database. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0016] In the attached picture: Figure 1 A schematic diagram of an application scenario of a method for determining data storage status in a database according to an embodiment of the present invention; Figure 2 A schematic diagram of a method for determining data storage conditions in a database according to an embodiment of the present invention; Figure 3 A schematic diagram of data skewness of a method for determining data storage conditions in a database according to an embodiment of the present invention; Figure 4 A schematic diagram of the range of data tables involved in a method for determining data storage conditions in a database according to an embodiment of the present invention; Figure 5 A schematic diagram of a device for determining data storage conditions in a database according to an embodiment of the present invention; Figure 6 A schematic diagram of an electronic device for determining data storage status in a database according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0018] In the description of the present invention, it should be further explained that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0019] The present invention discloses a method, device, equipment and product for determining data storage status in a database. Figure 1 As shown, when the storage status of a data table is counted under the prior art, statistics can only be performed in a unified and fixed manner, and the threads used for statistics are relatively single, and cannot be flexibly set according to the situation of the database cluster, wasting a lot of computing resources. Especially when facing a large-scale database, the statistical efficiency is low and cannot meet the real-time requirements. At the same time, the form of the statistical task is also relatively single and cannot be flexibly set, which makes it difficult for the database administrator to timely discover and handle the problems existing in the storage status of the data table, thereby affecting the query performance and overall operation efficiency of the database. The present invention discloses a method, device, equipment and product for determining the data storage status in a database, which can establish a corresponding number of threads according to the characteristics of the database cluster, fully mobilize computing resources, and execute data storage situation statistics tasks in parallel, effectively improving the statistical efficiency and meeting the real-time requirements. At the same time, the statistical tasks of the data table can be flexibly set to realize the statistics of the storage status of different types of data tables, so that the database administrator can timely understand the problems existing in the storage status of the data table in the database from different angles, timely maintain the database, and improve the operation performance of the database.
[0020] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0021] In one embodiment disclosed in the present invention, Figure 2 As shown, a method for determining data storage status in a database includes: Step S201, establishing a corresponding number of threads according to the concurrent configuration parameters of the database cluster; In this embodiment, the configuration file of the database cluster is read to determine the concurrency configuration parameter.
[0022] Exemplarily, various parameters in the configuration file config.ini of the database cluster are read, including database connection information, such as host, user, password, database, port, etc., statistical mode flag flag, display row number num, concurrency number conc, and table range parameter file tblist, etc.
[0023] The statistical mode flag is used to indicate the type of data storage statistics task selection; The concurrency number conc is used to indicate the concurrency number of cluster nodes, and a corresponding number of threads are established based on the concurrency number.
[0024] The table range parameter file tblist is used to indicate the database's entire database table and specified data tables.
[0025] Step S202, assigning data storage situation statistics tasks to various threads for parallel execution, where the data storage situation statistics tasks are used to count the storage status of a single data table in the database; In this embodiment, by adopting a multi-threaded concurrent execution mode, cluster resources are fully utilized to improve the statistical speed while ensuring statistical accuracy, while avoiding adverse effects on the normal operating environment of the system.
[0026] In this embodiment, if Figure 4 As shown, according to the preset file, the scope of the data tables involved in the data storage situation statistics task is determined, and the scope of the data tables includes: the database's entire database data table and the specified data table.
[0027] In this embodiment, the preset file is the table range parameter file tblist. If the table range parameter file tblist exists, the specified data table recorded in the table range parameter file tblist is allocated to each thread, and each thread is responsible for processing the statistical task of a part of the data table.
[0028] If the table range parameter file tblist does not exist, all database tables are evenly distributed to each thread for statistics.
[0029] Existing statistical tools usually use a sequential scanning method to count the storage status of each data table, traversing the data tables in the database one by one, and cannot specify the data table. When faced with massive data, it takes too long, which seriously affects the efficiency of database management. When administrators only need to pay attention to some important data tables, the existing statistical methods will lead to unnecessary waste of computing resources, which will have a great impact on the normal operation of the database, and may even cause a brief interruption of database services or performance degradation, affecting business continuity and user experience.
[0030] In this embodiment, the method supports two modes: statistics of the entire database table range and statistics of specified data tables, which meet the management needs in different scenarios and improve the flexibility and pertinence of statistics. For example, when conducting overall database performance evaluation, you can choose statistics of the entire database range; when paying attention to the storage situation and data distribution of a specific business module, you can specify the range of the data table for statistics by setting the table range parameter file tblist, avoiding unnecessary calculations and waste of resources, and improving the pertinence and practicality of statistics.
[0031] In an embodiment, data storage situation statistics tasks are evenly distributed to various threads.
[0032] In this embodiment, thread allocation depends on several management node IP address information configured in the host parameter.
[0033] In an embodiment, data storage situation statistics tasks are evenly distributed to each thread and executed concurrently, which can fully utilize the symbiotic characteristics of management nodes of the distributed cluster, fully utilize the computing resources of multiple management nodes, improve statistical efficiency, ensure load balancing between threads, and avoid situations such as mutual waiting.
[0034] In this embodiment, the data storage situation statistics task type is determined according to a preset parameter, and the preset parameter is a statistics mode flag.
[0035] Step S203, based on the execution results of all threads, calculate and generate the storage status of a single data table in the database.
[0036] In this embodiment, the display format of the storage status of a single data table in the database is as follows: the first three fields represent the virtual cluster name, library name and table name respectively; when the statistical mode flag flag = 1, the fourth field represents the disk space occupied by the storage of a data table, and its numerical unit is bytes; when the statistical mode flag flag = 2, the fourth field represents the data skew degree of the data table; Exemplarily, the values of the fourth field are arranged from high to low so that the administrator can quickly obtain the most important statistical information. According to the parameter num of the number of display rows, it is determined that only the first num rows of statistical results are retained to highlight the data table that occupies the most disk space or the data table with the most serious data skew. If the user needs to record the statistical information in a file, the result can be output to the specified my.log file by executing the . / balance>my.log command, which can also be used as a data source for secondary data analysis in downstream links.
[0037] The display method disclosed in this embodiment can highlight key statistical information, simplify statistical results, enable administrators to quickly obtain important data, and facilitate performance optimization and resource adjustment.
[0038] Based on the previous embodiment, in another embodiment of the present invention, the storage state includes the storage occupied disk space of the data table, such as Figure 2 As shown, step S203, based on the execution results of all threads, calculate and generate the storage status of a single data table in the database, including: determining the disk space occupied by the storage of the single data table based on the system table and data dictionary of the database.
[0039] Based on the previous embodiment, in another embodiment of the present invention, the storage state includes the data skew degree of the data table; Figure 2 and Figure 3 As shown, step S203, based on the execution results of all threads, calculate and generate the storage status of a single data table in the database, including: based on the storage occupied disk space of a single data table shard stored on all nodes of the database cluster, calculate the data skewness coefficient of the single data table, and the data skewness coefficient is used to characterize the degree of data skew.
[0040] In this embodiment, for example, for a data table that needs statistics, in a distributed database system, based on metadata information, the storage capacity information of the data table fragments stored in multiple nodes can be obtained, for example: A In a possession n In a database cluster with nodes, , ... ,in, Represents a data table A The shards are stored in i The storage of each node occupies disk space. n Indicates shared n nodes, the data skewness coefficient is expressed as , the calculation process is as follows:
[0041] For example, Less than or equal to 0.5 indicates that the data table is evenly distributed, and greater than 0.5 indicates that the data table has a certain degree of skewness. The larger the value, the more serious the skewness. The results are also stored in a temporary data structure to facilitate database administrators to understand.
[0042] In this embodiment, the data skewness coefficient eliminates the influence of the data dimension and the mean value, so that the discreteness between different data sets is comparable, and the data skewness of each data table is quantified more accurately.
[0043] like Figure 5 As shown, the present invention also discloses a device for determining data storage conditions in a database, comprising: A thread establishment module 501 is used to establish a corresponding number of threads according to the concurrent number configuration parameters of the database cluster; Task execution module 502, used to assign data storage situation statistics tasks to various threads for parallel execution, where the data storage situation statistics tasks are used to count the storage status of a single data table in the database; The result generation module 503 is used to calculate and generate the storage status of a single data table in the database according to the execution results of all threads.
[0044] The present invention also discloses an electronic device, such as Figure 6 As shown, an embodiment is disclosed, which is a block diagram of an electronic device suitable for determining the data storage situation in the database as mentioned above.
[0045] The electronic device 60 of this embodiment includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the ROM 602 or the program loaded from the storage part 608 to the RAM 603. The processor 601 may include, for example, a general-purpose microprocessor, an instruction set processor and / or a related chipset and / or a dedicated microprocessor, etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0046] In RAM603, various programs and data required for the operation of electronic device 60 are stored. Processor 601, ROM602 and RAM603 are connected to each other through bus 604, and processor 601 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM602 and / or RAM603. It should be noted that the program can also be stored in one or more memories other than ROM602 and RAM603, and processor 601 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in one or more memories.
[0047] According to an embodiment of the present invention, the electronic device 60 may further include an I / O interface 605, which is also connected to the bus 604. The electronic device 60 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube, a liquid crystal display, and a speaker; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 6010 is also connected to the I / O interface 605 as needed. A removable medium 6011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 6010 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.
[0048] The present invention also provides a computer-readable storage medium.
[0049] The computer-readable storage medium may be included in the electronic device / device system described in the above embodiment; or it may exist independently without being assembled into the electronic device / device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0050] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory RAM, a read-only memory ROM, an erasable programmable read-only memory EPROM or a flash memory, a portable compact disk read-only memory CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.
[0051] Embodiments of the present invention also include a computer program product.
[0052] The computer program product includes a computer program, which contains program codes for executing the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program codes are used to enable the electronic device to implement the method provided by the embodiment of the present invention.
[0053] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium. The program code included in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0054] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written by any combination of one or more programming languages, and specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages. Programming languages include but are not limited to programming languages such as Java, C++, python, C language or similar. The program code can be executed completely on the user computing device, partially on the user device, partially on the remote computing device, or completely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network or a wide area network, or can be connected to an external computing device.
[0055] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It can be understood by those skilled in the art that the features recorded in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such a combination or combination is not explicitly recorded in the present invention. In particular, without departing from the spirit and teaching of the present invention, the features described in the various embodiments and / or claims of the present invention may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present invention.
[0056] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination. The scope of the present invention is limited by the attached claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A method for determining data storage conditions in a database, characterized in that: include: Create a corresponding number of threads based on the concurrent configuration parameters of the database cluster; Allocating a data storage situation statistics task to each of the threads for parallel execution, the data storage situation statistics task being used to count the storage status of a single data table in the database; The storage state of the single data table in the database is generated by calculation according to the execution results of all the threads.
2. A method for determining data storage conditions in a database according to claim 1, characterized in that: The storage status includes the disk space occupied by the storage of the data table; the storage status of the single data table in the database is calculated and generated based on the execution results of all the threads, including: determining the disk space occupied by the storage of the single data table based on the system table and data dictionary of the database.
3. A method for determining data storage conditions in a database according to claim 1, characterized in that: The storage status includes the degree of data skew of the data table; the storage status of the single data table in the database is calculated and generated based on the execution results of all the threads, including: based on the storage occupied disk space stored by a single data table shard on all nodes of the database cluster, a data skew coefficient of the single data table is calculated, and the data skew coefficient is used to characterize the degree of data skew.
4. A method for determining data storage conditions in a database according to claim 1, characterized in that: The establishing of a corresponding number of threads according to the concurrent number configuration parameter of the database cluster includes: reading a configuration file of the database cluster to determine the concurrent number configuration parameter.
5. A method for determining data storage conditions in a database according to claim 1, characterized in that: The data storage situation statistics task is distributed to each of the threads for parallel execution, including: determining the range of the data tables involved in the data storage situation statistics task according to a preset file, and the range of the data tables includes: database full library data tables and designated data tables.
6. A method for determining data storage conditions in a database according to claim 1, characterized in that: The step of allocating the data storage situation statistics task to each of the threads for parallel execution includes: evenly allocating the data storage situation statistics task to each of the threads.
7. A method for determining data storage conditions in a database according to claim 1, characterized in that: Before allocating the data storage situation statistics task to each of the threads for parallel execution, the method further includes: determining the data storage situation statistics task type according to preset parameters.
8. A device for determining data storage conditions in a database, characterized in that: The device comprises: Establish a thread module to create a corresponding number of threads according to the concurrent configuration parameters of the database cluster; A task execution module, used for allocating data storage situation statistics tasks to each of the threads for parallel execution, wherein the data storage situation statistics tasks are used for counting the storage status of a single data table in the database; The result generation module is used to calculate and generate the storage state of the single data table in the database according to the execution results of all the threads.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to perform the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Distributed parallel database system based on Infiniband network and data processing method
CN109933631A
Data table performance detection method and system, computing equipment and computer readable storage medium
CN116775598A
Method and device for disk space management, equipment and storage medium
CN119396346A
In-memory cursor duration temp tables
US20170116266A1
Efficient hybrid parallelization for in-memory scans
US20180060399A1