Method, device and equipment for allocating storage resource and storage medium
By acquiring namespace image files and parsing distributed data warehouses, combined with table lineage methods, the problem of unreasonable storage resource allocation in existing technologies is solved, enabling reasonable allocation of storage resources and optimization of enterprise data governance.
Patent Information
- Application Number
- CN202211114641.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-09-14
AI Technical Summary
The existing system has problems with unreasonable allocation of storage resources, which affects subsequent business optimization and has a significant impact on the performance of the name service node, making it impossible to effectively optimize the allocation of storage resources.
The namespace image file is obtained through the preset interface of the name service node, the storage resources are parsed in combination with the distributed data warehouse, and the table lineage is obtained using the extended plugin. Based on the preset allocation rules, the storage resources are distributed to the upper and lower layer data tables.
It enables the rational allocation of storage resources, reduces the performance impact on name service nodes, improves system stability and analysis speed, supports enterprises in better identifying storage resource consumption, and promotes data governance and cost optimization.
Smart Images

Figure CN116126217B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, and in particular to a storage resource allocation method and device, equipment and a storage medium. BACKGROUND
[0002] With the continuous development of science and technology, business management in enterprises needs to rely on electronic systems. For enterprises with a high proportion of Internet business, data storage resource allocation is an important reference for business management.
[0003] The existing system has a greater impact on the performance of related servers when performing storage resource statistics, affecting the performance of other work of the system, and the storage resource consumption collected by the system is very rough, without establishing a structured association with related businesses, making it difficult to effectively optimize storage resource allocation during subsequent enterprise management.
[0004] That is, the existing system has the technical problem of unreasonable allocation of storage resources, affecting subsequent business optimization. SUMMARY
[0005] The present application provides a storage resource allocation method, device, equipment and storage medium to solve the technical problem of unreasonable allocation of storage resources by the existing system, affecting subsequent business optimization.
[0006] In a first aspect, the present application provides a storage cost allocation method, comprising:
[0007] obtaining a namespace image file through a preset interface corresponding to the name service node, the namespace image file including table information corresponding to each of a plurality of first data tables;
[0008] determining, according to the table information of each first data table in the namespace image file, a storage resource of each first data table and a target data table having a table blood relationship with each first data table, the target data table including an upper layer data table of the first data table and a lower layer data table of the first data table, data in the first data table being used to provide for the lower layer data table, or the target data table including an upper layer data table of the first data table, wherein data in the upper layer data table is used to provide for the first data table;
[0009] obtaining a target storage resource corresponding to each lower layer data table of each first data table or a target storage resource corresponding to each first data table based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper layer data table of the first data table and / or the storage resource corresponding to the lower layer data table of the first data table, and a preset allocation rule.
[0010] In a second aspect, the present application provides a storage resource allocation device, comprising:
[0011] The acquisition module is configured to acquire, through a preset interface corresponding to the name service node, a namespace mirror file, wherein the namespace mirror file comprises table information corresponding to each first data table.
[0012] The processing module is configured to determine, according to the table information of each first data table in the namespace mirror file, a storage resource of each first data table and a target data table having a table blood relationship with each first data table, wherein the target data table comprises an upper-layer data table of the first data table and a lower-layer data table of the first data table, data in the first data table is used to provide for the lower-layer data table, or the target data table comprises the upper-layer data table of the first data table, wherein data in the upper-layer data table is used to provide for the first data table.
[0013] The computing module is configured to obtain, based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper-layer data table of the first data table, and / or the storage resource corresponding to the lower-layer data table of the first data table and a preset allocation rule, a target storage resource corresponding to the lower-layer data table of each first data table or a target storage resource corresponding to each first data table.
[0014] In a third aspect, the present application provides an electronic device, comprising:
[0015] The memory is configured to store program instructions.
[0016] The processor is configured to invoke and execute the program instructions in the memory, and execute any one of the possible methods provided in the first aspect.
[0017] In a fourth aspect, the present application provides a storage medium, wherein the readable storage medium stores a computer program, and the computer program is used to execute any one of the possible methods provided in the first aspect.
[0018] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement any one of the possible methods provided in the first aspect.
[0019] The application provides a storage resource allocation method and device, equipment and a storage medium. The name service node corresponds to a preset interface, and a namespace image file is obtained, so that the storage resource statistics is avoided to be directly performed on the name service node, and the working performance of the name service node is affected. Then, according to table information of each first data table in the namespace image file, storage resources of each first data table and target data tables having a table blood relationship with each first data table are determined. The target data tables include upper-layer data tables of the first data table and lower-layer data tables of the first data table. Data in the first data table is used to provide for the lower-layer data tables, or the target data tables include the upper-layer data tables of the first data table, wherein data in the upper-layer data tables is used to provide for the first data table. Finally, based on the storage resources corresponding to each first data table, the storage resources corresponding to the upper-layer data tables of the first data table and / or the storage resources corresponding to the lower-layer data tables of the first data table and a preset allocation rule, target storage resources corresponding to the lower-layer data tables of each first data table or target storage resources corresponding to each first data table are obtained. Thus, the technical problem that the existing system is not reasonable in storage resource allocation and affects subsequent business optimization is solved. The storage resources are allocated to the data tables directly related to the business department, so as to facilitate subsequent data governance and storage resource optimization allocation of the entire company based on the department and business. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.
[0021] Figure 1 A flowchart of a storage resource allocation method provided by the application is shown in the figure.
[0022] Figure 2 A flowchart of another storage resource allocation method provided by the embodiment of the application is shown in the figure.
[0023] Figure 3 An application scenario diagram of the storage resource allocation method provided by the embodiment of the application is shown in the figure.
[0024] Figure 4 A structure diagram of the storage resource allocation device provided by the embodiment of the application is shown in the figure.
[0025] Figure 5 A structure diagram of an electronic device provided by the application is shown in the figure.
[0026] The present application has been shown and described with reference to the preferred embodiments. Equivalent mechanisms and treatments exist and are apparent to those skilled in the art in light of this disclosure and are intended to be within the scope of the application. DETAILED DESCRIPTION
[0027] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work, including but not limited to combinations of multiple embodiments, belong to the scope of protection of the present application.
[0028] The terms "first", "second", "third", "fourth" and the like in the description and the claims of the present application and the above drawings, if any, are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] The professional terms involved in the present application are explained as follows:
[0030] Hadoop: a distributed system infrastructure developed by the Apache Foundation. Users can develop distributed programs without knowing the details of the underlying distributed system. Take full advantage of the power of clusters for high-speed computing and storage. Hadoop implements a distributed file system (Distributed File System), one of which is HDFS (Hadoop Distributed File System).
[0031] HDFS (Hadoop Distributed File System): A distributed file storage system in Hadoop with high fault tolerance and designed to be deployed on low-cost hardware. It provides high throughput to access data for applications with large data sets. HDFS relaxes the POSIX requirements and allows streaming access to data in the file system.
[0032] NameNode: A software that runs on a separate machine from the HDFS instance. It is responsible for managing the file system namespace and controlling access from external clients. The NameNode decides where to map a file to replicated blocks on DataNodes. For the most common case of 3 replicated blocks, the first block is stored on different nodes in the same rack, and the last block is stored on a node in a different rack. Actual I / O transactions do not go through the NameNode, only the metadata representing the file mapping to DataNodes and blocks go through the NameNode. When an external client sends a request to create a file, the NameNode responds with the block identifiers and the DataNode IP address of the first replica of the block. The NameNode also informs other DataNodes that will receive replicas of the block. The NameNode stores all information about the file system namespace in a file called FsImage. This file and a log file containing all transactions (here, EditLog) are stored on the local file system of the NameNode. The FsImage and EditLog files also need to be replicated for protection against file corruption or loss of the NameNode system.
[0033] DataNode: Also a software that runs on a separate machine, usually in an HDFS instance. A Hadoop cluster contains one NameNode and a large number of DataNodes. DataNodes are usually organized in racks, which are connected by a switch. One assumption of Hadoop is that the transfer speed between nodes in a rack is faster than the transfer speed between nodes in different racks. DataNodes respond to read and write requests from HDFS clients. They also respond to commands from the NameNode to create, delete, and replicate blocks. The NameNode relies on periodic heartbeat messages from each DataNode. Each message contains a block report, from which the NameNode can verify the block mapping and other file system metadata. If a DataNode fails to send a heartbeat message, the NameNode will take remedial action to replicate the lost blocks on that node.
[0034] getContentSummary: An interface in the NameNode to return a summary of the namespace containing statistics such as storage size.
[0035] fsimage: A file in HDFS that stores the namespace structure. When the NameNode starts, it reloads this file to load the current namespace.
[0036] Hive: A distributed data warehouse built on Hadoop that provides SQL to facilitate users to do distributed processing and analysis on data.
[0037] MySQL: A relational database that is widely used in the industry.
[0038] LDAP: Lightweight Directory Access Protocol. Enterprises usually use this service to record the organizational structure of the enterprise and provide unified authentication.
[0039] SQL (Structured Query Language): A special-purpose programming language for database queries and programming, used to access data and query, update, and manage relational database systems.
[0040] The existing system, such as a cost accounting system, when performing storage resource accounting, first takes a table path as a parameter, acquires each table storage space by calling a getContentSummary interface of a Namenode, and writes the storage of each table into a table storage space record table in a MySQL. Then, a management page of the database is accessed, and then a SQL operation statement executed in the database and related information of an executor of the SQL operation statement are parsed from the management page. After the SQL is acquired, a table blood relationship of the query, that is, a table blood relationship between each data table, is parsed by using a SQL parser. After the table blood relationship is parsed, the system writes the information into the MySQL database.
[0041] After the table storage space analysis of the previous day and the storage into the database are completed in the early morning of each day, the storage resource of the previous day can be started to be counted, and the storage resource can be used in various application scenarios. For example, a system administrator sets a storage cost unit price in the system, and the cost accounting system uses the set storage cost unit price to multiply the consumed storage space of a table, so as to obtain a table storage cost. After the counting is completed, the system writes the table storage cost into a table storage cost table in the MySQL database, so as to generate a report based on the data or send a related storage cost statistical email to a related user and a department leader.
[0042] The existing system has the following problems:
[0043] (1) When the table storage space is counted by using the getContentSummary interface, if there are many stored files in the table, the influence on the Namenode is large.
[0044] (2) The use of the polling access to the Hive management page to acquire the SQL has a certain influence on the performance of the Hive service, and the number of SQLs displayed in the Hive management page is limited. If the cost management system does not acquire the SQL in the page for a long time, the old SQL can not be collected again after the new SQL is run.
[0045] (3) The cost accounting system directly counts the storage cost of each table, and does not allocate the underlying storage cost to the upper business and the corresponding department. It is not convenient for subsequent data governance and storage cost optimization of the entire company based on the department and the business.
[0046] In summary, the existing system has the technical problem of unreasonable allocation of storage resources, which affects the subsequent business optimization.
[0047] To solve the above problems, the inventive concept of the present application is:
[0048] The storage resource of the content summary interface is directly counted without using the Namenode, and the corresponding file is downloaded from the Namenode for analysis, so that the performance of the Namenode is not affected. Further, the storage resource consumption of each data table is counted by means of distributed analysis, so that the problem of large resource consumption caused by single machine analysis is avoided. Moreover, not only the storage cost of the data table alone is counted, but also the storage cost of the lower layer table or the bottom layer table providing the original business data of the upper layer table corresponding to the business department is allocated, so as to facilitate subsequent data governance based on the department and business and storage cost optimization of the whole company.
[0049] It should be noted that the application scenario of the storage cost allocation method provided in the present application includes: being integrated in an enterprise cost management system, or an enterprise data analysis system, or any system for enterprise business data analysis and management.
[0050] Figure 1 A flowchart of a storage resource allocation method provided in an embodiment of the present application is shown in FIG. 1. Figure 1 As shown in the figure, the specific steps of the storage cost allocation method include:
[0051] S101, acquiring a namespace image file through a preset interface corresponding to a name service node.
[0052] In this step, the name service node includes a Namenode, which is a service in HDFS responsible for storing file namespace and metadata information such as file blocks. The preset interface includes an HTTP interface. The namespace image file includes a fsimage file. It should be noted that the namespace image file includes table information corresponding to each first data table.
[0053] In the present application, the namespace image file is downloaded through the preset interface, instead of counting the storage space by calling the statistical interface in the name service node in the prior art, so as to reduce the influence of counting the storage space on the computing performance of the name service node, and increase the system stability.
[0054] S102, determining the storage resource of each first data table and the target data table having a table blood relationship with each first data table according to the table information of each first data table in the namespace image file.
[0055] In the step, the target data table includes the upper data table of the first data table and the lower data table of the first data table, the data in the upper data table is used to provide for the first data table, and the data in the first data table is used to provide for the lower data table, or the target data table includes the upper data table of the first data table, and the data in the upper data table is used to provide for the first data table. For example, the data table B is obtained by using the data in the data table A for statistics and arrangement, then the data table A is the upper data table of the data table B, and the data table B is the lower data table of the data table A. The upper and lower corresponding relationship is the table blood relationship.
[0056] In the embodiment, the storage resource of each first data table and the target data table having the table blood relationship of each first data table in the step can be determined respectively, that is, the storage resource of each first data table is determined first, and then the target data table having the table blood relationship of each first data table is determined, or the target data table having the table blood relationship of each first data table is determined first, and then the storage resource of each first data table is determined.
[0057] (1) According to the table information of each first data table in the namespace mirror file, the storage resource of each first data table is determined, specifically including:
[0058] Firstly, an analysis data table corresponding to the namespace mirror file is created in the distributed database. The distributed data warehouse includes: a hive data warehouse.
[0059] Then, the namespace mirror file is read and parsed to obtain the table information corresponding to each first data table in the namespace mirror file, and the table information corresponding to each first data table is loaded into the analysis data table. For example, the parsing work of different data tables in the namespace mirror file is respectively completed on multiple servers, and then the statistical summary is performed, so that the large resource consumption in single machine analysis can be avoided, and the analysis speed is accelerated.
[0060] Finally, the storage resource corresponding to each first data table is determined by querying the analysis data table through a data query instruction.
[0061] (2) According to the table information of each first data table in the namespace mirror file, the target data table having the table blood relationship with each first data table is determined, specifically including:
[0062] Firstly, the namespace mirror file is sent to the distributed data warehouse, so that the distributed data warehouse extracts the table information of each first data table from the namespace mirror file,
[0063] Then, by an extension plug-in integrated in the distributed data warehouse, the data query instruction is executed in the distributed data warehouse to query the table information of each first data table, and extract a target data table having a table blood relationship with each first data table queried by the data query instruction.
[0064] Specifically, the extension plug-in reads the execution object of the SQL query instruction, i.e., the table name of each first data table and the table name of the upper layer table and the lower layer table corresponding to the first data table, and stores them in a tree storage mode, such as a binary tree form, to obtain a table blood relationship tree, i.e., a so-called table blood relationship.
[0065] The embodiment of the application obtains the table blood relationship through the extension plug-in, avoids obtaining the table blood relationship by polling and analyzing the hive page to obtain the SQL execution object in the prior art, reduces the impact on the hive data warehouse, and reduces the development cost.
[0066] It should be noted that (1) and (2) above can also be performed simultaneously, and there is no order requirement for the two.
[0067] S103, based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper layer data table of the first data table, and / or the storage resource corresponding to the lower layer data table of the first data table, and the preset allocation rule, obtaining the target storage resource corresponding to each lower layer data table of the first data table or the target storage resource corresponding to each first data table.
[0068] In this step, based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper layer data table of the first data table, and / or the storage resource corresponding to the lower layer data table of the first data table, and the preset allocation rule, the target storage resource corresponding to each lower layer data table of the first data table is obtained; or based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper layer data table of the first data table, and the preset allocation rule, the target storage resource corresponding to each first data table is obtained.
[0069] In this step, the preset allocation rule includes: allocating the storage resource of the upper layer table into the storage resource of the lower layer table, and the lower layer table generally corresponds to a business department of the enterprise, so as to help the enterprise to better identify the business on which the storage resource is consumed, so as to reduce the operating cost of the enterprise, such as the storage cost of data, by means of data and asset management. The embodiment of the application provides a storage resource allocation method, which downloads a namespace image file through a preset interface, instead of calling a statistical interface in a name service node to perform storage space statistics in the prior art, reduces the influence on the name service node when the storage space is counted, and increases the system stability. The distributed data warehouse is used to analyze the namespace image file, so as to avoid a large amount of resource consumption when a single machine is analyzed, and to speed up the analysis speed. The table blood relationship is obtained through an extension plug-in, so as to avoid obtaining the table blood relationship by polling and analyzing the hive page in the prior art, reduce the influence on the hive data warehouse, and reduce the development cost. The storage cost of the lower layer table is allocated to the storage cost of the upper layer table, so as to help the enterprise to better identify the business on which the storage cost is consumed, so as to reduce the storage cost of the enterprise by means of data and asset management.
[0070] Figure 2 The flowchart of another storage cost allocation method provided by the embodiment of the application is shown in FIG. 3. Figure 2 The specific steps of the storage cost allocation method include:
[0071] S201, obtaining a namespace image file through a preset interface corresponding to a name service node.
[0072] In this step, the current namespace image file is downloaded through the preset interface corresponding to the name service node at a preset time. The name service node includes: a Namenode, which is a service in the HDFS responsible for storing file name space and metadata information such as file blocks. The preset interface includes: an HTTP interface. The namespace image file includes: a fsimage file.
[0073] Figure 3 The application scenario diagram of the storage resource allocation method provided by the embodiment of the application is shown in FIG. 4. Figure 3 As shown in FIG. 4, the storage resource allocation method provided by the embodiment is integrated in a cost accounting system 100. In this embodiment, the cost accounting system 100 downloads the fsimage file, that is, the namespace image file, from the Namenode name service node 201 of the HDFS distributed file storage system 200 through the HTTP interface at a preset time point, such as 23:40 every night.
[0074] As shown in FIG. 4, the storage resource allocation method provided by the embodiment is integrated in a cost accounting system 100. In this embodiment, the cost accounting system 100 downloads the fsimage file, that is, the namespace image file, from the Namenode name service node 201 of the HDFS distributed file storage system 200 through the HTTP interface at a preset time point, such as 23:40 every night. Figure 3As shown, the HDFS distributed file storage system 200 includes a plurality of Datanode data service nodes 202 and at least one Namenode name service node 201. The cost accounting system 100 interacts with the Hive data warehouse 300, the MySQL relational database 400, and the LDAP directory access service 500. The specific interaction logic is described below.
[0075] S202, creating an analysis data table corresponding to the namespace mirror file in the distributed database.
[0076] In this step, the analysis data table is created in the distributed database by connecting to the distributed database in a preset manner.
[0077] S203, reading and parsing the namespace mirror file to obtain table information corresponding to each first data table in the namespace mirror file, and loading the table information corresponding to each first data table into the analysis data table.
[0078] In this step, the namespace mirror file is read and parsed, the parsing result is determined, and the parsing result is saved in a preset file format. The parsing result contains table information corresponding to each first data table, for example, the parsing result includes: file path, file size, file block number, etc.
[0079] The table field name in each first data table in the preset format is the same as the table field name in the first data table. That is, the analysis data table and the file format of the parsing result are the same, and the field order in the analysis data table is determined according to the file format.
[0080] Specifically, the table information corresponding to each data table is loaded into the analysis data table, including: uploading the table information corresponding to each first data table in the preset format to the temporary directory of the distributed file storage system; loading the table information corresponding to each first data table into the analysis data table through the loading instruction of the distributed data warehouse.
[0081] In this embodiment, as shown in the figure, Figure 3 After the cost accounting system 100 completes the download of the fsimage file, the namespace mirror file is read and parsed. The data after parsing, i.e. the parsing result, contains fields such as file path, file size, and block number. The data after parsing, i.e. the parsing result, is saved in the format of a csv file for subsequent processing.
[0082] Then, the cost accounting system 100 connects the Hive data warehouse 300 by means of JDBC, and creates an analysis data table corresponding to the analysis fsimage file in the Hive data warehouse 300, wherein the storage format of the analysis data table is csv, and the field order in the analysis data table is the same as that in the csv file.
[0083] Then, the cost accounting system 100 uploads the parsed fsimage file stored in the csv format to the temporary directory of the data service node 202 in the HDFS distributed file storage system 200, and then loads the data in the parsed namespace mirror image file, i.e., the parsing result, into the analysis data table by using the LOAD DATA statement of the Hive data warehouse 300.
[0084] S204, querying the analysis data table by means of a data query instruction to determine the storage resource corresponding to each first data table.
[0085] In this step, the data query instruction is sent to the distributed data warehouse to enable the distributed data warehouse to perform storage resource statistics on the table information corresponding to each first data table in the analysis data table, and store the storage resource corresponding to each first data table into a storage resource record table.
[0086] For ease of understanding, in this embodiment, as shown in Figure 3 The cost accounting system completes the table-level space statistics by sending a SQL database query statement to the Hive data warehouse 300, wherein the grouping field in the SQL is the path, and the table prefix part in the path is intercepted, for example, the file path of the data file corresponding to each table information in the analysis data table is “ / Hive unified prefix / library name / table name / file name”. Therefore, “ / Hive unified prefix / library name / table name” is used as the grouping field. After the storage resource, such as the storage space, occupied by each first data table is analyzed, the cost accounting system 100 writes the information into the table storage resource record table in the MySQL relational database 400.
[0087] S205, querying the table information of each first data table by means of the data query instruction in the distributed data warehouse, and extracting a target data table having a table blood relationship with each first data table.
[0088] In this step, the target data table includes an upper-layer data table of the first data table and a lower-layer data table of the first data table, or the target data table includes an upper-layer data table of the first data table, the data in the upper-layer data table is used to provide for the first data table, and the data in the first data table is used to provide for the lower-layer data table.
[0089] It should be noted that S204 and S205 are performed synchronously, that is, while the distributed data warehouse performs the data query instruction to query the table information of each first data table to obtain the storage resource corresponding to each first data table, the distributed data warehouse synchronously extracts the target data table having the table blood relationship with each first data table through the integrated extension plug-in of the distributed data warehouse.
[0090] For ease of understanding, in the embodiment, as shown in Figure 3 The extension plug-in of the Hive data warehouse 300 is implemented according to the specification of the Hive, and the plug-in is integrated into the Hive data warehouse 300. The function of the plug-in is to obtain the SQL executor and the table blood relationship between the data tables after the execution of the SQL database operation instruction, and send the information to the cost accounting system 100 through the HTTP mode. After the cost accounting system 100 receives the information of the SQL database operation instruction sent from the Hive data warehouse 300, the cost accounting system 100 will perform the following two operations:
[0091] 1) The cost accounting system 100 writes the information of the table blood relationship between the data tables into the table blood relationship record table, that is, the second record table, in the MySQL relational database 400.
[0092] 2) The cost accounting system 100 queries the LDAP directory access service 500 by using the information of the SQL executor to obtain the department to which the executor belongs, and writes the information into the table responsibility person record table, that is, the third record table, in the MySQL relational database 400.
[0093] S206, load the table blood relationship in the form of a dictionary into the memory.
[0094] In this step, the table blood relationship of each first data table includes: each lower data table, each upper data table, and the department information of the business department corresponding to each lower data table.
[0095] In the embodiment, the table blood relationship record table in the MySQL relational database 400 loads the blood relationship record into the memory in the cost accounting system 100, so as to avoid repeatedly querying the relational database 400 when performing cost allocation calculation, and to affect the calculation efficiency. The blood relationship record is stored in the memory of the cost accounting system 100 in the form of a dictionary, wherein the key of the dictionary is the input data table obtained by querying the second record table, and the value of the dictionary is the output data table obtained by querying the table blood relationship record table.
[0096] S207, searching in the dictionary according to the name of each first data table to obtain the storage resource of at least one upper data table and / or the storage resource of at least one lower data table corresponding to each first data table.
[0097] Preferably, the storage resource of all upper data tables and / or the storage resource of all lower data tables corresponding to each first data table is obtained by searching in the dictionary according to the name of each first data table.
[0098] In the embodiment, the cost accounting system 100 reads the data in the table storage space record table in the MySQL relational database 400 in a paging manner, associates each piece of data in the page with the blood relationship in the memory, i.e., the table blood relationship and the dictionary in the memory, and traverses all upper data tables and / or lower data tables corresponding to each first data table in a breadth-first manner.
[0099] S208, obtaining the target storage resource of each lower data table of each first data table or the target storage resource of each first data table based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper data table of the first data table and / or the storage resource corresponding to the lower data table of the first data table, and a preset allocation rule.
[0100] In this step, the storage resource of the upper data table is evenly allocated to the storage resource of each lower data table corresponding to the same upper table.
[0101] If a first data table has no corresponding lower data table, the data table is a leaf data table; if a first data table has no corresponding upper data table, the data table is a root data table; if a first data table has both corresponding upper data table and corresponding lower data table, the data table is an intermediate layer data table.
[0102] In the embodiment, the intermediate layer allocation resource and / or the bottom layer allocation resource are added in the storage resource of the leaf data table. The intermediate layer allocation resource is determined according to the first quantity of all leaf data tables corresponding to the intermediate layer data table and the storage resource of the intermediate layer table; the bottom layer allocation resource is determined according to the second quantity of all leaf data tables corresponding to the root data table and the storage resource of the root data table.
[0103] Specifically, for example, the table blood relationship records the processing relationship between each data table, for example, for the table blood relationship: A→B, A→C and B→D, we can know that the data table B and the data table C are processed by the data of the data table A, and the data table D is processed by the data of the data table B. Since the data table B and the data table C are processed by the data of the data table A, we can consider that the data table A is the data source of the data table D. Therefore, when we allocate storage resources later, we can allocate part of the storage resources related to the data table A and all the storage resources of the data table B to the data table D, because the purpose of the data table A and the data table B is to process the D table. In the above example, the data table A is the root data table, the data table C and the data table D are the leaf data tables, and the data table B is the intermediate layer data table.
[0104] It should be noted that one application of the storage resource is to calculate the storage cost, for example, by pre-setting the unit cost corresponding to the storage resource, the product of the storage resource and the unit cost can be used as the storage cost.
[0105] For example, the storage cost P of the data table C C which can be represented by formula (1):
[0106]
[0107] where p represents the storage cost unit price set by the administrator in the system, V A represents the storage resource, such as storage space, occupied by the data table A, and because the data table A corresponds to 2 leaf data tables in total, the coefficient is 1 / 2, V C represents the storage resource, such as storage space, occupied by the data table C.
[0108] The storage cost P of the data table D D which can be represented by formula (2):
[0109]
[0110] where p represents the storage cost unit price set by the administrator in the system, V A represents the storage resource, such as storage space, occupied by the data table A, and because the data table A corresponds to 2 leaf data tables in total, the coefficient is 1 / 2, V B represents the storage resource, such as storage space, occupied by the data table B, V D represents the storage resource, such as storage space, occupied by the data table D.
[0111] The cost accounting system 100 uses the set storage cost unit price to multiply the table allocated storage to obtain the allocated table storage cost. After the statistics are completed, the system writes the table storage cost into the table storage cost table, i.e., the fourth record table, in the MySQL relational database 400, to generate a report based on the data or send a relevant storage cost statistics email to relevant users and department heads.
[0112] The embodiment of the application provides a storage cost allocation method, which analyzes table storage space by analyzing a fsimage file of a Namenode. The fsimage file is downloaded from the Namenode through an HTTP mode, which has no influence on the performance of the Namenode, reduces the influence of the table storage space analysis on the Namenode in the HDFS, analyzes the file in a distributed mode to accelerate the analysis speed, analyzes the fsimage file converted into a csv format through a Hive SQL in a distributed mode to avoid consuming a large amount of single computer resources during the analysis of the fsimage, collects the SQL through an extended plug-in integrated in the Hive to effectively avoid the influence of polling the Hive page on the Hive and development workload, reduces the influence of collecting the Hive SQL on the Hive and reduces the development cost of collection. The storage cost allocation is realized from a bottom table to a top business related table based on table blood relationship analysis, which is beneficial to subsequent cost related data management of an enterprise.
[0113] Figure 4 A structural schematic diagram of a storage resource allocation device provided by the embodiment of the application is provided. The storage resource allocation device 400 can be realized by software, hardware or a combination of both.
[0114] As shown in the Figure 4 The storage resource allocation device 400 includes:
[0115] The acquisition module 401 is configured to acquire a namespace mirror file through a preset interface corresponding to a name service node, and the namespace mirror file includes table information corresponding to each first data table.
[0116] The processing module 402 is configured to determine storage resources of each first data table and a target data table having a table blood relationship with each first data table according to the table information of each first data table in the namespace mirror file, the target data table including an upper data table of the first data table and a lower data table of the first data table, data in the first data table being used to provide for the lower data table, or the target data table including the upper data table of the first data table, data in the upper data table being used to provide for the first data table.
[0117] The computing module 403 is configured to obtain target storage resources corresponding to lower-layer data tables of each first data table or target storage resources corresponding to each first data table based on the storage resources corresponding to each first data table, the storage resources corresponding to upper-layer data tables of the first data table and / or the storage resources corresponding to lower-layer data tables of the first data table, and a preset allocation rule.
[0118] In a possible design, the processing module 402 is configured to:
[0119] create an analysis data table corresponding to the namespace mirror file in the distributed database;
[0120] read and parse the namespace mirror file to obtain table information corresponding to each first data table in the namespace mirror file, and load the table information corresponding to each first data table into the analysis data table;
[0121] query the analysis data table through a data query instruction, and determine the storage resources corresponding to each first data table.
[0122] In a possible design, the processing module 402 is configured to:
[0123] upload the table information corresponding to each first data table in the preset format into a temporary directory of the distributed file storage system;
[0124] load the table information corresponding to each first data table into the analysis data table through a loading instruction of the distributed data warehouse.
[0125] In a possible design, the table field name in each first data table in the preset format is the same as the table field name in the first data table.
[0126] In a possible design, the processing module 402 is configured to:
[0127] send the data query instruction to the distributed data warehouse, so that the distributed data warehouse performs storage resource statistics on the table information corresponding to each first data table in the analysis data table;
[0128] store the storage resources corresponding to each first data table into a storage resource record table.
[0129] In a possible design, the processing module 402 is configured to:
[0130] send the namespace mirror file to the distributed data warehouse, so that the distributed data warehouse extracts the table information of each first data table from the namespace mirror file;
[0131] The distributed data warehouse executes the data query instruction to query the table information of each first data table through an extension plug-in integrated in the distributed data warehouse, and extracts a target data table having a table blood relationship with each first data table queried by the data query instruction.
[0132] In a possible design, the processing module 402 is further configured to:
[0133] The table blood relationship is loaded into the memory in the form of a dictionary.
[0134] The storage resource of at least one upper-layer data table and / or the storage resource of at least one lower-layer data table corresponding to each first data table is obtained by traversing and searching in the dictionary according to the name of each first data table.
[0135] It is worth noting that, Figure 4 The apparatus provided in the embodiments shown can execute the method provided in any of the method embodiments, and the specific implementation principles, technical features, professional term explanations and technical effects are similar, and will not be repeated here.
[0136] Figure 5 A structural schematic diagram of an electronic device is provided in the embodiments of the present application. As shown in the figure, Figure 5 The electronic device 500 can include at least one processor 501 and a memory 502. Figure 5 As shown, the electronic device is an example of an electronic device with one processor.
[0137] The memory 502 is used to store programs. Specifically, the programs can include program codes, and the program codes include computer operation instructions.
[0138] The memory 502 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.
[0139] The processor 501 is used to execute the computer execution instructions stored in the memory 502 to implement the methods described in the above method embodiments.
[0140] The processor 501 can be a central processing unit (CPU) or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0141] Optionally, the memory 502 can be independent or integrated with the processor 501. When the memory 502 is independent of the processor 501, the electronic device 500 can further include:
[0142] A bus 503 is used to connect the processor 501 and the memory 502. The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like, but does not mean that there is only one bus or one type of bus.
[0143] Optionally, in a specific implementation, if the memory 502 and the processor 501 are integrated on a chip, the memory 502 and the processor 501 can communicate through an internal interface.
[0144] The embodiments of the present application further provide a computer readable storage medium, which can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like various media that can store program codes. Specifically, the computer readable storage medium stores program instructions, and the program instructions are used for the method in each method embodiment.
[0145] The embodiments of the present application further provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the method in each method embodiment.
[0146] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The specification and examples given are exemplary only and the true scope and spirit of the application are indicated by the claims. The true scope and spirit of the application are indicated by the claims.
[0147] It should be understood that the present application is not limited to the precise construction that has been described and illustrated herein and that various modifications and changes can be made therein without departing from the scope thereof. The scope of the application is indicated by the appended claims.
Claims
1. A method for allocating storage resources, characterized in that: include: Obtaining a namespace mirror file through a preset interface corresponding to the name service node, wherein the namespace mirror file includes table information corresponding to each of the plurality of first data tables; determining, based on the table information of each first data table in the namespace image file, a storage resource for each first data table and a target data table having a table lineage relationship with each first data table, wherein the target data table includes an upper-layer data table of the first data table and a lower-layer data table of the first data table, and data in the first data table is provided to the lower-layer data table, or the target data table includes an upper-layer data table of the first data table, and data in the upper-layer data table is provided to the first data table; Based on the storage resources corresponding to each of the first data tables, the storage resources corresponding to the upper data table of the first data table and / or the storage resources corresponding to the lower data table of the first data table and the preset allocation rules, the target storage resources corresponding to the lower data table of each of the first data tables or the target storage resources corresponding to each of the first data tables are obtained.
2. The method for allocating storage resources according to claim 1, wherein: The determining, according to the table information of each first data table in the namespace mirror file, the storage resource of each first data table includes: Creating an analysis data table corresponding to the namespace image file in a distributed database; Reading and parsing the namespace mirror file, obtaining table information corresponding to each first data table in the namespace mirror file, and loading the table information corresponding to each first data table into the analysis data table; The analysis data table is queried through a data query instruction to determine the storage resources corresponding to each of the first data tables.
3. The method for allocating storage resources according to claim 2, wherein: The step of loading the table information corresponding to each data table into the analysis data table includes: Uploading table information corresponding to each of the first data tables in a preset format to a temporary directory of a distributed file storage system; The table information corresponding to each of the first data tables is loaded into the analysis data table through a loading instruction of the distributed data warehouse.
4. The method for allocating storage resources according to claim 2, wherein: The querying of the analysis data table by using a data query instruction to determine the storage resource corresponding to each first data table includes: Sending the data query instruction to a distributed data warehouse, so that the distributed data warehouse performs storage resource statistics on table information corresponding to each first data table in the analysis data table; The storage resources corresponding to each of the first data tables are stored in a storage resource record table.
5. The method for allocating storage resources according to claim 1, wherein: The determining, based on the table information of each first data table in the namespace mirror file, a target data table having a table lineage relationship with each first data table includes: Sending the namespace mirror file to a distributed data warehouse, so that the distributed data warehouse extracts table information of each first data table from the namespace mirror file; Through the extension plug-in integrated in the distributed data warehouse, a data query instruction is executed in the distributed data warehouse to query the table information of each first data table, and the target data table queried by the data query instruction and having a table lineage relationship with each first data table is extracted.
6. The method for allocating storage resources according to claim 1, wherein: Before obtaining the target storage resource corresponding to each lower-layer data table of the first data table based on the storage resource corresponding to each first data table, the storage resource corresponding to the upper-layer data table of the first data table, the storage resource corresponding to the lower-layer data table of the first data table, and the preset allocation rule, the method further includes: Loading the table of kinship relations into memory in the form of a dictionary; A traversal search is performed in the dictionary according to the name of each first data table to obtain storage resources of at least one upper-layer data table and / or storage resources of at least one lower-layer data table corresponding to each first data table.
7. The method for allocating storage resources according to claim 3, wherein: The table field names in each of the first data tables in the preset format are the same as the table field names in the first data table.
8. A storage resource allocation device, characterized in that: include: an acquisition module, configured to acquire a namespace mirror file through a preset interface corresponding to a name service node, wherein the namespace mirror file includes table information corresponding to each of the plurality of first data tables; a processing module, configured to determine, based on table information of each first data table in the namespace image file, a storage resource for each first data table and a target data table having a table lineage relationship with each first data table, wherein the target data table includes an upper-layer data table of the first data table and a lower-layer data table of the first data table, and data in the first data table is provided to the lower-layer data table, or the target data table includes an upper-layer data table of the first data table, and data in the upper-layer data table is provided to the first data table; A calculation module is used to obtain the target storage resources corresponding to the lower-level data table of each first data table or the target storage resources corresponding to each first data table based on the storage resources corresponding to each first data table, the storage resources corresponding to the upper-level data table of the first data table and / or the storage resources corresponding to the lower-level data table of the first data table and preset allocation rules.
9. An electronic device, characterized in that: include: processor; as well as, a memory for storing a computer program for the processor; The processor is configured to execute the storage resource allocation method according to any one of claims 1 to 7 by executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the storage resource allocation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Dispersed autonomous storage resource aggregation method with unified name space
CN110213352A