A method for building and retrieving a hierarchical database based on a tree structure

By converting complex ER graph relationships into tree structures and using hierarchical indexing and optimization algorithms, the problem of low recursive query efficiency in the database is solved, and efficient and accurate data retrieval is achieved.

CN119621730BActive Publication Date: 2025-05-30CHINESE PEOPLES LIBERATION ARMY UNIT 31002
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411854912.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-30
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

When processing complex structured data in a database, the prior art recursive query methods consume a lot of resources and the accuracy rate is difficult to guarantee.

Method used

The hierarchical database library construction and search method based on tree structure is adopted to convert complex ER graph relationships into tree structures, and the search efficiency is improved through the steps of designing data models, building hierarchical indexes, data extraction, extraction and search optimization, data storage and management.

Benefits of technology

Through the hierarchy and directionality of the tree structure, the efficiency and accuracy of data retrieval are significantly improved, and efficient hierarchical traversal algorithms are supported, such as DFS and BFS, which are suitable for scenarios such as permission inheritance and organizational structure expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621730B_ABST
    Figure CN119621730B_ABST
Patent Text Reader

Abstract

The present invention provides a hierarchical database building and retrieval method based on a tree structure, which designs a data model, constructs a hierarchical index, performs data extraction, optimizes extraction and retrieval, and stores and manages data. Due to the inherent hierarchy and directionality of the tree structure, traversal and retrieval can be made more intuitive and efficient. The tree structure supports efficient hierarchical traversal algorithms such as depth-first search (DFS) and breadth-first search (BFS), and these algorithms are very effective for scenarios such as permission inheritance and organizational structure expansion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data retrieval, and in particular, to a method for building and retrieving a hierarchical database based on a tree structure. Background Art

[0002] When processing complex structured data stored in the form of an entity-relationship (ER) graph in a database, the complex dependencies between the data make retrieval complicated. In this case, recursive queries are adopted. In a large database, recursive queries are very resource-consuming and the accuracy rate is difficult to guarantee. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for building and retrieving a hierarchical database based on a tree structure, which can convert the complex ER graph relationship into a tree structure, thereby significantly improving the retrieval efficiency, aiming at the deficiencies of the above-mentioned prior art.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] The present invention provides a method for building and retrieving a hierarchical database based on a tree structure, including the following steps:

[0006] S1. Design a data model: Define the dependency relationship between the root node of the tree structure and other nodes;

[0007] S2. Build a hierarchical index: Starting from the top layer of the tree, index the data layer by layer downward to establish a complete hierarchical index structure;

[0008] S3. Data extraction: Use the hierarchical index table to efficiently extract the required data from the database;

[0009] S4. Extraction and retrieval optimization: Use an improved DFS / BFS algorithm with jump pointers to traverse the tree and perform extraction operations, and at the same time analyze the time complexity of the improved DFS / BFS algorithm with jump pointers;

[0010] S5. Data storage and management: Ensure that the extracted data can be efficiently stored and is convenient for subsequent access and query.

[0011] Further, the S1 includes:

[0012] S101. Determine the root node of the tree structure:

[0013] Modeling of the root node:

[0014] Let be the set of all data entities in the database, where n is the number of data entities;

[0015] Let Represents a data entity A set of other data entities that it depends on;

[0016] Root node Is defined as: ;

[0017] S102. Define the dependency relationship between nodes: Except for the root node, each node should have one or more dependent nodes;

[0018] The modeling of the dependency relationship:

[0019] Let Represents the set of child nodes of the dependent data entity ;

[0020] For any non-root node , there must exist at least one such that , where is an element of the set composed of all non-root nodes ;

[0021] For any data entity , there exists , where is an element of the set composed of the data entity ;

[0022] Such that the data entity , indicating that each non-root node is dependent on at least one other node.

[0023] Furthermore, the S2 includes:

[0024] S201. Establish a hierarchical index: Starting from the top-level node of the tree structure, index the data layer by layer downward;

[0025] Recursive index query:

[0026] Let Represent any node of the tree structure;

[0027] Let Represent all child nodes of any node ;

[0028] Let Represent the level of any node ;

[0029] For the root node , set the root node level ;

[0030] For any node and its child nodes , set up child nodes level ;

[0031] S202. Build a hierarchical index table

[0032] Organize the index information into a hierarchical index table for subsequent data extraction;

[0033] For the index table , let each entry be , include:

[0034] representing the node ID;

[0035] representing the node level;

[0036] representing the parent node ID;

[0037] representing the list of child nodes;

[0038] S203. Index retrieval and level clarification:

[0039] Recursive index construction process: Starting from the root node, perform index operations for each node, query the child nodes of the node, assign appropriate levels to each child node, and record the relationship between the node and its corresponding child nodes in the index table.

[0040] Furthermore, the said S3 includes:

[0041] S301. Mathematical expression of the data extraction order:

[0042] The data extraction is carried out in the order from the bottom layer to the upper layer, using the level information in the hierarchical index table to determine the extraction order;

[0043] Define the maximum level:

[0044] Let be the maximum level of all entries in the index table , and the formula is:

[0045] ;

[0046] Level reverse extraction:

[0047] For each level from the maximum level to 0, the node set corresponding to the level is defined as:

[0048] ;

[0049] Extract data layer by layer to ensure that the dependent nodes of each node are extracted before itself;

[0050] S302. Mathematical expression of dependency check:

[0051] Check whether the node dependencies have been extracted:

[0052] The dependency set of each node is defined as:

[0053] ;

[0054] Check whether all dependencies of the node have been extracted, and use an indicator function to represent whether the child node has been extracted:

[0055] ;

[0056] S303. Avoid duplicate extraction:

[0057] Mark and check the extraction status:

[0058] Every time any node is extracted, update the status: ;

[0059] Before preparing to extract any node check:

[0060] .

[0061] Furthermore, in the above S4, the improved DFS / BFS algorithm with jump pointers is specifically as follows:

[0062] Use of jump pointers:

[0063] The jump pointer directly jumps from a certain node to a certain node in the subtree , and spans layers;

[0064] Calculation of the number of jumps:

[0065] The depth of the tree is , and each time it spans layers. When querying the target node, the total number of jumps required is:

[0066] ;

[0067] By using the jump pointer, the time complexity of querying the target node is ;

[0068] For paths that cannot be directly reached by jumping, the number of node accesses is:

[0069] ;

[0070] Among them, represents the total number of nodes in the tree; represents the number of branching factors; as the number of crossed layers increases, the jump pointer can cover more levels, and the number of remaining nodes that need to be accessed layer by layer will decrease;

[0071] After introducing the jump pointer, the total time complexity is expressed as the sum of the complexity of the jumping part and the complexity of the uncovered node part;

[0072] The complexity of the jumping part: The number of jumps is , and the complexity of each jump is O(1), then:

[0073] ;

[0074] The complexity of the uncovered node part: For nodes not covered by jumping, the number of node accesses is , and the complexity of access is:

[0075] ;

[0076] Then, the total time complexity is:

[0077] .

[0078] Furthermore, the S5 includes:

[0079] S501, Data storage structure: Stores node and level information;

[0080] For each any node , use to store the node level , the parent node , and the list of child nodes , and the data storage model is expressed as:

[0081] ;

[0082] Among them, represents the specific data containing the node;

[0083] S502, Index creation and optimization: Creates an index for common query paths;

[0084] If a single level or node type is queried, create an index for the level or the node type:

[0085] ;

[0086] Among them, represents the type or characteristic of a node;

[0087] S503. Optimize the data access path: According to the data access pattern, adjust the physical layout of data storage or the indexing strategy to reduce the data access time;

[0088] Let the data access function be for efficiently accessing any node related data;

[0089] ;

[0090] represents the actual data access operation, including reading data from the storage medium.

[0091] The beneficial effects of the present invention are as follows: Due to the inherent hierarchy and directionality of the tree structure, traversal and retrieval can be made more intuitive and efficient. The tree structure supports efficient hierarchical traversal algorithms, such as depth-first search DFS and breadth-first search BFS, which are very effective for scenarios such as permission inheritance and organizational structure expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 is a flowchart of a method for building and retrieving a hierarchical database based on a tree structure;

[0093] Figure 2 is a diagram of the hierarchy and dependency relationship in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0094] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0095] Please refer to Figure 1 , a method for building and retrieving a hierarchical database based on a tree structure, including the following steps:

[0096] S1. Design the data model: Define the dependency relationship between the root node of the tree structure and other nodes;

[0097] S2. Build a hierarchical index: Starting from the top layer of the tree, index the data layer by layer downward to establish a complete hierarchical index structure;

[0098] The relationship between each layer can be clarified, and the subsequent data extraction process can be optimized.

[0099] S3. Data extraction: Efficiently extract the required data from the database using the hierarchical index table;

[0100] Specifically, this process emphasizes data extraction from the bottom up, ensuring that the upper-layer data is processed only after the lower-layer data it depends on has been extracted, to avoid duplicate extraction and improve extraction efficiency.

[0101] S4. Extraction and retrieval optimization: Use the improved DFS / BFS algorithm with jump pointers to traverse the tree and perform extraction operations, and analyze the time complexity of the improved DFS / BFS algorithm with jump pointers;

[0102] S5. Data storage and management: Ensure that the extracted data can be stored efficiently and is convenient for subsequent access and query.

[0103] The S1 includes:

[0104] Specifically, the root node is the starting point of the tree structure, representing the bottom-layer data, which does not depend on any other data. In the database model, the root node can be regarded as the main entity, and all other entities directly or indirectly depend on this entity. Select those entities that do not depend on other data and on which other data depends as the root node.

[0105] S101. Determine the root node of the tree structure:

[0106] The modeling of the root node:

[0107] Let be the set of all data entities in the database, where is the number of data entities;

[0108] Let represent the set of other data entities that the data entity depends on;

[0109] The root node is defined as: ;

[0110] S102. Define the dependency relationship between nodes: Each node except the root node should have one or more dependent nodes;

[0111] Specifically, these dependency relationships define the hierarchical structure between data, where each upper-layer node depends on one or more lower-layer nodes.

[0112] The modeling of the dependency relationship:

[0113] Let represent the set of child nodes of the dependent data entity ;

[0114] For any non-root node , there must exist at least one such that , where is an element of the set composed of all non-root nodes ;

[0115] For any data entity , there exists , where is an element of the set composed of data entities ;

[0116] such that the data entity indicates that each non-root node is dependent on at least one other node.

[0117] The said S2 includes:

[0118] S201, Hierarchical index establishment: Starting from the top-level node of the tree structure, index data layer by layer downward; implemented through recursive queries, and each query will collect the index information of the current-level nodes and their child nodes.

[0119] Recursive index query:

[0120] Let represent any node of the tree structure;

[0121] Let represent all child nodes of any node ;

[0122] Let represent the level of any node ;

[0123] For the said root node , set the root node level ;

[0124] For any node and its child node , set the level of the child node ;

[0125] S202, Construct a hierarchical index table

[0126] Organize the index information into a hierarchical index table for subsequent data extraction;

[0127] Specifically, the index table includes node identifiers, the levels where they are located, and dependency relationship information. Each row in the table represents a node, including the node ID, level, parent node ID, and list of child nodes;

[0128] For the index table , let each entry be , including:

[0129] representing the node ID;

[0130] representing the node level;

[0131] representing the parent node ID;

[0132] representing the list of child nodes;

[0133] S203. Index retrieval and level clarification:

[0134] Recursive index construction process: Starting from the root node, perform index operations for each node, query the child nodes of the node, assign appropriate levels to each child node, and record the relationship between the node and the corresponding child nodes in the index table.

[0135] Specifically, through the index table, any node and related level and child node information can be quickly accessed, greatly improving the efficiency and accuracy of data retrieval.

[0136] The said S3 includes:

[0137] S301. Mathematical expression of the data extraction order:

[0138] The data extraction is carried out in the order from the bottom layer to the upper layer, and the level information in the hierarchical index table is used to determine the extraction order;

[0139] Define the maximum level:

[0140] Let be the maximum level of all entries in the index table , and the formula is:

[0141] ;

[0142] Level reverse extraction:

[0143] For each level from the maximum level to 0, the node set of the corresponding level is defined as:

[0144] ;

[0145] Extract data layer by layer to ensure that the dependent nodes of each node are extracted before itself;

[0146] S302. Mathematical expression of the dependency check:

[0147] Check whether the node dependencies have been extracted:

[0148] The dependency set of each node is defined as:

[0149] ;

[0150] Check whether all dependencies of the node have been extracted, and use an indicator function to represent whether the child node has been extracted:

[0151] ;

[0152] S303. Avoid duplicate extraction:

[0153] Specifically, during the extraction process, if some data has been extracted at its lower layer, it will not be extracted repeatedly. This is achieved by marking the extracted nodes in the index table. For each node, an extraction status is maintained, and after each data extraction, the extraction status of the node is updated.

[0154] Mark and check the extraction status:

[0155] After any node is extracted, update the status: ;

[0156] Before preparing to extract any node , check:

[0157] .

[0158] Specifically, handle special dependency relationships:

[0159] To prevent data extraction from entering an infinite loop, when the data is in the following three special dependency relationships, the processing method is:

[0160] Processing of lower-layer data dependencies:

[0161] The rule for processing lower-layer data dependencies means that when data depends on lower-layer data, extraction should start directly from the bottom layer upwards. The advantage is that it can effectively utilize the lower-layer data to ensure that when extracting upper-layer data, the relevant lower-layer information has been obtained and is available for use. This rule can effectively manage hierarchical data. For example, in an organizational structure, product classification, or file system, the lower-layer data is usually an important part of the upper-layer data. Extracting the lower-layer data first can ensure the integrity and accuracy of the upper-layer data.

[0162] This rule is applicable to scenarios where data depends on lower-level data and the data is updated frequently. If the underlying data changes frequently while the upper-level data is relatively stable, bottom-up extraction can ensure that the latest underlying data is extracted, and then the upper-level data is updated based on this data, thereby maintaining data consistency and real-time performance.

[0163] The processing rule for lower-level data dependencies can reduce redundant operations. When the underlying data has been effectively extracted, the upper-level data can directly utilize this information, thereby improving extraction efficiency.

[0164] Processing of same-level data dependencies:

[0165] The processing rule for same-level data dependencies is applicable to complex dependency relationships among data at the same level. When extracting data at the same level, if there are dependency relationships among them and they may also form complex dependency relationships with other data, then special steps must be taken to handle these dependency relationships to protect data consistency and integrity.

[0166] This rule manages the dependency relationships among data at the same level. When relevant data needs to be extracted, if they have dependency relationships at the same level and may form complex interconnections with other data, to ensure data consistency and integrity, they must be isolated to avoid potential data conflicts during the extraction process;

[0167] Implementing this rule first requires identifying the dependency relationships among the data. Then, back up the data. Temporary tables can be created or relevant records can be copied to prevent the loss of any important information during the subsequent process. After the backup is completed, cut off the relationships among the data and set the relevant fields to null, thereby ensuring that they no longer depend on each other during extraction. After completing the step of cutting off the relationships, data extraction can be safely performed without affecting other relevant data. After successful data extraction, restore the relationships among the data from the backup and supplement them back to the original state. Finally, perform data verification to ensure the integrity and consistency of the extracted data and confirm that the relationships between the data and other data have been correctly restored.

[0168] This rule can effectively manage the complex dependency relationships among data at the same level through steps of backup, cutting off relationships, and final restoration, improve the security of data extraction, reduce the risk of data conflicts, and ensure the maintenance of data integrity throughout the extraction process.

[0169] Processing of upper-level data dependencies:

[0170] The processing rule for upper-level data dependencies is applicable to complex dependency relationships between upper and lower levels of data. When data depends on upper-level data, the relationships between the upper and lower level lines need to be effectively managed;

[0171] Identify the lines that need to be retrieved and specify the initial node of the line. Specifying the initial node provides a clear starting point for the extraction of the line, ensuring that the integrity of the data hierarchy is maintained in subsequent operations.

[0172] Cut off the relationship between the node and the line. Set the line name field to MULL to avoid conflicts with other node data transmissions during the extraction process. Since the extraction is performed from bottom to top, the lower-layer data cannot be set to NULL, while the upper-layer data can be set to NULL. This can ensure that the lower-layer data remains available during extraction and avoid the loss of node information due to line changes.

[0173] After cutting off the relationship, start retrieving the line name from the bottom layer to ensure that the latest line information is obtained during the extraction process. The data input must follow the order of line first and then node to ensure that the corresponding line can be correctly associated when extracting node data.

[0174] After completing the extraction, verify the integrity and consistency of the extracted line and node data. Ensuring the correct relationship between the line information and the node information is helpful for subsequent data analysis and use. Through this processing method, the dependency relationship between the line and the node can be effectively managed, avoiding potential data conflicts and inconsistencies.

[0175] In S4, the improved DFS / BFS algorithm with jump pointers is specifically as follows:

[0176] In the tree-layer database, by introducing jump pointers, the query time complexity can be significantly reduced.

[0177] Usage of jump pointers:

[0178] The jump pointer directly jumps from a certain node to a certain node in the subtree and spans layers;

[0179] Specifically, when querying the target node, the jump pointer can be used only once in each layer instead of traversing all levels;

[0180] Calculation of the number of jumps:

[0181] The depth of the tree is , and each time it spans layers. When querying the target node, the total number of jumps required is:

[0182] ;

[0183] By using the jump pointer, the time complexity of querying the target node is ;

[0184] Specifically, although skip pointers can directly skip layers, in some cases, it is still necessary to access nodes that are not covered by skip pointers. Assume that skip pointers cover the main layers, and the remaining nodes still need to be accessed layer by layer in the normal way.

[0185] For paths that cannot be directly reached by skipping, the number of node accesses is:

[0186] ;

[0187] where represents the total number of nodes in the tree; represents the number of branching factors; as the number of skipped layers increases, skip pointers can cover more layers, and the number of remaining nodes that need to be accessed layer by layer will decrease;

[0188] After introducing skip pointers, the total time complexity is expressed as the sum of the complexity of the skipped part and the complexity of the uncovered node part;

[0189] The complexity of the skipped part: The number of skips is , and the complexity of each skip is O(1), then:

[0190] ;

[0191] The complexity of the uncovered node part: For nodes not covered by skips, the number of node accesses is , and the complexity of access is:

[0192] ;

[0193] Then, the total time complexity is:

[0194] .

[0195] Specifically, the traditional ER diagram recursive retrieval method:

[0196] In a traditional complex ER diagram structure database, there are multiple associations and dependencies between data. Recursive queries need to traverse the same nodes and relationships multiple times, resulting in low query efficiency. The key points for analyzing the time complexity of recursive queries are:

[0197] Query depth: The depth of a recursive query refers to the number of levels of recursive calls. In some cases, the depth can be very large, especially when dealing with large-scale hierarchical data;

[0198] Number of nodes per layer: This refers to the number of nodes that may be touched at each recursive level. For a very connection-dense data graph, this number may grow rapidly;

[0199] Index efficiency: If the data in the database is well supported by indexes, the efficiency of recursive queries will be significantly improved. Recursive queries without indexes will result in full table scans, greatly increasing the query time.

[0200] Give a specific complexity estimate. The simplified model is as follows:

[0201] Assume that each node has an average of branching factors, and the recursive depth is . If starting from a node, the number of nodes touched in the first layer should be equal to the number of branches, so the first layer will touch nodes, the second layer will touch nodes, until the th layer with nodes. Therefore, the total number of node accesses is:

[0202] ;

[0203] When the number of branching factors , that is, when the number of branches is greater than 1, this sequence is dominated by the last term, so the time complexity is approximately ; is any number from 1 to .

[0204] The time complexity of recursive queries highly depends on the branching factor and the depth of recursion. Therefore, when dealing with large and complex ER diagrams, recursive queries are very resource-consuming. In deep recursions or data-intensive queries, the complexity increases rapidly.

[0205] The comparison between the traditional ER diagram recursive retrieval method and the improved DFS / BFS algorithm with jump pointers is as follows: The complexity of recursive queries in traditional ER diagrams shows exponential growth. For a query with a recursive depth of , the complexity is , which is exponential growth.

[0206] In an actual database, if there are many interrelated entities and relationships in the diagram, the depth of recursive queries may be very large, and the average branching factor per layer may also be large. Therefore, the time complexity increases exponentially; in addition, due to the presence of loops and multiple paths in the graph structure, the same nodes will be accessed multiple times in recursive queries, resulting in extremely low efficiency.

[0207] In contrast, the time complexity of the improved DFS / BFS algorithm with jump pointers is 。The improved DFS / BFS algorithm with jump pointers has a significant improvement in query efficiency. This method is applicable to data storage with clear hierarchical relationships. In the case of large data depth and relatively concentrated access frequencies, it can significantly improve query efficiency.

[0208] S5 includes:

[0209] S501. Data storage structure: storing node and hierarchical information;

[0210] For each and every node , use to store the node level , the parent node , and the list of child nodes . The data storage model is represented as:

[0211] ;

[0212] Wherein, represents the specific data containing the node;

[0213] S502. Index creation and optimization: creating indexes for common query paths;

[0214] If a single level or node type is queried, create an index for the level or the node type:

[0215] ;

[0216] Wherein, represents the type or characteristic of the node;

[0217] S503. Optimize the data access path: adjust the physical layout of data storage or the index strategy according to the data access pattern to reduce the data access time;

[0218] Let the data access function be , which is used to efficiently access the relevant data of any node ;

[0219] ;

[0220] represents the actual data access operation, including reading data from the storage medium.

[0221] Example:

[0222] Consider a specific complex case: an employee management system in a large organization, which includes different departments and sub - departments, as well as employee information under each department. In this case, the relationship between departments and sub - departments can form a complex tree - like hierarchical structure. Each department may have multiple sub - departments, and each sub - department may also have its own sub - departments, and so on. This structure is very common in large enterprises or government agencies, especially those with complex administrative hierarchies. The specific explanations of the hierarchy and dependencies in this case are as follows:

[0223] Please refer to Figure 2 , the root node is the headquarters:

[0224] The root node is the top - most layer of the entire organizational structure. Usually, it is the headquarters or the main administrative unit. It does not depend on any other departments, but all other departments directly or indirectly depend on it. As the main entity, the root node supports the basic structure of the entire organization. In the case of this employee management system, the root node is set as the central part of the company. The headquarters is the top - most layer of the entire organizational structure and is the direct or indirect support and dependency point for all departments and sub - departments.

[0225] The headquarters serves as the root node in this database, and all hierarchical relationships extend from the headquarters. Each department can be located step by step through the root node. This unified hierarchical structure facilitates the orderly management and extraction of department information, ensuring a clear search path.

[0226] The headquarters - level index contains the sub - node information of all levels of departments, such as the first - level, second - level, third - level departments, etc. When it is necessary to query the information of a certain department, the system can start from the headquarters - level index and quickly locate the target department through jump pointers, without the need for layer - by - layer recursion. This index structure makes data access more direct and effectively reduces the process of multiple scans and recursive queries in the database.

[0227] The second layer is the main sub - departments:

[0228] In the example of this employee management system, the second layer represents the first - level departments directly subordinate to the headquarters. These departments are the main functional departments in the organization, responsible for managing specific business areas, such as the Finance Department, Human Resources Department, Technology Department, and Marketing Department. Each main sub - department has its own independent functional scope, but generally accepts the leadership of the headquarters and operates within the strategic and policy framework of the headquarters.

[0229] The second layer directly depends on the root node, that is, each first-level department directly depends on the headquarters. This dependency means that the headquarters can directly exert influence on these first-level departments, including decision-making, resource allocation, and policy guidance. At the same time, the second layer is responsible for managing more subdivided subordinate departments. Each first-level department contains several second-level departments to further refine and implement the functions of the department. For example, the Technology Department may include sub-departments such as the Software Development Department, the Hardware Maintenance Department, and the Network Security Department.

[0230] In the database, the node structure of the second layer is as follows:

[0231] Node information: Each first-level department is designed as a node in the database, recording its department ID, name, and association information with the headquarters, such as the superior department ID;

[0232] Sub-node reference: The first-level department node contains references to its second-level departments, so that all second-level departments under the first-level department can be quickly located through hierarchical indexing.

[0233] Employee data association: Each first-level department records the associated employee information, usually including the basic information of the department head and senior management.

[0234] The third layer and below are even more subdivided sub-departments:

[0235] In the embodiment, the more detailed sub-departments of the third layer and below represent the specific departments or groups under each main sub-department. These more subdivided sub-departments have more specialized functions and responsibilities and are the operation execution layer of an organization, responsible for specific business and project implementation. For example, the Technology Department, which is a main sub-department of the second layer, contains multiple third-layer sub-departments such as the Software Development Department, the Hardware Maintenance Department, and the Network Security Department, and the Software Development Department may be further subdivided into smaller teams such as the Front-End Development Group and the Back-End Development Group.

[0236] In the hierarchical design of the database, the sub-departments of the third layer and below are defined as more detailed node levels, and the information of these nodes in the database includes:

[0237] Department ID: Each sub-department has a unique department ID, which is used to distinguish different departments in the database.

[0238] Department level: Records the level of the department, which is the third layer or lower layer in this embodiment.

[0239] Parent department ID: Points to the superior department of this department. For example, the parent department of the Software Development Department is the Technology Department.

[0240] Sub-department list: If there are smaller sub-departments under this department, it contains references to these sub-departments. For example, the references to the Front-End Development Group and the Back-End Development Group are in the Software Development Department.

[0241] Employee data: Record the basic information of all employees in the department, such as name, position, and contact information. More detailed sub - departments are dependent on their superior departments, forming a hierarchical dependency relationship;

[0242] Direct dependence on superior departments: Sub - departments at the third level and below directly rely on the management and resources of their superior departments. For example, the front - end development group and the back - end development group of the software development department rely on the software development department and receive task assignments and guidance from it.

[0243] Direct dependence on the headquarters and root nodes: Through the transmission of the first - level and second - level departments, these sub - departments indirectly rely on the headquarters. The policies and strategies designated by the headquarters are conveyed through multiple levels of departments and finally implemented at these execution levels.

[0244] Each more detailed sub - department has a clear management relationship, thus ensuring that task and resource allocation are carried out in an orderly manner and are passed down layer by layer to the execution units.

[0245] Comparison of algorithm efficiency in implementation cases:

[0246] When performing data extraction operations, assume that it is necessary to query the employee information of all sub - departments at the third level and below under the technology department. The specific organizational structure is as follows: The technology department is at the second level, that is, the first - level department; there are multiple second - level departments under the technology department, that is, the third level, such as the software development department, the hardware maintenance department, and the network security department. Each second - level sub - department contains multiple third - level sub - departments, such as the front - end development group and the back - end development group, forming more specific teams or groups.

[0247] Traditional recursive query method:

[0248] Using the traditional recursive query method, the system needs to recursively traverse from the root node layer by layer to the bottom - most sub - departments, which will bring many repeated queries. Specifically:

[0249] Query technology department information: Locate the technology department node and query the basic information;

[0250] Recursively query second - level sub - departments: Find the second - level sub - departments of the technology department, such as the software development department, the hardware maintenance department, and the network security department, and then enter each second - level department;

[0251] Recursively query third - level sub - departments: Enter the software development department and then recursively query the third - level sub - departments, such as the front - end development group and the back - end development group;

[0252] Gradually extract employee data: After reaching the bottom - most sub - departments, extract all employee information and return layer by layer to the upper level until all sub - department information has been extracted;

[0253] Starting from the technical department, query all sub - departments layer by layer downwards until the employee information of the third - level and lower - level sub - departments is collected. The number of queries grows exponentially with the depth of the hierarchy and the branching factor at each layer. The time complexity of traditional recursive query is:

[0254] ;

[0255] If we assume , and each department has an average of four sub - departments, then the complexity of recursive query is:

[0256] ;

[0257] This means that in traditional recursive query, the system needs to perform up to 64 query operations to complete the traversal of the technical department and all its sub - departments.

[0258] Using the new method of hierarchical indexing and skip pointers:

[0259] Through the optimized method of hierarchical indexing and skip pointers, the recursive depth of the query can be reduced, and the efficiency can be improved through operations. The specific steps are as follows:

[0260] Directly locate the technical department: Through the hierarchical index of the root node, quickly find the technical department node;

[0261] Use skip pointers to directly access second - level sub - departments: Starting from the technical department node, use skip pointers to directly jump to the second - level sub - departments of the technical department, such as the software development department, the hardware maintenance department, and the network security department.

[0262] Parallel query third - level sub - departments through indexing: The second - level sub - departments of the technical department, such as the software development department, record the indexes of the subordinate third - level sub - departments, such as the front - end development group and the back - end development group. Therefore, the system can directly query all third - level sub - departments in parallel without recursively entering layer by layer.

[0263] Parallel extraction of all employee data: Starting from the technical department, parallelly extract the employee data of each sub - department, summarize and return the final result.

[0264] The time complexity of the method is:

[0265] ;

[0266] That is, the sum of the complexity of the skip part and the complexity of the un - covered node part.

[0267] The system only needs one skip operation at each level. Therefore, the complexity of the skip part is , that is, the linear complexity of the number of levels;

[0268] Assume is the number of nodes not covered. For nodes not directly covered by skip pointers, they still need to be accessed layer by layer, and the complexity of this part is ;

[0269] In this example , and assuming that skip pointers cover most nodes in each layer, take . Then the query complexity will be close to:

[0270] ;

[0271] The comparison results show that the number of queries of the new algorithm is about 7 times, significantly lower than the 64 queries of the traditional algorithm, and the efficiency difference is extremely significant;

[0272] From a mathematical point of view, the time complexity of the traditional recursive algorithm grows exponentially, and the number of queries and time consumption will increase significantly. The time complexity of the new algorithm grows linearly, and the number of accesses to uncovered nodes is reduced by skip pointers . Therefore, the overall query time increases relatively slowly with the increase of the hierarchical depth.

[0273] The above-described embodiments only express the implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the appended claims.

Claims

1. A hierarchical database construction and retrieval method based on a tree structure, characterized in that: The following steps are involved: S1. Design data model: define the dependency relationship between the root node of the tree structure and other nodes; S2. Build a hierarchical index: Starting from the top level of the tree, index the data layer by layer to build a complete hierarchical index structure; S3, data extraction: use hierarchical index tables to efficiently extract required data from the database; S4, extraction and retrieval optimization: Use the improved DFS / BFS algorithm with jump pointer to traverse the tree and perform extraction operations, and analyze the time complexity of the improved DFS / BFS algorithm with jump pointer; S5. Data storage and management: ensure that the extracted data can be stored efficiently and can be easily accessed and queried later; The S3 includes: S301. Mathematical expression of data extraction order: Data is extracted from the bottom layer to the top layer, and the extraction order is determined by using the level information in the level index table; Define the maximum value of the level: set up Is the index table The maximum level of all entries in , the formula is: ; Extract in reverse order: For each level From the maximum level To 0, the node set of the corresponding level is defined as: ; Extract data layer by layer to ensure that the dependent nodes of each node are extracted before itself; S302. Mathematical expression of dependency check: Check whether the node dependencies have been extracted: The set of dependencies for each node is defined as: ; Check if all dependencies of a node have been extracted, using indicator functions to represent child nodes Has it been extracted? ; S303, avoid repeated extraction: Mark and check extraction status: Every time any node After extraction, update the status: ; Before extracting any node Before, check: 。 2. A tree-structured hierarchical database building and retrieval method according to claim 1, characterized in that: The S1 includes: S101, determine the root node of the tree structure: The root node modeling: set up is the set of all data entities in the database, where is the number of data entities; set up Represents data entity A collection of other data entities that it depends on; Root Node Defined as: ; S102. Define the dependency relationship between nodes: each node, except the root node, should have one or more dependent nodes; The dependency modeling: set up Represents dependent data entities The collection of child nodes; For any non-root node , there must be at least one , so that ,in, For all non-root nodes An element that makes up a set; For any data entity ,exist ,in, For data entities An element that makes up a set; Make data entity , indicating that each non-root node is depended on by at least one other node.

3. The tree-structured hierarchical database building and retrieval method according to claim 2, characterized in that: The S2 includes: S201, hierarchical index establishment: starting from the top node of the tree structure, indexing data layer by layer downward; Recursive index query: set up Represents any node of a tree structure; set up Represents any node All child nodes of ; set up Represents any node Level of For the root node , set the root node level ; For any node and child nodes , set the child node Level ; S202: Building a hierarchical index table Organize the index information into a hierarchical index table for subsequent data extraction; for the index table , let each entry be , Include: Represents the node ID; Represents the node level; Represents the parent node ID; Represents a list of child nodes; S203, Index retrieval and hierarchy clarification: Recursive index building process: Starting from the root node, perform indexing operations for each node, query the node's child nodes, assign appropriate levels to each child node, and record the relationship between the node and the corresponding child nodes in the index table.

4. The tree-structured hierarchical database building and retrieval method according to claim 3, characterized in that: In S4, the improved DFS / BFS algorithm with jump pointer is specifically as follows: Use of jump pointer: The jump pointer directly from a node Jump to a node in the subtree , and across layer; Jump count calculation: The depth of the tree is , each time you cross Layer, when querying the target node, the total number of jumps required is: ; By using the jump pointer, the time complexity of querying the target node is ; For paths that cannot be reached directly by jumping, the number of node visits is: ; in, Represents the total number of nodes in the tree; Represents the number of branching factors; as the number of spanned layers increases, the jump pointer can cover more levels, and the number of remaining nodes that need to be visited layer by layer will decrease; After the introduction of jump pointers, the total time complexity is expressed as the sum of the complexity of the jump part and the complexity of the uncovered node part; The complexity of the jumping part: the number of jumps is , the complexity of each jump is O(1), then: ; The complexity of the uncovered node part: For nodes that are not covered by jumps, the number of node visits is , the access complexity is: ; Then, the total time complexity is: 。 5. The tree-structured hierarchical database building and retrieval method according to claim 4, characterized in that: The S5 includes: S501, data storage structure: storage node and level information; For every any node ,use Storage Node Tier , parent node , child node list , the data storage model is expressed as: ; in, Indicates the specific data of the node; S502, index creation and optimization: create indexes for common query paths; If a single level or node type is queried, create an index on that level or node type: ; in, Indicates the type or characteristics of a node; S503, optimizing data access path: adjusting the physical layout of data storage or index strategy according to the data access mode to reduce data access time; Let the data access function be , used to efficiently access any node relevant data; ; Represents actual data access operations, including reading data from storage media.

Citation Information

Patent Citations

  • Multi-column joint storage method based on column storage

    CN110413624A

  • Data reading method and device, electronic equipment, medium and product

    CN118567561A