Data structure optimization method and device for querying IP address location information
By optimizing the data structure of the initial data, including replacement and deduplication, the problem of low efficiency in the binary tree query IP address location information in the prior art is solved, and more efficient query performance and smaller storage space occupation are achieved.
Patent Information
- Application Number
- CN202411979110.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the prior art, the efficiency of querying the location information corresponding to the IP address based on the binary tree is low, especially when the data volume is huge, it occupies a large memory space in the computer system, resulting in low query efficiency.
By optimizing the initial data, including replacing the node information in the non-data node with location information and deduplication of the non-data nodes, reducing the complexity of the data structure and storage space occupancy, thereby improving query efficiency.
The optimized data structure reduces the use of computer storage space, improves the efficiency of querying IP address location information, and reduces the number of times the memory space is accessed during the query process.
Smart Images

Figure CN119377265B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data structure optimization method and device for querying IP address location information. Background Art
[0002] With the development of computer technology, in order to achieve flexible network division, IP addresses are usually classified based on Classless Inter-Domain Routing (CIDR for short).
[0003] An IP address is actually a binary sequence. For example, the IP address (Internet Protocol Address) 127.0.0.1 can be converted into binary form as "01111111.00000000.00000000.00000001".
[0004] CIDR uses slash notation to represent IP address blocks. For example, the " / 8" in 127.0.0.1 / 8 indicates that the length of the network address is 8, that is, in the binary form of the IP address, the first 8 bits are the network address segment, and the last 24 bits are the host address. Therefore, 127.0.0.1 / 8 actually refers to an address block containing 256 IP addresses, and the network address segment of the IP addresses in this address block is "01111111".
[0005] In the process of realizing network communication connection, there is a need to determine the location information corresponding to the IP address, and the location information may be the country, city, and other regions to which the IP address belongs.
[0006] Since each address block has a unique network address segment, a network address segment corresponds to unique regional location information. In order to implement IP address location information query, the data used to query IP address location information is usually stored in the form of a binary tree according to a preset IP address location information query rule.
[0007] For example, the network address segment of the address block 127.0.0.1 / 8 is 11111110, and the location information corresponding to 11111110 is abc. Create the root node of the binary tree, and traverse each binary parameter in 11111110 from left to right. When the binary parameter is 0, create a left child node, and when the binary parameter is 1, create a right child node. After creating the corresponding child node according to the last binary parameter, write abc into the child node. Repeat the above method to write the location information corresponding to multiple network address segments into the binary tree to obtain the target binary tree.
[0008] After determining the target binary tree, the IP address query can be implemented based on the target binary tree. The specific query process is as follows:
[0009] Get the IP address to be queried. If the network address segment of the IP address to be queried consists of a 32-bit binary number, traverse the target binary tree based on the 32-bit binary number. When the binary parameter is 0, walk the left node during the target binary tree query. When the binary parameter is 1, walk the right node during the target binary tree query. After traversing to the last node according to the last binary parameter, read the position information corresponding to the last node to obtain the position information corresponding to the IP address to be queried.
[0010] Based on the above description, when the computer goes to a node corresponding to the target binary tree according to 0 or 1, it needs to access the memory space once during the query process. Therefore, the existing query process needs to access the memory space 32 times. However, when the amount of data corresponding to the target binary tree is huge, it occupies a large amount of memory space of the computer system. The current method of querying the location information corresponding to the IP address based on the target binary tree will result in low query efficiency. Summary of the invention
[0011] Based on this, a data structure optimization method and device for querying IP address location information are provided to solve the problem of low query efficiency of querying location information corresponding to an IP address based on a binary tree in the prior art.
[0012] In a first aspect, a data structure optimization method for querying IP address location information is provided, the method comprising:
[0013] Acquire initial data, the initial data being used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node for storing the location information is a data node, the node for storing two node information is a non-data node, and the node information is used to determine a child node of the non-data node;
[0014] Replacing two of the node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two child nodes of the first target node is the same, and the first position information is the position information in the child node of the first target node; and / or
[0015] The non-data nodes are deduplicated to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
[0016] By using the above method, the data structure of the initial data is optimized, and the volume of the computer's storage space occupied by the initial data is reduced, thereby improving the efficiency of determining the location information of the IP address to be queried.
[0017] In one embodiment, replacing the two node information in the first target node in the initial data with the first location information includes:
[0018] Traverse the initial data and perform the following steps on each non-data node:
[0019] Determine two first nodes according to information of two first nodes in the non-data nodes;
[0020] If the same location information is stored in two of the first nodes, the non-data node is determined as the first target node, the location information in the first node is determined as the first location information, and the two first node information in the first target node are replaced with the first location information.
[0021] Through the above method, the first node information of the non-data node is replaced with the first position information of the first node, thereby achieving the purpose of deleting the two first nodes in the initial data, ensuring that the volume of the initial data can be reduced.
[0022] In one embodiment, before deduplication of the non-data nodes, the method further includes:
[0023] Generating a corresponding node identifier for each of the nodes in the initial data or the first target data;
[0024] and / or,
[0025] After deduplication of the non-data nodes, the method further includes:
[0026] A corresponding node identifier is generated for each of the nodes in the second target data.
[0027] Through the above method, each data node and each non-data node in the initial data or the first template data is marked with a node identifier, ensuring that the nodes in the initial data or the first template data can be quickly located based on the node identifier.
[0028] In one embodiment, the node identifier corresponds to the non-data node within a first preset range;
[0029] The deduplication of the non-data nodes includes:
[0030] Creating a first mapping table, wherein the first mapping table is used to store a mapping relationship between the node information and the node identification group;
[0031] Traverse the initial data or the first target data, and perform the following steps on each of the non-data nodes therein:
[0032] Determine a first child node of the non-data node according to first node information in the non-data node;
[0033] If the first sub-node includes two pieces of second node information, determining the first sub-node as the second target node;
[0034] According to the node identifiers of the two second child nodes of the second target node, a target node identifier group corresponding to the second target node is formed;
[0035] If the target node identification group does not exist in the first mapping table, writing the target node identification group and the first node information used to determine the second target node into the first mapping table as a key-value pair;
[0036] If the target node identification group exists in the first mapping table, the second node information corresponding to the target node identification group is obtained from the first mapping table, and the first node information used to determine the second target node in the non-data node is replaced with the second node information.
[0037] Through the above method, the node identification group corresponding to the non-data node in the initial data or the first target data is deduplicated based on the first mapping table, so that different data are stored in different storage spaces, avoiding access to each storage space during the data query process, thereby ensuring the efficiency of querying location information.
[0038] In one embodiment, the node is further used to store an access identifier, wherein the access identifier is an initial value indicating that the node is in a to-be-accessed state, and the access identifier is a target value indicating that the node is in a visited state;
[0039] Generating a corresponding node identifier for each of the nodes in the initial data, the first target data, or the second target data includes:
[0040] Traversing the initial data or the first target data, and setting the access identifier of each of the nodes to the initial value;
[0041] Traversing the initial data or the first target data, determining the target non-data nodes in the to-be-accessed state in sequence according to the access identifier, setting the first node identifier for the target non-data node, and setting the access identifier of the target non-data node to the target value, wherein the value of the first node identifier increases with the number of the determined target non-data nodes;
[0042] Determine a starting value of the second node identifier of the data node according to the largest first node identifier, wherein the starting value is greater than the largest first node identifier;
[0043] Creating a second mapping table and a target location information array, wherein the second mapping table is used to store a mapping relationship between the node identifier and the location information, and the target location information array is used to store a mapping relationship between the location information and the node identifier;
[0044] Traverse the initial data and perform the following steps for each data node:
[0045] Acquire second location information of the data node;
[0046] If the second location information does not exist in the second mapping table, write the second location information into the target location information array, determine the subscript of the second location information in the target location information array, set the second node identifier of the data node according to the subscript and the starting value, and write the second location information and the corresponding second node identifier into the second mapping table in the form of a key-value pair;
[0047] If the second location information exists in the second mapping table, the second node identifier of the data node is set according to the node identifier corresponding to the second location information in the second mapping table.
[0048] Through the above method, the first node identifier of the non-data node and the second node identifier of the data node are generated, and the starting value is used as the value of the first and second node identifiers, ensuring that the node identifier values of the non-data node and the data node are continuous, which is conducive to improving the accuracy of queries based on initial data.
[0049] In a second aspect, a method for querying IP address location information is provided, the method comprising:
[0050] Get the IP address to be queried;
[0051] Determine the target network address segment according to the IP address to be queried;
[0052] Based on the target network address segment and the target data, the target location information corresponding to the IP address to be queried is determined.
[0053] Through the above method, the target location information of the IP address to be queried is queried based on the network address segment and the target data, thereby ensuring that the query efficiency can be improved.
[0054] In one embodiment, determining the target location information corresponding to the IP address to be queried includes:
[0055] If the IP address to be queried is an IPv4 address, a quick query array is obtained, the quick query array includes node identifiers determined from the target data in sequence according to the representation range of the first byte of the IPv4 address, a target node identifier is determined from the quick query array according to the first byte of the IP address to be queried, and the target location information is determined from a target location information array and the target data according to the target node identifier, the target location information array is used to store a mapping relationship between location information and node identifiers;
[0056] If the IP address to be queried is an IPv6 address, the target location information is determined from the target data according to the IP address to be queried.
[0057] Through the above method, different query methods are selected according to the IP type of the IP address to be queried to determine the target location information, thereby ensuring the accuracy of the target location information.
[0058] In one embodiment, obtaining a fast query array includes:
[0059] Get the representation range of the first byte of the IPv4 address;
[0060] Create a quick query array;
[0061] According to the binary form of each numerical value in the representation range, the corresponding third nodes are determined from the target data in turn, and the node identifiers in the third nodes are written into the fast query array.
[0062] Through the above method, a quick query array storing node identifiers is generated. When the target location information corresponding to the IP address to be queried is determined, the corresponding node identifier can be obtained by accessing the quick query array once, thereby ensuring that the number of times the memory space is accessed can be reduced.
[0063] In one embodiment, determining the target location information from a target location information array and the target data according to the target node identifier includes:
[0064] If the target node identifier is within the first preset range, determine a third target node corresponding to the target node identifier from the target data, and perform a query based on the second byte, the third byte, and the fourth byte of the IP address to be queried, taking the third target node in the target data as the starting point to obtain the target location information;
[0065] If the target node identifier is within a second preset range, the target location information is determined from the target location information array according to the target node identifier.
[0066] Through the above method, the location information of the IP address to be queried is queried based on the fast query array and the target data, thereby ensuring the accuracy of the obtained target location information.
[0067] In a third aspect, the present application provides a data structure optimization device for querying IP address location information, the device comprising:
[0068] An acquisition module, used to acquire initial data, wherein the initial data is used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node used to store the location information is a data node, the node used to store two node information is a non-data node, and the node information is used to determine a child node of the non-data node;
[0069] A replacement module, used for replacing two of the node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two subnodes of the first target node is the same, and the first position information is the position information in the subnode of the first target node;
[0070] A deduplication module is used to deduplicate the non-data nodes to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
[0071] In one embodiment, the replacement module is specifically used to traverse the initial data and perform the following steps for each non-data node: determine two first nodes based on two first node information in the non-data node, if the same position information is stored in the two first nodes, determine the non-data node as the first target node, determine the position information in the first node as the first position information, and replace the two first node information in the first target node with the first position information.
[0072] In one embodiment, the acquisition module or replacement module is specifically used for the data node and the non-data node to also store a node identifier. If the node identifier is within a first preset range, the node identifier corresponds to the non-data node. If the node identifier is within a second preset range, the node identifier corresponds to the data node.
[0073] In one embodiment, the replacement module is also used to, before deduplication of the non-data nodes, also include: generating a corresponding node identifier for each of the nodes in the initial data or the first target data, and / or, after deduplication of the non-data nodes, also include: generating a corresponding node identifier for each of the nodes in the second target data.
[0074] In one embodiment, the replacement module is also used to deduplicate the non-data nodes, including: creating a first mapping table, the first mapping table is used to store the mapping relationship between the node information and the node identification group, traversing the initial data or the first target data, and performing the following steps for each of the non-data nodes therein: determining the first child node of the non-data node according to the first node information in the non-data node, if the first child node includes two second node information, determining the first child node as the second target node, and forming a target node identification group corresponding to the second target node according to the node identifications of the two second child nodes of the second target node, if the target node identification group does not exist in the first mapping table, writing the target node identification group and the first node information used to determine the second target node as a key-value pair into the first mapping table, if the target node identification group exists in the first mapping table, obtaining the second node information corresponding to the target node identification group from the first mapping table, and replacing the first node information used to determine the second target node in the non-data node with the second node information.
[0075] In one embodiment, the replacement module, the node is also used to store an access identifier, the access identifier is an initial value indicating that the node is in a to-be-accessed state, and the access identifier is a target value indicating that the node is in an accessed state. It is also used to generate a corresponding node identifier for each of the nodes in the initial data, the first target data or the second target data, including: traversing the initial data or the first target data, setting the access identifier of each node to the initial value, traversing the initial data or the first target data, determining the target non-data nodes in the to-be-accessed state in turn according to the access identifier, setting the first node identifier for the target non-data node, and setting the access identifier of the target non-data node to the target value, the value of the first node identifier increases with the number of the determined target non-data nodes, and determining the starting value of the second node identifier of the data node according to the largest first node identifier, and the starting value is greater than the largest first node identifier. A node identifier is provided, a second mapping table and a target location information array are created, the second mapping table is used to store the mapping relationship between the node identifier and the location information, the target location information array is used to store the mapping relationship between the location information and the node identifier, the initial data is traversed, and the following steps are performed for each data node: the second location information of the data node is obtained, if the second location information does not exist in the second mapping table, the second location information is written into the target location information array, the subscript of the second location information in the target location information array is determined, and the second node identifier of the data node is set according to the subscript and the starting value, and the second location information and the corresponding second node identifier are written into the second mapping table in the form of a key-value pair, if the second location information exists in the second mapping table, according to the node identifier corresponding to the second location information in the second mapping table, it is set as the second node identifier of the data node.
[0076] In a fourth aspect, the present application provides a device for querying IP address location information, the device comprising:
[0077] The acquisition module is used to obtain the IP address to be queried;
[0078] A determination module, used to determine a target network address segment according to the IP address to be queried;
[0079] The query module is used to determine the target location information corresponding to the IP address to be queried based on the target network address segment and the target data.
[0080] In one embodiment, the query module is specifically used to obtain a quick query array if the IP address to be queried is an IPv4 address, the quick query array including node identifiers determined from the target data in sequence according to the representation range of the first byte of the IPv4 address, determining the target node identifier from the quick query array according to the first byte of the IP address to be queried, and determining the target location information from the target location information array and the target data according to the target node identifier, the target location information array being used to store a mapping relationship between location information and node identifiers, and if the IP address to be queried is an IPv6 address, determining the target location information from the target data according to the IP address to be queried.
[0081] In one embodiment, the query module is also used to obtain the representation range of the first byte of the IPv4 address, create a quick query array, determine the corresponding third node from the target data in turn according to the binary form of each numerical value in the representation range, and write the node identifier in the third node into the quick query array.
[0082] In one embodiment, the query module is also used to determine a third target node corresponding to the target node identifier from the target data. If the target node identifier is within a first preset range, the query is performed based on the second byte, third byte and fourth bytes of the IP address to be queried, with the third target node in the target data as the starting point to obtain the target location information. If the target node identifier is within a second preset range, the location information in the third target node is used as the target location information.
[0083] In one embodiment, the query module is also used to determine a third target node corresponding to the target node identifier from the target data if the target node identifier is within a first preset range, and to perform a query based on the second byte, third byte and fourth byte of the IP address to be queried, taking the third target node in the target data as a starting point to obtain the target location information; if the target node identifier is within a second preset range, determine the target location information from the target location information array based on the target node identifier.
[0084] In a fifth aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data structure optimization method for querying IP address location information of the first aspect and the method for querying IP address location information of the second aspect when executing the computer program.
[0085] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the data structure optimization method for querying IP address location information of the first aspect and the method for querying IP address location information of the second aspect are implemented.
[0086] The above-mentioned data structure optimization method, device, computer equipment and storage medium for querying IP address location information obtains initial data, the initial data is used to store location information in the form of a binary tree according to a preset IP address location information query rule, replaces two node information in the first target node in the initial data with the first location information, obtains the first target data, the location information in the two child nodes of the first target node is the same, the first location information is the location information in the child node of the first target node, and / or deduplicates the non-data nodes to obtain the second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data. The initial data is optimized for data structure, the node information in the non-data node is replaced with the first location information, so as to achieve the purpose of deleting the two child nodes of the non-data node, and the non-data node is deduplicated, which solves the problem that the same data exists in different storage spaces, causing multiple accesses to the storage space, ensures that the number of accesses to the memory space can be reduced, and thus can improve the efficiency of determining the geographical location of the IP address to be queried.
[0087] The above-mentioned method, device, computer equipment and storage medium for querying IP address location information obtain an IP address to be queried, determine a target network address segment according to the IP address to be queried, and determine the target location information corresponding to the IP address to be queried based on the target network address segment and target data. The location information of the IP address to be queried is queried based on the optimized target data, which ensures that the time for querying the target data can be reduced, and the location information is determined based on the network address segment and the target data, which ensures that the efficiency of determining the target location information can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 A schematic diagram of a flow chart of a data structure optimization method for querying IP address location information in one embodiment;
[0089] Figure 2 A schematic diagram of a flow chart of deleting repeated position information in initial data in one embodiment;
[0090] Figure 3 A schematic diagram of a flow chart of a data structure optimization method for querying IP address location information in one embodiment;
[0091] Figure 4 A schematic diagram of a process of deduplicating non-data nodes in one embodiment;
[0092] Figure 5 A schematic diagram of a flow chart of a data structure optimization method for querying IP address location information in one embodiment;
[0093] Figure 6 A flowchart of a method for querying IP address location information in one embodiment;
[0094] Figure 7 A structural block diagram of a data structure optimization device for querying IP address location information in one embodiment;
[0095] Figure 8 A structural block diagram of a device for querying IP address location information in one embodiment;
[0096] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0097] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0098] When querying the location information corresponding to the IP address based on the binary tree, since the binary tree occupies a large storage space of the computer system and requires multiple accesses to the memory space, the efficiency of querying the location information corresponding to the IP address is low. Therefore, in order to improve the efficiency of querying the location information corresponding to the IP address, the embodiment of the present application adopts a method of optimizing the data structure of the initial data, so as to reduce the storage space occupied by the initial data in the computer, and thus query the location information of the IP address based on the optimized initial data.
[0099] The present application provides a data structure optimization method for querying IP address location information, which is used to optimize the data structure and improve the efficiency of querying the geographical location corresponding to the IP address based on a binary tree.
[0100] In one embodiment, Figure 1 As shown, the data structure optimization method includes steps S1 and S2:
[0101] Step S1: Obtain initial data.
[0102] The initial data is used to store the location information in the form of a binary tree according to the preset IP address location information query rule. The preset IP address location information query rule in the embodiment of the present application can be to query the location information of the IP address by traversing the binary tree storing the above initial data, or to query the location information of the IP address by an array storing the above initial data. Since the initial data stores the location information in the form of a binary tree, the location information is stored in the binary tree, and the location information needs to be written into the binary tree. The specific writing process is as follows:
[0103] To obtain the CIDRs of multiple different network address segments, you can obtain the location information corresponding to each CIDR from the operator platform, determine the network address segment of each CIDR, and the binary parameters in each network address segment, create the root node of the initial data, traverse the binary parameters in the network address segment from left to right, and when the binary parameter is 0, create a left child node; when the binary parameter is 1, create a right child node, and after creating the corresponding child node according to the last binary parameter, write the location information corresponding to the network address segment into the child node. Repeat the above process until the location information of all network address segments is written into the corresponding nodes in the initial data.
[0104] The node storing the position information in the above initial data is determined as a data node, and the node storing the node information of two child nodes in the above initial data is determined as a non-data node. The child nodes of the non-data node can be located according to the node information, and the node information can be a pointer.
[0105] The initial data in the embodiment of the present application includes multiple nodes. In order to improve the convenience of query, a node identifier is generated for each node in the initial data. The specific process of generating the first node identifier of the non-data node is as follows:
[0106] An access identifier is stored in each node in the initial data. The access identifier is an initial value that indicates that the node is in a to-be-accessed state, and the access identifier is a target value that indicates that the node has been accessed. All nodes in the initial data can be traversed to set access identifiers for data nodes and non-data nodes in the initial data, and the access identifier of each node is set to an initial value, which can be 0. The specific process of generating the first node identifier of a non-data node is as follows:
[0107] The initial data is traversed, and the target non-data nodes to be accessed are determined in turn according to the access identifiers, a first node identification is set for the target non-data node, and the access identifier of the target non-data node is set to a target value. The value of the first node identification increases with the number of determined target non-data nodes. Thus, when the last target non-data node is traversed, the value of the first node identification of the last target non-data node is the total number of target non-data nodes.
[0108] For example: when the visit identifier is the hasVisit function, the function result of the hasVisit function is "falae" or "true", "falae" means that the node has not been visited, and "true" means that the node has been visited. The function results of all nodes can be set to "falae", and all non-data nodes are traversed. If the function result of the non-data node is "true", the node is not traversed; if the function result of the non-data node is "falae", a first node identifier is marked for the current non-data node. The first node identifier of the first non-data node is 1, and the value of the first node identifier increases with the number of determined non-data nodes until the first node identifier is generated for the last non-data node.
[0109] The specific process of generating the second node identifier of a data node is as follows:
[0110] Create a second mapping table and a target location information array. The second mapping table stores the mapping relationship between the node identifier and the location information. In order to determine the second node identifier of the data node, it is necessary to traverse the initial data, determine the second location information corresponding to each data node in the initial data, and determine whether the second location information exists in the second mapping table.
[0111] If there is second position information in the second mapping table, the node identifier corresponding to the second position information in the second mapping table is set as the second node identifier of the current data node, and the second node identifier is written into the data node. When there is a node identifier in the data node, the node identifier in the second mapping table is used to replace the node identifier stored in the data node; when there is no node identifier in the data node, the second node identifier is directly written into the data node.
[0112] If the second position information does not exist in the second mapping table, the second position information needs to be written into the target position information array. The target position information array is used to store the mapping relationship between the position information and the node identifier. In order to determine the second node identifier, the subscript and starting value of the second position information array need to be determined. The subscript represents the serial number of the data node stored in the second mapping table. The starting value is greater than the total number of target non-data nodes. The starting value is the value of the first second node identifier, and the second node identifier of the data node is incremented based on the starting value. The second position information and the second node identifier are written into the second mapping table in the form of a key-value pair.
[0113] If the serial numbers of the data nodes in the second mapping table increase from 0, the starting value of the data nodes in the second mapping table is the total number of target non-data nodes plus one; if the serial numbers of the data nodes in the second mapping table increase from 1, the starting value of the data nodes in the second mapping table is the total number of target non-data nodes plus one.
[0114] For example: the total number of target non-data nodes is 100. If the subscript of the first data node in the second mapping table is 0, the starting value is 101, and the second node identifier of the data node is 0+101=101. If the subscript of the first data node in the second mapping table is 1, the starting value is 101, and the second node identifier of the data node is 100+1=101.
[0115] Since the starting value of the second identifier of the data node is the total number of target non-data nodes plus one, when the node identifier is within the first preset range, the node identifier corresponds to the non-data node; when the node identifier is within the second preset range, the node identifier corresponds to the data node, the maximum value of the first preset range is the total number of non-data nodes, and the maximum value of the second preset range is the total number of all data nodes and all non-data nodes.
[0116] For example: the first preset range is [1, 100], and the second preset range is [101, 200], which means that there are 100 non-data nodes and 100 data nodes, the first node identifier of the non-data nodes ranges from 1 to 100, the starting value of the data node is 101, and the second node identifier of the data node ranges from 101 to 200.
[0117] Generating the first node identifier of the non-data node and the second node identifier of the data node based on the above method is conducive to quickly finding the corresponding node in the initial data based on the node identifier in the computer, so as to improve the efficiency of optimizing the initial data.
[0118] In one embodiment, step S2 includes:
[0119] Step S21: Replace the two node information in the first target node in the initial data with the first position information to obtain the first target data.
[0120] Since the initial data contains the same location information in different nodes, in order to delete the repeated location information in the initial data, it is necessary to use scheme one to optimize the data structure of the initial data. The specific scheme one is as follows:
[0121] In order to avoid repeatedly visiting the same node in the process of traversing the initial data, it is necessary to adopt a depth-first search, first traversing the child nodes and then traversing the parent nodes, so as to determine the first target node. The first target node is a non-data node. Since the first target node stores the first node information of two first nodes, the corresponding first nodes can be determined respectively according to each first node information, that is, the two child nodes of the first target node are determined. When the position information stored in the two child nodes is the same, the two first node information in the first target node needs to be replaced by the position information of the two child nodes, so as to obtain the first target data.
[0122] Through the above method, the first node storing the same location information in the initial data is deleted, thereby achieving the purpose of deleting duplicate location information. The scheme one reduces the depth of the data nodes in the initial data and reduces the complexity of the initial data, thereby reducing the memory space occupied by the initial data in the computer system, thereby improving the search efficiency.
[0123] The implementation process of Solution 1 can also be:
[0124] After the first target node is determined from the initial data, since the first target node stores the first node information of two first nodes, that is, the first target node stores the first node information of each of the two child nodes, when the two child nodes of the first target node have the same second node identifier, it means that the two child nodes store the same first position information. Therefore, the two first node information in the first target node are replaced by the first position information to obtain the first target data.
[0125] For example: The flowchart for deleting duplicate location information in the initial data is as follows Figure 2 As shown, in Figure 2 There are five nodes in total, namely: node 1, node 2, node 3, node 4 and node 5. The non-data nodes are node 1 and node 2, and the data nodes are node 3, node 4 and node 5. Node 1 stores the node information of two child nodes, namely: node 1 stores the node information of node 2 and node 3. Based on the node information stored in node 1, node 2 and node 3 can be located. Node 2 stores the node information of two child nodes. According to the two node information in node 2, node 4 and node 5 can be located respectively. Node 4 and node 5 store the same location information, and the node identifier of the location information is 8. The node information of node 4 and node 5 stored in node 2 is replaced with the node identifier 8 of the location information. This is only used as an example.
[0126] In the above process of obtaining the first target data, duplicate data nodes are deleted. In order to ensure the continuity of the node identifiers in the first target data, it is necessary to generate the node identifiers of the nodes in the first target data according to the above method of generating the node identifiers of the nodes of the initial data. Since the process of generating the node identifiers of the nodes of the initial data is consistent with the process of generating the node identifiers of the nodes in the first target data, the process of generating the node identifiers of the nodes in the first target data refers to the above process of generating the node identifiers of the nodes of the initial data, which is not elaborated in detail here.
[0127] In another embodiment, step S2 comprises:
[0128] Step S22: De-duplicate non-data nodes in the initial data to obtain second target data.
[0129] Since the information in each node in the initial data occupies a computer storage location, in order to avoid the same information occupying multiple storage locations in the computer system, it is necessary to use the second scheme to optimize the initial data. The specific second scheme is as follows:
[0130] The same node identification group exists in the initial data, and the node identification group is the node identifications of two first nodes of the non-data node, that is, the node identifications of two child nodes of the non-data node.
[0131] For example, the two first child nodes of the non-data node are node A and node B, the node identifier of node A is identifier 1, and the node identifier of node B is identifier 2. At this time, the node identifier group of the non-data node is [identifier 1, identifier 2].
[0132] In order to merge the same node identification groups, it is necessary to create a first mapping table, which is used to store the mapping relationship between node information and node identification groups. After determining the first mapping table, traverse the non-data nodes in the initial data and perform the following steps for each non-data node:
[0133] According to the first node information in the non-data node, the first child node of the non-data node is determined. When two first node information are stored in the first child node, the first child node is determined to be the second target node, and the node identifiers corresponding to the two child nodes of the second target node are determined respectively. According to the node identifiers of the two child nodes, a target node identifier group corresponding to the second target node is formed.
[0134] If the target node identification group exists in the first mapping table, the second node information corresponding to the target node identification group is determined from the first mapping table, and the first node information used to determine the second target node in the non-data node is replaced with the second node information.
[0135] If the target node identification group does not exist in the first mapping table, the target node identification group and the first node information for determining the second target node are written into the first mapping table as a key-value pair, where the key is the target node identification group and the value is the first node information corresponding to the target node identification group.
[0136] According to the above method, all non-data nodes are traversed to obtain the second target data.
[0137] In the above process of obtaining the second target data, the node identification group of non-data nodes is deduplicated. In order to ensure the continuity of the node identification in the second target data, it is necessary to generate the node identification of the nodes in the second target data according to the above method of generating the node identification of the nodes of the initial data. Since the process of generating the node identification of the nodes of the initial data is consistent with the process of generating the node identification of the nodes in the second target data, the process of generating the node identification of the nodes in the second target data refers to the above process of generating the node identification of the nodes of the initial data, which is not elaborated here.
[0138] Based on the above method, by deduplicating non-data nodes in the initial data, the data structure of the initial data is optimized so that different node information can be located at different nodes, so that different node information is stored in different storage locations in the computer.
[0139] When the node information is a pointer, since the non-data node stores pointers to child nodes, when the pointers of two child nodes point to the same node identification group, it is necessary to modify the pointer stored in the non-data node from pointing to the child node to pointing to "null", so that different pointers point to different nodes, so that different storage locations in the computer store different information of the node.
[0140] Through the above method, it is possible to ensure that different nodes in the initial data correspond to only one storage location, and avoid the same node occupying multiple storage locations, thereby reducing the volume of the initial data.
[0141] In one embodiment, Figure 3 As shown, the data structure optimization method includes steps S101-103:
[0142] S101, obtaining initial data;
[0143] S102, replacing two node information in the first target node in the initial data with first position information to obtain first target data;
[0144] S103: De-duplicate non-data nodes in the first target data to obtain second target data.
[0145] The complete implementation steps of steps S101-S102 are shown in the above-described solution 1 to obtain the first target data.
[0146] Since the information in each node in the first target data occupies a computer storage location, in order to avoid the same information occupying multiple storage locations in the computer system, the first target data needs to be optimized. The specific optimization process is as follows:
[0147] The same node identification group exists in the first target data. In order to merge the same node identification group, a first mapping table needs to be created. The first mapping table is used to store the mapping relationship between node information and node identification group. After determining the first mapping table, the non-data nodes in the first target data are traversed, and the following steps are performed for each non-data node:
[0148] According to the first node information in the non-data node, the first child node of the non-data node is determined. When two first node information are stored in the first child node, the first child node is determined to be the second target node, and the node identifiers corresponding to the two child nodes of the second target node are determined respectively. According to the node identifiers of the two child nodes, a target node identifier group corresponding to the second target node is formed.
[0149] If the target node identification group exists in the first mapping table, the second node information corresponding to the target node identification group is determined from the first mapping table, and the first node information used to determine the second target node in the non-data node is replaced with the second node information.
[0150] If the target node identification group does not exist in the first mapping table, the target node identification group and the first node information for determining the second target node are written into the first mapping table as a key-value pair, where the key is the target node identification group and the value is the first node information corresponding to the target node identification group.
[0151] According to the above method, all non-data nodes are traversed to obtain the second target data.
[0152] For example: The process diagram for deduplication of non-data nodes is as follows Figure 4 As shown, in Figure 4In the example, if the first target data is displayed in the form of a binary tree, the identifier corresponding to node 1 is 1, the identifier corresponding to node 2 is 2, the identifier corresponding to node 3 is 3, the identifier corresponding to node 4 is 4, and the identifier corresponding to node 5 is 5. The pointer stored in node 2 points to node 4 and node 5, and the pointer stored in node 3 also points to node 4 and node 5, that is, node 4 and node 5 occupy multiple storage spaces in the computer. In the process of traversing the first target data, node 2 is traversed first, and when traversing to node 3, it is found that there are duplicate nodes. Therefore, it is necessary to modify the pointers pointing to node 4 and node 5 in node 3 to point to "null", and write the pointers pointing to node 4 and node 5 stored in node 2 to the right child node of node 1 (that is, node 3), so that node 4 and node 5 can also be located through the right child node of node 1, thereby achieving the purpose of deleting the storage space occupied by node 4 and node 5 pointed to by node 3. This is only used as an example.
[0153] Based on the above method, by deduplicating non-data nodes in the first target data, the data structure of the first target data is optimized so that different node information can be located at different nodes, so that different node information is stored in different storage locations in the computer.
[0154] In one embodiment, Figure 5 As shown, the data structure optimization method includes steps S201-203:
[0155] S201, obtaining initial data;
[0156] S202, deduplicating non-data nodes in the initial data to obtain second target data;
[0157] S203: Replace two node information in the first target node in the second target data with the first position information to obtain the first target data.
[0158] The complete implementation steps of steps S201-S202 are shown in the above-described scheme 2 to obtain the second target data.
[0159] After the second target data is determined, since the second target data has different nodes containing the same position information, in order to delete the repeated position information in the second target data, the data structure of the second target data needs to be optimized. The specific optimization process is as follows:
[0160] In order to avoid repeatedly visiting the same node during the traversal of the second target data, it is necessary to adopt a depth-first search, first traversing the child nodes and then traversing the parent nodes, so as to determine the first target node. The first target node is a non-data node. Since the first target node stores the first node information of two first nodes, the corresponding first nodes can be determined respectively according to each first node information, that is: determine the two child nodes of the first target node. When the position information stored in the two child nodes is the same, the two first node information in the first target node needs to be replaced by the position information of the two child nodes, so as to obtain the first target data.
[0161] The implementation process of step S203 may also be:
[0162] After the first target node is determined from the second target data, since the first target node stores the first node information of two first nodes, that is, the first target node stores the first node information of each of the two child nodes, when the two child nodes of the first target node have the same second node identifier, it means that the two child nodes store the same first position information. Therefore, the two first node information in the first target node are replaced by the first position information to obtain the first target data.
[0163] In the above process of obtaining the first target data, duplicate data nodes are deleted. In order to ensure the continuity of the node identifiers in the first target data, it is necessary to generate the node identifiers of the nodes in the first target data according to the above method of generating the node identifiers of the nodes of the initial data. Since the process of generating the node identifiers of the nodes of the initial data is consistent with the process of generating the node identifiers of the nodes in the first target data, the process of generating the node identifiers of the nodes in the first target data refers to the above process of generating the node identifiers of the nodes of the initial data, which is not elaborated in detail here.
[0164] Through the above method, the first node storing the same position information in the second target data is deleted, thereby achieving the purpose of deleting duplicate position information, reducing the depth of the data nodes in the second target data, reducing the complexity of the data nodes in the second target data, thereby reducing the memory space occupied by the second target data in the computer system, and thus improving the search efficiency. Based on the above method, the data structure of the initial data is optimized, and the node information in the non-data node is replaced with the first position information, thereby achieving the purpose of deleting the two child nodes of the non-data node, reducing the depth of the initial data during the query process, and deduplicating the node identification group corresponding to the non-data node in the initial data, ensuring that different pointers point to different nodes, solving the problem of the same information occupying multiple storage spaces, reducing the storage space occupied by the initial data, and ensuring that the efficiency of querying based on the initial data can be improved.
[0165] In one embodiment, Figure 6 As shown, after optimizing the initial data to obtain the target data, the location information of the IP address can be queried based on the target data. The present application also provides a method for querying the location information of the IP address to improve the efficiency of querying the location information of the IP address, comprising the following steps:
[0166] Step S301: Obtain the IP address to be queried.
[0167] Get the IP address to be queried, which can be an IPv4 address or an IPv6 address.
[0168] Step S302: Determine the target network address segment according to the IP address to be queried.
[0169] After the IP address to be queried is determined, the network address segment corresponding to the IP address to be queried can be obtained, and the binary parameters of the first byte, the second byte, the third byte and the fourth byte can be obtained from the network address segment.
[0170] Step S303: Determine the target location information corresponding to the IP address to be queried based on the target network address segment and the target data.
[0171] When the IP address to be queried is an IPv4 address, in order to reduce the number of times the storage space is accessed during the process of querying the location information of the IP address to be queried, a fast query array needs to be created. The specific creation process is as follows:
[0172] Obtain an IPv4 address and the first byte of the IPv4 address represented in big-endian format, the first byte representing the representation range of the IPv4 address, traverse the nodes in the target data in sequence according to the binary parameters in the first byte, determine the node to which the binary parameter is traversed from the target data as the third node, if the third node is a data node, write the node identifier corresponding to the data node into a quick query array; if the third node is a non-data node, continue to traverse the next node, and when the third node corresponding to the last binary parameter of the first byte is a non-data node, write the node identifier of the non-data node into the quick query array.
[0173] According to the method described above, the representation ranges of the first bytes of the plurality of IPv4 addresses are sequentially used to determine node identifiers from the target data, and the determined node identifiers are sequentially written into the fast query array.
[0174] According to the correspondence between the first byte in the fast query array and the node identifier, the target node identifier corresponding to the first byte of the IP address to be queried is determined from the fast query array.
[0175] When the target node identifier is within the first preset range, the third target node corresponding to the target node identifier is determined from the target data. The third target node is a non-data node. Since the target location information is not determined in the quick query array, it is necessary to obtain the second byte, the third byte and the fourth byte corresponding to the IP address to be queried. The third target node is used as the starting point in the target data, and traversal is started from the binary parameter in the second byte. If traversal reaches a data node, the location information corresponding to the data node is determined according to the above target location information array, and the location information is determined to be the target location information corresponding to the IP address to be queried; if traversal reaches a non-data node, continue to traverse the next binary parameter.
[0176] When the target node identifier is within the second preset range, since the second location information array stores the location information of all data nodes in the target data and the node identifier of each location information, the corresponding target location information can be determined from the target location information array according to the target node identifier.
[0177] For example: create an empty quick query array called Ipv4FirstByteNodeCacheList, set the array length of Ipv4FirstByteNodeCacheList to 256, obtain 256 IPv4 addresses with different first bytes from the operator database, traverse the target data based on the first byte of each IPv4 address, the binary parameter in the first byte is 0, go to the left node in the target data, the binary parameter is 1, go to the right node in the target data, when traversing to the data node in the target data, write the identifier of the data node to Ipv4FirstByteNodeCacheList; when traversing to a non-data node, traverse the next binary parameter, traverse up to 8 times, the 8th node is still a non-data node, and the node identifier of the non-data node traversed for the 8th time is written into Ipv4FirstByteNodeCacheList.
[0178] IPv4 address: The location information of 127.0.0.1 / 8 is the intranet, the location information of 8.8.8.0 / 24 is the United States, and the location information of 8.8.7.0 / 24 is China. In Ipv4FirstByteNodeCacheList, 127 points to the data node, and the node identifier of the data node is found to be the intranet in the target location information array.
[0179] When the IP address to be queried is an IPv6 address, determine the network address segment, that is, 128-bit binary represented by big endian method, traverse the target data according to the 128-bit binary parameter, if traversing to the data node, determine the location information corresponding to the data node according to the above target location information array, and use the location information as the target location information corresponding to the IP address to be queried; if traversing to a non-data node, continue to traverse the next binary parameter.
[0180] Through the above method, the location information of the IP address to be queried is queried based on the optimized target data. Since the target data occupies a smaller storage space, the time for accessing the storage space during the query process can be reduced, thereby improving the query efficiency. Moreover, when the IP address to be queried is an IPv4 address, the location information can be quickly queried based on the fast query array. Therefore, the number of times the storage space is accessed during the query process can be reduced, thereby improving the query efficiency.
[0181] It should be understood that although Figure 1 , 3 The steps in the flowcharts of , 5, and 6 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 , 3 At least part of the steps in , 5, and 6 may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequentially, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0182] In one embodiment, Figure 7 As shown, a data structure optimization device for querying IP address location information is provided, including: an acquisition module 701, a replacement module 702 and a deduplication module 703, wherein:
[0183] The acquisition module 701 is used to obtain initial data, and the initial data is used to store the location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node used to store the location information is a data node, the node used to store two node information is a non-data node, and the node information is used to determine the child node of the non-data node;
[0184] A replacement module 702 is used to replace the two node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two child nodes of the first target node is the same, and the first position information is the position information in the child node of the first target node;
[0185] The deduplication module 703 is used to deduplicate the non-data nodes to obtain the second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
[0186] The specific definition of the data structure optimization device for querying IP address location information can be found in the definition of the data structure optimization method for querying IP address location information above, which will not be repeated here. Each module in the above-mentioned data structure optimization device for querying IP address location information can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0187] In one embodiment, Figure 8 As shown, a device for querying IP address location information is provided, including: an acquisition module 801, a determination module 402 and a query module 403, wherein:
[0188] The acquisition module 801 is used to acquire the IP address to be queried;
[0189] A determination module 802 is used to determine a target network address segment according to the IP address to be queried;
[0190] The query module 803 is used to determine the target location information corresponding to the IP address to be queried based on the target network address segment and the target data.
[0191] For the specific definition of the query device for IP address location information, please refer to the definition of the query method for IP address location information above, which will not be repeated here. Each module in the above-mentioned query device for IP address location information can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0192] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data structure optimization data for querying IP address location information and query data for IP address location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data structure optimization method for querying IP address location information and a method for querying IP address location information are implemented.
[0193] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0194] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0195] An acquisition module, used to acquire initial data, wherein the initial data is used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node used to store the location information is a data node, the node used to store two node information is a non-data node, and the node information is used to determine a child node of the non-data node;
[0196] A replacement module, used for replacing two of the node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two subnodes of the first target node is the same, and the first position information is the position information in the subnode of the first target node;
[0197] A deduplication module is used to deduplicate the non-data nodes to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
[0198] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0199] Get the IP address to be queried;
[0200] Determine the target network address segment according to the IP address to be queried;
[0201] Based on the target network address segment and target data, the target location information corresponding to the IP address to be queried is determined, and the target data is obtained based on the method described in any one of claims 1-4.
[0202] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0203] An acquisition module, used to acquire initial data, wherein the initial data is used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node used to store the location information is a data node, the node used to store two node information is a non-data node, and the node information is used to determine a child node of the non-data node;
[0204] A replacement module, used for replacing two of the node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two subnodes of the first target node is the same, and the first position information is the position information in the subnode of the first target node;
[0205] A deduplication module is used to deduplicate the non-data nodes to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
[0206] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0207] Get the IP address to be queried;
[0208] Determine the target network address segment according to the IP address to be queried;
[0209] Based on the target network address segment and target data, the target location information corresponding to the IP address to be queried is determined, and the target data is obtained based on the method described in any one of claims 1-4.
[0210] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0211] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0212] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A data structure optimization method for querying IP address location information, characterized in that: include: Acquire initial data, the initial data being used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node for storing the location information is a data node, the node for storing two node information is a non-data node, and the node information is used to determine a child node of the non-data node; Replace the two node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two child nodes of the first target node is the same, and the first position information is the position information in the child node of the first target node; Deduplication of the non-data nodes to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data; The step of replacing the two node information in the first target node in the initial data with the first location information includes: Traverse the initial data and perform the following steps on each non-data node: Determine two first nodes according to information of two first nodes in the non-data nodes; If the same location information is stored in two of the first nodes, the non-data node is determined as the first target node, the location information in the first node is determined as the first location information, and the two first node information in the first target node are replaced with the first location information.
2. The method according to claim 1, characterized in that: Before deduplication of the non-data nodes, the method further includes: Generating a corresponding node identifier for each of the nodes in the initial data or the first target data; and / or, After deduplication of the non-data nodes, the method further includes: A corresponding node identifier is generated for each of the nodes in the second target data.
3. The method according to claim 2, characterized in that The node identifier corresponds to the non-data node within a first preset range; The deduplication of the non-data nodes includes: Creating a first mapping table, wherein the first mapping table is used to store a mapping relationship between the node information and the node identification group; Traverse the initial data or the first target data, and perform the following steps on each of the non-data nodes therein: Determine a first child node of the non-data node according to first node information in the non-data node; If the first sub-node includes two pieces of second node information, determining the first sub-node as the second target node; According to the node identifiers of the two second child nodes of the second target node, a target node identifier group corresponding to the second target node is formed; If the target node identification group does not exist in the first mapping table, writing the target node identification group and the first node information used to determine the second target node into the first mapping table as a key-value pair; If the target node identification group exists in the first mapping table, the second node information corresponding to the target node identification group is obtained from the first mapping table, and the first node information used to determine the second target node in the non-data node is replaced with the second node information.
4. The method according to claim 2, characterized in that: The node is further used to store an access identifier, wherein the access identifier is an initial value indicating that the node is in a to-be-accessed state, and the access identifier is a target value indicating that the node is in a visited state; Generating a corresponding node identifier for each of the nodes in the initial data, the first target data, or the second target data includes: Traversing the initial data or the first target data, and setting the access identifier of each of the nodes to the initial value; Traversing the initial data or the first target data, determining the target non-data nodes in the to-be-accessed state in sequence according to the access identifier, setting the first node identifier for the target non-data node, and setting the access identifier of the target non-data node to the target value, wherein the value of the first node identifier increases with the number of the determined target non-data nodes; Determine a starting value of the second node identifier of the data node according to the largest first node identifier, wherein the starting value is greater than the largest first node identifier; Creating a second mapping table and a target location information array, wherein the second mapping table is used to store a mapping relationship between the node identifier and the location information, and the target location information array is used to store a mapping relationship between the location information and the node identifier; Traverse the initial data and perform the following steps for each data node: Acquire second location information of the data node; If the second location information does not exist in the second mapping table, write the second location information into the target location information array, determine the subscript of the second location information in the target location information array, set the second node identifier of the data node according to the subscript and the starting value, and write the second location information and the corresponding second node identifier into the second mapping table in the form of a key-value pair; If the second location information exists in the second mapping table, the second node identifier of the data node is set according to the node identifier corresponding to the second location information in the second mapping table.
5. A method for querying IP address location information, characterized in that: include: Get the IP address to be queried; Determine the target network address segment according to the IP address to be queried; Based on the target network address segment and target data, the target location information corresponding to the IP address to be queried is determined, and the target data is obtained based on the method described in any one of claims 1-4.
6. The method according to claim 5, characterized in that The determining the target location information corresponding to the IP address to be queried includes: If the IP address to be queried is an IPv4 address, a quick query array is obtained, the quick query array includes node identifiers determined from the target data in sequence according to the representation range of the first byte of the IPv4 address, a target node identifier is determined from the quick query array according to the first byte of the IP address to be queried, and the target location information is determined from a target location information array and the target data according to the target node identifier, the target location information array is used to store a mapping relationship between location information and node identifiers; If the IP address to be queried is an IPv6 address, the target location information is determined from the target data according to the IP address to be queried.
7. The method according to claim 6, characterized in that The obtaining of the quick query array includes: Get the representation range of the first byte of the IPv4 address; Create a quick query array; According to the binary form of each numerical value in the representation range, the corresponding third nodes are determined from the target data in turn, and the node identifiers in the third nodes are written into the fast query array.
8. The method according to claim 6, characterized in that The step of determining the target location information from a target location information array and the target data according to the target node identifier includes: If the target node identifier is within the first preset range, determine a third target node corresponding to the target node identifier from the target data, and perform a query based on the second byte, the third byte, and the fourth byte of the IP address to be queried, taking the third target node in the target data as the starting point to obtain the target location information; If the target node identifier is within a second preset range, the target location information is determined from the target location information array according to the target node identifier.
9. A data structure optimization device for querying IP address location information, characterized in that: The device comprises: An acquisition module, used to acquire initial data, wherein the initial data is used to store location information in the form of a binary tree according to a preset IP address location information query rule, wherein the initial data includes a plurality of nodes, the node used to store the location information is a data node, the node used to store two node information is a non-data node, and the node information is used to determine a child node of the non-data node; A replacement module, used for replacing two of the node information in the first target node in the initial data with the first position information to obtain the first target data, wherein the position information in the two subnodes of the first target node is the same, and the first position information is the position information in the subnode of the first target node; The replacement module is used to traverse the initial data and perform the following steps on each of the non-data nodes: determine two first nodes according to two first node information in the non-data node, if the same position information is stored in the two first nodes, determine the non-data node as the first target node, determine the position information in the first node as the first position information, and replace the two first node information in the first target node with the first position information; A deduplication module is used to deduplicate the non-data nodes to obtain second target data, so that at least one different node information exists in any two non-data nodes in the initial data or the first target data.
10. A device for querying IP address location information, characterized in that: The device comprises: The acquisition module is used to obtain the IP address to be queried; A determination module, used to determine a target network address segment according to the IP address to be queried; A query module is used to determine the target location information corresponding to the IP address to be queried based on the target network address segment and target data, and the target data is obtained based on the device described in claim 9.
11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Information determination method, device and equipment and computer readable storage medium
CN116233064A
Hybrid network system, communication method and network node
US20170187766A1