User behavior data storage optimization method based on approximate neighbor retrieval algorithm

By constructing a weighted undirected graph and generating a maximum spanning tree, the storage path of user behavior data is optimized, which solves the problems of slow user behavior data retrieval and slow system response in smart home systems, and achieves more efficient data retrieval and stable operation.

CN120803371AActive Publication Date: 2025-10-17QINGDAO TAPER ROBOTICS CO LTD +3
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511284798.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-17
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

In smart home systems, the scale of user behavior data continues to grow. Traditional approximate nearest neighbor retrieval algorithms lead to increased read and write overhead between memory and hard disk, affecting retrieval efficiency and system response speed.

Method used

By constructing a weighted undirected graph, using the Kruskal algorithm to generate a maximum spanning tree, and performing parallel computing and depth-first search, the storage sequence of data blocks is optimized to determine the optimal storage path, reducing the number of hard disk reads and network latency.

Benefits of technology

It significantly improves the retrieval speed of user behavior data and the system response speed, and enhances the stability and user experience of the smart home system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803371A_ABST
    Figure CN120803371A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage optimization, in particular to a user behavior data storage optimization method based on an approximate neighbor retrieval algorithm, which specifically comprises the following steps of: acquiring user behavior data through various sensors and intelligent equipment, analyzing data blocks related to the user behavior data in a historical query process, and storing the data blocks into a database; the method comprises the following steps: constructing a historical use matrix of a data block, further constructing a weighted undirected graph according to matrix information, constructing a maximum spanning tree based on the weighted undirected graph by using a Kruskal Krukaran algorithm, generating a longest path of the tree by adopting two times of depth-first search, extracting an optimal data block storage sequence sub-sequence from the longest path, and storing the optimal data block storage sequence sub-sequence in the longest path. And removing related nodes of the sub-sequence from the weighted undirected graph, and repeating searching and removing operations until the weighted undirected graph is empty, so as to generate a complete data block optimal storage sequence. According to the method, by introducing the approximate neighbor retrieval and graph theory optimization algorithm, the storage efficiency of the user behavior data can be improved, and the access delay during query is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data storage optimization, and particularly relates to a user behavior data storage optimization method based on an approximate nearest neighbor search algorithm. BACKGROUND

[0002] With the maturity of technologies such as the Internet of Things and 5G, the smart home industry has entered a stage of vigorous development, and more and more families choose smart homes, driving the rapid growth of the market. However, in order to improve the user experience and increase user stickiness, it is necessary to collect, process and store user behavior data. Under the background of a large population base in China and the continuous use of smart homes by users, a large amount of user behavior data will be generated. How to efficiently and reasonably store these data and improve the speed of terminal reading data has become one of the problems that need to be solved at present. User behavior data is influenced by many factors, so there is a large difference in user behavior data. In the field of smart homes, user behavior data can be represented as a vector containing multiple dimensions, where each dimension represents a different aspect, in order to improve the compactness of storing user data and improve computing efficiency. Facing massive high-dimensional user behavior data, the approximate nearest neighbor algorithm (ANN) can find the possible nearest neighbor (NN) of the target node with high retrieval efficiency under the premise of ensuring retrieval accuracy. As a general algorithm for implementing large-scale data retrieval, ANN has been widely used in various information retrieval research and data mining research such as classification and clustering, and has obvious technical universality.

[0003] But with the overall outbreak of smart home, the size of user behavior data continues to grow, and user behavior data and its index structure cannot be stored in the relatively limited terminal memory at the same time, and only the index structure and a small part of the data can be stored in the terminal memory, and the majority of the data can be stored in the cloud hard disk. The traditional approximate nearest neighbor search algorithm will first represent the multi-modal user behavior data as a plurality of high-dimensional vectors when processing large-scale user data, then select the core data in the high-dimensional vector through clustering and other methods (hereinafter referred to as the head node, and the non-core data is referred to as the non-head node), then use the head node to build an index, then divide the non-head node according to the distance between the head node, and obtain the same number of non-head node data blocks as the head node. The index built by the head node is called head index, and the data block of the non-head node is called hard disk index. After the division is completed, the head index is stored in the terminal memory, and the hard disk index is stored in the cloud hard disk. In the process of finding the target user data, the query statement is represented as a high-dimensional vector in the previous manner, and then a plurality of head nodes are queried in the head index, and the data block corresponding to the head node is loaded into the terminal memory through hard disk reading and writing or network communication. Compared with the pure memory search method, this method significantly relieves the memory pressure and greatly improves the data processing scale, but the read-write overhead between the memory and the hard disk caused by the large-scale data set and the network communication will affect the overall efficiency of the search, and then affect the response speed of the smart home terminal system.

[0004] Therefore, the present application proposes a user behavior data storage optimization method based on approximate nearest neighbor search algorithm to solve the above problems. SUMMARY

[0005] The present application is directed to the deficiencies of the prior art, and a user behavior data storage optimization method based on approximate nearest neighbor search algorithm is developed. The present application can improve the speed and stability of the approximate nearest neighbor search algorithm by determining the storage order of the data block in the hard disk according to the usage frequency, access order and spatiotemporal local features of the data block and generating a new efficient storage sequence, thereby greatly improving the user experience and system reliability of the smart home system.

[0006] The technical scheme for solving the technical problems of the present application is a user behavior data storage optimization method based on approximate nearest neighbor search algorithm, comprising the following steps: S1, collecting the behavior data of the user in different scenes through the sensors and intelligent devices in the smart home system, and then preprocessing the collected behavior data to construct a user behavior database; S2, analyzing the data blocks involved in the historical query process of the user behavior data in the user behavior database in the smart home system, extracting the access frequency of each data block and the correlation between the data blocks, and then constructing a historical usage matrix of the data blocks. S3. Based on the history of the data blocks, a weighted undirected graph is constructed using a matrix. All data blocks are considered nodes in the graph. If two data blocks are read simultaneously in the historical records, an edge is established between the two data blocks. The weight of the edge is determined by the frequency of co-occurrence. A maximum spanning tree is then constructed using the Kruskal algorithm. Parallel computing is used to optimize the calculation process of the maximum spanning tree. S4. Use two depth-first search DFSs on the maximum spanning tree to calculate the longest path of the maximum spanning tree, and extract the optimal data block storage sequence subsequence based on the longest path. Then remove the nodes and edges related to the subsequence from the weighted undirected graph. Repeat the above process multiple times until the weighted undirected graph is empty, and finally form a complete data block optimal storage sequence.

[0007] S1 is as follows: Sensors include temperature sensors, light sensors, motion sensors, and door and window sensors; Smart devices include smart light bulbs, smart speakers, and smart thermostats; Preprocessing operations include data cleaning and normalization operations; The preprocessed data is converted into a high-dimensional vector form, and then a user behavior database is constructed based on the data of the high-dimensional vector.

[0008] S2 is as follows: The historical query information is obtained through the system log of the smart home system. The relevant data blocks are obtained from the historical query information according to the user behavior data. The number of historical query information is expressed as , , the number of data blocks read for each historical query information is expressed as , ,The historical usage matrix of the data block obtained based on the historical query information is expressed as , Representation matrix Medium data block Number.

[0009] The operations for constructing a weighted undirected graph using a matrix based on the history of a data block are as follows: The data block matrix Each data block in Add it as a node to the weighted undirected graph to be constructed, and then select the matrix row by row A bidirectional edge is constructed between any two nodes. When a bidirectional edge already exists between the two points, the weight of the bidirectional edge is automatically increased by 1, and finally a weighted undirected graph is obtained. , Represents a node set, that is, a data block matrix Data blocks in a set of edges, represents the read relationship of two different nodes, represents the weight, that is, the probability of two nodes being read at the same time.

[0010] The operation of constructing the maximum spanning tree by Kruskal algorithm is as follows: The operation of constructing the maximum spanning tree by Kruskal algorithm is as follows: The operation of constructing the maximum spanning tree by Kruskal algorithm is as follows: First, sort all bidirectional edges in the edge set according to the weight, and create an empty set to store the edges of the maximum spanning tree; Then, create a union-find set data structure to manage the connectivity of the vertices in the node set ; Iterate through each edge, and for the current edge , where and belong to the vertex set in the graph, check whether the vertices and belong to the same set, that is, they belong to the same set if there is a loop; If they do not belong to the same set, store the edge in the set and merge the sets of and in the union-find set; If they belong to the same set, skip this edge; When the number of edges in the set is equal to the number of vertices in the weighted undirected graph minus one, all edges in the set together form the maximum spanning tree, that is, the construction of the maximum spanning tree is completed.

[0011] The operation of optimizing the calculation process of the maximum spanning tree is as follows: (1) Use parallel computing to optimize the process of constructing the maximum spanning tree: Each node in the weighted undirected graph is regarded as an independent sub-tree, and the weight of the edge is taken as the opposite, and each node broadcasts the weight of its own edge to its neighbor nodes. If a node receives a smaller edge weight, it updates its own edge weight to a smaller value and continues to broadcast the updated edge weight to the outside; If both nodes hope to be a section of an edge, the two nodes are merged on the edge, and after merging, the sub-trees where the two nodes are located are merged into a larger sub-tree; during the merging process, if a loop is formed, the node deletes the edge with the smallest weight on the loop; When all nodes are added to the spanning tree, and no more edges can be merged, the algorithm terminates, and the final generated tree is the maximum spanning tree of the entire graph; (2) The parallel computing method is used to accelerate the generation of the diameter of the maximum spanning tree: A node is selected as the root node, and the entire maximum spanning tree is divided into multiple sub-trees, each sub-tree containing the root node and its descendant nodes, an MPI message passing interface process is started, the sub-trees are allocated to each process, and a depth-first search is performed in each process to calculate the diameter of the sub-tree, and the start and end points of the sub-tree are recorded; For each process, the calculated sub-tree diameter and start and end points are sent to the root process, and the root process collects the diameters and start and end points of all sub-trees; The root process traverses all the start and end points to find the farthest two points as the start and end points of the diameter of the entire tree, and calculates the path length between them, which is the diameter of the entire tree; The MPI process is ended, and the diameter of the entire tree is returned.

[0012] S4 is as follows: The diameter of the set is obtained by performing two depth-first searches, and the diameter of the set is the longest diameter of the maximum spanning tree; First depth-first search: an initial node is arbitrarily selected , and then a depth-first search is performed to find the farthest node from the node ; Second depth-first search: based on the farthest node from the node , a depth-first search is performed again to find the farthest node from the node ; The longest diameter of the maximum spanning tree is obtained, and the longest diameter of the maximum spanning tree is the path of all nodes and edges from the node to the node ; The nodes contained in the longest diameter of the maximum spanning tree are saved in the order of the tree, which is a sub-sequence of the final storage sequence, and then the nodes and all edges related to the nodes contained in the sub-sequence are deleted from the weighted undirected graph , and a new weighted undirected graph is obtained; According to the new weighted undirected graph, a new maximum spanning tree is generated by Kruskal algorithm, and then the above depth first search operation is performed to obtain a new weighted undirected graph, and the operation is repeated multiple times until the weighted undirected graph is empty, and the final storage sequence is formed by the subsequence obtained by each repeated operation.

[0013] The effects provided in the summary are only the effects of the embodiments, not all the effects of the application, and the above technical solutions have the following advantages or beneficial effects: The application discloses a user behavior data storage optimization method based on an approximate neighbor search algorithm, converts user behavior information into high-dimensional vector data in a large-scale approximate neighbor search algorithm scene, and can efficiently and quickly search for required data of a smart home system by searching for related user information through the approximate neighbor search algorithm, solves problems of slow user behavior data reading speed and low accuracy, and the like, and through optimization of a storage path and parallel computing technology, the system significantly reduces reading times and reading time in the process of reading user vector data from a hard disk and transmitting data from the cloud; meanwhile, a complete smart home field search system can be constructed, the search algorithm is continuously improved, the running speed of a terminal can be further improved, real-time problems can be solved, user experience can be optimized, and the sales of smart home can be increased. The algorithm disclosed by the application greatly reduces the reading time of data in a hard disk in the process of searching for user vector data by the system and network delay in the process of transmitting data from the cloud to the terminal, further solves real-time problems, improves user experience, and can stably operate under high load and complex environment. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and serve to explain the application, and do not constitute a limitation of the application.

[0015] Figure 1 The application is a method flowchart.

[0016] Figure 2 The application is a schematic diagram in actual application. DETAILED DESCRIPTION

[0017] In order to clearly illustrate the technical features of the present application, the application will be described in detail below with reference to the specific embodiments and in conjunction with the accompanying drawings.

[0018] Embodiment 1 As shown in the figure, a user behavior data storage optimization method based on an approximate neighbor search algorithm comprises the following steps: Figure 1 ​S1, collecting behavior data of a user in different scenarios through sensors and smart devices in a smart home system, and then constructing a user behavior database after preprocessing the collected behavior data; S2, analyzing data blocks involved in a historical query process of user behavior data in the smart home system, extracting access frequencies of the data blocks and correlations between the data blocks, and then constructing a historical usage matrix of the data blocks; The access frequency of a data block is the number of times the data block is read in the smart home system, and the correlation between data blocks is the possibility of two data blocks being read simultaneously. The greater the possibility of being read simultaneously, the greater the correlation between the data blocks; S3, constructing a weighted undirected graph according to the historical usage matrix of the data blocks, regarding all data blocks as nodes in the graph, establishing an edge between two data blocks if they are read simultaneously in historical records, and determining the weight of the edge according to the co-occurrence frequency, and then constructing a maximum spanning tree through Kruskal's algorithm and optimizing the calculation process of the maximum spanning tree through parallel computing; S4, calculating the longest path of the maximum spanning tree through two depth-first searches (DFS), extracting the optimal data block storage sequence subsequence according to the longest path, removing the related nodes and edges of the subsequence from the weighted undirected graph, repeating the above process multiple times until the weighted undirected graph is empty, and finally forming a complete optimal data block storage sequence.

[0019] In the specific implementation, S1 is as follows: The sensors include temperature sensors, light sensors, motion sensors, and door and window sensors; The smart devices include smart light bulbs, smart speakers, and smart thermostats; The preprocessing operations include data cleaning and normalization operations; The preprocessed data is converted into a high-dimensional vector form, and then a user behavior database is constructed according to the high-dimensional vector data.

[0020] In the specific implementation, S2 is as follows: The historical query information is obtained through the system log of the smart home system, the relevant data blocks are obtained according to the user behavior data in the historical query information, the number of historical query information is represented as , The number of data blocks read by each piece of historical query information is represented as , The historical usage matrix of the data blocks obtained from the historical query information is represented as , represents the number of data blocks in matrix . ​

[0021] The operations for constructing a weighted undirected graph using a matrix based on the history of a data block are as follows: The data block matrix Each data block in Add it as a node to the weighted undirected graph to be constructed, and then select the matrix row by row A bidirectional edge is constructed between any two nodes. When a bidirectional edge already exists between the two points, the weight of the bidirectional edge is automatically increased by 1, and finally a weighted undirected graph is obtained. , Represents a node set, that is, a data block matrix Data blocks in A collection of Represents an edge set, indicating the read relationship between two different nodes, Represents the weight, that is, the probability that two nodes are read at the same time.

[0022] In a specific implementation, the operation of constructing a maximum spanning tree using the Kruskal algorithm is as follows: Using Kruskal algorithm based on weighted undirected graph Construct a maximum spanning tree; First, the edge set Sort all bidirectional edges in by weight and create an empty set Store the edges of the maximum spanning tree; Then, create a union-find data structure to manage the node set Connectivity of vertices in the Traverse each edge, for the current edge ,in and The set of vertices belonging to the graph , check the vertex and Whether they belong to the same set, if there is a loop, they belong to the same set; If they do not belong to the same set, the edge Save to Collection and merged in the combined search and A collection of If they belong to the same set, skip this edge to avoid forming a loop; When the collection The number of edges in the weighted undirected graph is equal to When the number of vertices in the set is reduced by one, All the edges in together constitute the maximum spanning tree, that is, the construction of the maximum spanning tree is completed.

[0023] In the embodiment, the operation of optimizing the calculation process of the maximum spanning tree is specifically as follows: (1) The process of constructing the maximum spanning tree is optimized by using parallel computing: Each node in the weighted undirected graph is regarded as an independent sub-tree, and the weight of the edge is taken as the opposite, and each node broadcasts the weight of its own edge to the neighbor nodes, if a node receives a smaller edge weight, the weight of its own edge is updated to a smaller value, and the updated edge weight is continuously broadcasted outward; If two nodes both hope to be a part of an edge, the two nodes are merged on the edge, and after the merging, the sub-trees where the two nodes are located are merged into a larger sub-tree; in the merging process, if a loop is monitored, the node deletes the edge with the smallest weight on the loop; When all the nodes are joined into the spanning tree, and there is no more edge to be merged, the algorithm terminates, and the finally generated tree is the maximum spanning tree of the entire graph; (2) The generation of the diameter of the maximum spanning tree is accelerated by using parallel computing: A node is selected as a root node, and the entire maximum spanning tree is divided into multiple sub-trees, each sub-tree containing the root node and its descendant nodes, an MPI (Message Passing Interface) process is started, the sub-trees are allocated to each process, and a depth-first search is performed in each process to calculate the diameter of the sub-tree, and the start point and the end point of the sub-tree are recorded; For each process, the calculated diameter of the sub-tree and the start point and the end point are sent to the root process, and the root process collects the diameters and the start points and the end points of all the sub-trees; The root process traverses all the start points and the end points, finds the two most distant points as the start point and the end point of the diameter of the entire tree, and calculates the path length between them, which is the diameter of the entire tree; The MPI process is ended, and the diameter of the entire tree is returned.

[0024] In the embodiment, S4 is specifically as follows: The diameter of the set is obtained by performing two depth-first searches, and the diameter of the set is the longest diameter of the maximum spanning tree; First depth-first search: an initial node is selected arbitrarily , and then a depth-first search is performed to find the most distant node from the node ; Second depth-first search: based on the most distant node from the node , a depth-first search is performed again to find the most distant node from the node ; Get the longest diameter of the maximum spanning tree, the longest diameter of the maximum spanning tree is the node To Node All nodes and edges of the path; Then save the nodes contained in the longest diameter of the maximum spanning tree in the order of the tree. This order is the subsequence of the final storage sequence. Then, remove the nodes contained in the subsequence and all the edges related to the nodes from the weighted undirected graph. Delete it and get a new weighted undirected graph; Then, according to the new weighted undirected graph, a new maximum spanning tree is generated by Kruskal algorithm, and then the above depth-first search operation is performed to obtain a new weighted undirected graph. Repeat this operation many times until the weighted undirected graph is It is an empty graph, and the subsequences obtained by each repeated operation constitute the final storage sequence Example 2 To better demonstrate the technical effects of this invention, the method of this invention is compared with the existing method IVFPQFS in different scenarios based on data from the user behavior database constructed by this invention. VFPQFS is a large-scale approximate nearest neighbor retrieval algorithm. Its full name is Inverted File System with Product Quantization and Fast Scan. This algorithm combines the inverted file system (IVF), product quantization (PQ), and fast scanning (FS) technologies, and can be used to efficiently process and retrieve large-scale high-dimensional vector data. In the process of constructing vector indexes using the method of the present invention and the IVFPQFS method, the entire vector space is first divided into n_list cluster regions. The number of cluster regions actually searched during the query is n_probe, where n_list represents the number of cluster units and n_probe represents the number of cluster units actually scanned. As can be seen from Table 1, when reading the storage sequence in the MST set of the present invention and the default storage sequence used by the IVFPQFS algorithm, and comparing the two, when obtaining the same data and completing the same query, the number of hard disk reads required to read the storage sequence generated by the present invention is significantly lower than that of the storage sequence generated by IVFPQFS. As the number of reads decreases, the time required to read the data is also greatly reduced. This proves that the method of the present invention can improve the efficiency of large-scale approximate neighbor retrieval algorithms.

[0025] Table 1 Comparison of the effects of the method of the present invention and the existing method Example 3 like Figure 2 As shown, the actual application process of the method in the present invention is as follows: The smart home user behavior data includes smart home single-modal query data (i.e. data to be queried) and smart home multi-modal data (i.e. historical records of the smart home system), the smart home user behavior data is vector represented, the instruction to be queried is "turn off the light and the TV through the mobile phone", the instruction includes "mobile phone", "light" and "TV", therefore the instruction to be queried is a cross-modal query; the multi-modal data in the smart home system includes images, texts, videos and voices; According to the data block involved in the historical records of the smart home system according to the user behavior data contained in the instruction to be queried, a history matrix of randomly stored data is obtained, then a vector index is constructed, a weighted undirected graph is generated, a maximum spanning tree is constructed through Kruskal algorithm, the longest path of the maximum spanning tree is calculated through twice depth-first search (DFS), and the optimal data block storage sequence sub-sequence is extracted according to the longest path, then the related nodes and edges of the sub-sequence are removed from the weighted undirected graph, the above process is repeated for multiple times until the weighted undirected graph is empty, and finally a complete data block optimal storage sequence is formed, i.e. the regularly stored data after rearrangement is obtained, the randomly stored data and the regularly stored data are stored in the computer hard disk; Then the regularly stored data is searched to obtain a vector set composed of TopK query results, and a result set composed of the first K query results is selected.

Claims

1. A user behavior data storage optimization method based on an approximate nearest neighbor retrieval algorithm, characterized in that the steps are: as follows: S1. Collect user behavior data in different scenarios through sensors and smart devices in the smart home system, and then pre-process the collected behavior data to build a user behavior database; S2. Analyze the data blocks involved in the historical query process of user behavior data in the smart home system in the user behavior database, extract the access frequency of each data block and the correlation between data blocks, and then construct a historical usage matrix of the data blocks; S3. Based on the history of the data blocks, a weighted undirected graph is constructed using a matrix. All data blocks are considered nodes in the graph. If two data blocks are read simultaneously in the historical records, an edge is established between the two data blocks. The weight of the edge is determined by the frequency of co-occurrence. A maximum spanning tree is then constructed using the Kruskal algorithm. Parallel computing is used to optimize the calculation process of the maximum spanning tree. S4. Use two depth-first search DFSs on the maximum spanning tree to calculate the longest path of the maximum spanning tree, and extract the optimal data block storage sequence subsequence based on the longest path. Then remove the nodes and edges related to the subsequence from the weighted undirected graph. Repeat the above process multiple times until the weighted undirected graph is empty, and finally form a complete data block optimal storage sequence.

2. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 1 is characterized in that: S1 is as follows: Sensors include temperature sensors, light sensors, motion sensors, and door and window sensors; Smart devices include smart light bulbs, smart speakers, and smart thermostats; Preprocessing operations include data cleaning and normalization operations; The preprocessed data is converted into a high-dimensional vector form, and then a user behavior database is constructed based on the data of the high-dimensional vector.

3. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 2 is characterized in that: S2 is as follows: The historical query information is obtained through the system log of the smart home system. The relevant data blocks are obtained from the historical query information according to the user behavior data. The number of historical query information is expressed as , , the number of data blocks read for each historical query information is expressed as , ,The historical usage matrix of the data block obtained based on the historical query information is expressed as , Representation matrix Medium data block Number.

4. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 3 is characterized in that: The operations for constructing a weighted undirected graph using a matrix based on the history of a data block are as follows: The data block matrix Each data block in Add it as a node to the weighted undirected graph to be constructed, and then select the matrix row by row A bidirectional edge is constructed between any two nodes. When a bidirectional edge already exists between the two points, the weight of the bidirectional edge is automatically increased by 1, and finally a weighted undirected graph is obtained. , Represents a node set, that is, a data block matrix Data blocks in A collection of Represents an edge set, indicating the read relationship between two different nodes, Represents the weight, that is, the probability that two nodes are read at the same time.

5. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 4 is characterized in that: The operations for constructing a maximum spanning tree using the Kruskal algorithm are as follows: Using Kruskal algorithm based on weighted undirected graph Construct a maximum spanning tree; First, the edge set Sort all bidirectional edges in by weight and create an empty set Store the edges of the maximum spanning tree; Then, create a union-find data structure to manage the node set Connectivity of vertices in the Traverse each edge, for the current edge ,in and The set of vertices belonging to the graph , check the vertex and Whether they belong to the same set, if there is a loop, they belong to the same set; If they do not belong to the same set, the edge Save to Collection and merged in the combined search and A collection of If they belong to the same set, skip this edge; When the collection The number of edges in the weighted undirected graph is equal to When the number of vertices in the set is reduced by one, All the edges in together constitute the maximum spanning tree, that is, the construction of the maximum spanning tree is completed.

6. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 5 is characterized in that: The operations for optimizing the calculation process of the maximum spanning tree are as follows: (1) Use parallel computing to optimize the process of constructing the maximum spanning tree: Treat each node in the weighted undirected graph as an independent subtree, and invert the edge weights. At the same time, each node broadcasts its own edge weights to its neighboring nodes. If a node receives a smaller edge weight, it updates its own edge weight to the smaller value and continues to broadcast the updated edge weights. If two nodes both want to be part of an edge, they are merged on this edge. After the merge, the subtrees containing the two nodes are merged into a larger subtree. During the merge process, if a loop is detected, the node will delete the edge with the smallest weight on the loop. When all nodes are added to the spanning tree and there are no more edges to merge, the algorithm terminates and the resulting tree is the maximum spanning tree of the entire graph; (2) Use parallel computing to accelerate the generation of the diameter of the maximum spanning tree: Select a node as the root node and divide the entire maximum spanning tree into multiple subtrees. Each subtree contains the root node and its descendant nodes. Start the MPI message passing interface process, assign the subtree to each process, and perform a depth-first search in each process to calculate the diameter of the subtree and record the starting and ending points of the subtree. For each process, the calculated subtree diameter, starting point, and end point are sent to the root process, and the root process collects the diameters, starting points, and end points of all subtrees; The root process traverses all starting points and end points, finds the two farthest points as the starting point and end point of the diameter of the entire tree, and calculates the path length between them, which is the diameter of the entire tree; End the MPI process and return the diameter of the entire tree.

7. The user behavior data storage optimization method based on the approximate nearest neighbor retrieval algorithm according to claim 6 is characterized in that: S4 is as follows: Get the set by performing two depth-first searches The diameter of the set The diameter is the longest diameter of the maximum spanning tree; First depth-first search: choose an arbitrary starting node , then perform a depth-first search to find the node The farthest node ; Second depth-first search: node-based The farthest node Perform depth-first search again to find the distance node The farthest node ; Get the longest diameter of the maximum spanning tree, the longest diameter of the maximum spanning tree is the node To Node All nodes and edges of the path; Then save the nodes contained in the longest diameter of the maximum spanning tree in the order of the tree. This order is the subsequence of the final storage sequence. Then, remove the nodes contained in the subsequence and all the edges related to the nodes from the weighted undirected graph. Delete it and get a new weighted undirected graph; Then, according to the new weighted undirected graph, a new maximum spanning tree is generated by Kruskal algorithm, and then the above depth-first search operation is performed to obtain a new weighted undirected graph. Repeat this operation many times until the weighted undirected graph is It is an empty graph, and the final storage sequence is composed of the subsequences obtained by each repeated operation.

Citation Information

Patent Citations

  • Approximate nearest neighbor search method combining VP tree and guide nearest neighbor graph

    CN112287185A

  • Field image splicing method based on Delauny algorithm and Kruskal maximum spanning tree algorithm

    CN117173224A

  • Graph embedding method and device based on big data generic construction

    CN119357191A

  • Method of processing relational queries in a database system and corresponding database system

    EP2682878A1

  • System and method for sequencing XML documents for tree structure indexing

    US20060161575A1