Node Search Method, Device, Electronic Device and Storage Medium in Network Diagram

By dynamically adjusting the sampling strategy, the sampling probability of neighbor nodes is optimized, and the problem of high temporal and spatial complexity of alias sampling method in large-scale network graphs is solved, and the sampling efficiency and storage efficiency of random walk sequences are improved.

CN112559847BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011438052.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-10
Publication Date
2025-07-11
Estimated Expiration
2040-12-10

AI Technical Summary

Technical Problem

In the prior art, the alias sampling method has the problem of high spatial and temporal complexity when building Alias Table, especially in large-scale network diagrams, which requires a large amount of storage space, resulting in insufficiency of sampling time and space.

Method used

By adjusting the sampling strategy, dynamically adjust the sampling probability of neighbor nodes, increase the sampling probability of rejected sampling nodes, optimize the sampling efficiency of random walk sequences, and reduce sampling time.

Benefits of technology

The sampling efficiency of random walk sequences is improved, sampling time is reduced, storage demand is reduced, and the efficiency of node search in the network graph is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112559847B_ABST
    Figure CN112559847B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, electronic device, and storage medium for node search in a network graph. The method includes: obtaining a current node and a list of neighbor nodes of the current node in the pre-stored network graph, where the list of neighbor nodes includes at least one neighbor node, and the current node serves as the starting node of a random walk sequence; determining a sampling node of the current node from the list of neighbor nodes based on the sampling strategy of the current node; if the sampling node refuses to be the next-hop node of the current node, adjusting the sampling strategy; determining a new sampling node of the current node from the list of neighbor nodes based on the adjusted sampling strategy until a new sampling node is accepted as the next-hop node of the current node; updating the current node, repeating the above steps, and traversing the network graph to obtain a random walk sequence, thereby effectively improving the sampling efficiency of the random walk sequence and reducing the sampling time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer cloud computing, and in particular, to a method, apparatus, electronic device, and storage medium for searching nodes in a network graph. Background Art

[0002] With the development and popularization of network technology, people increasingly use the network search function to obtain information that meets their needs. When a computer and / or server processes a search request, it is necessary to select nodes in the message network graph to select the target message and push it to the user. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, and storage medium for searching nodes in a network graph to improve the efficiency of random walk sampling.

[0004] According to one aspect of this application, a method for searching nodes in a network graph is provided, including: obtaining a list of the current node and neighbor nodes of the current node in a pre-stored network graph, where the neighbor node list includes at least one neighbor node, and the current node serves as the starting node of a random walk sequence; determining a sampling node of the current node from the neighbor node list based on the sampling strategy of the current node; if the sampling node refuses to be the next-hop node of the current node, adjusting the sampling strategy; determining a new sampling node of the current node from the neighbor node list based on the adjusted sampling strategy until the new sampling node accepts being the next-hop node of the current node; updating the current node, repeating the above steps, and traversing the network graph to obtain the random walk sequence.

[0005] In some embodiments, determining a new sampling node of the current node from the neighbor node list based on the adjusted sampling strategy includes: determining the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle based on the adjusted sampling strategy; determining a new sampling node of the current node from the neighbor node list based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle.

[0006] In some embodiments, determining a new sampling node of the current node from the neighbor node list based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle includes: obtaining the maximum value of the dynamic transition probability of the current node, where the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node accepting or rejecting as the next-hop node of the current node; determining a decision threshold based on the maximum value of the dynamic transition probability, where the decision threshold is used to determine whether the sampling node accepts or rejects as the next-hop node of the current node; randomly selecting a probability value within the sampling range determined by the maximum value of the dynamic transition probability; when the difference between the probability value and the maximum value of the dynamic transition probability is greater than or equal to the decision threshold, it indicates that the sampling node accepts as the next-hop node of the current node; when the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, then based on the sampling strategy of the current node, determine the sampling node of the current node from the neighbor node list.

[0007] In some embodiments, determining the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node includes: determining the sampling effective area of each neighbor node in the neighbor node list; when there is a to-be-clipped sampling effective area exceeding a preset threshold among the sampling effective areas of the neighbor nodes, determining the clipped area region of the to-be-clipped sampling effective area, where the clipped area region is the area region exceeding the preset threshold in the to-be-clipped sampling effective area; adjusting the maximum value of the dynamic transition probability of the current node based on the clipped area region.

[0008] In some embodiments, adjusting the maximum value of the dynamic transition probability of the current node based on the clipped area region includes: obtaining the number of neighbor nodes included in the neighbor node list; adjusting the maximum value of the dynamic transition probability of the current node according to the clipped area region and the number of neighbor nodes.

[0009] In some embodiments, if the sampling node rejects as the next-hop node of the current node, then adjusting the sampling strategy includes: obtaining the sampling effective area corresponding to the sampling node; adjusting the ratio between the sampling effective area of each neighbor node of the current node and the total area of the rectangle based on the sampling effective area corresponding to the sampling node.

[0010] In some embodiments, adjusting the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle based on the sampling effective area corresponding to the sampling node includes: adjusting the maximum dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node, where the maximum dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node accepting or rejecting as the next-hop node of the current node.

[0011] In some embodiments, adjusting the maximum dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node includes: determining the sampling effective area corresponding to the sampling node; obtaining the number of neighbor nodes included in the neighbor node list; and adjusting the maximum dynamic transition probability of the current node based on the number of neighbor nodes and the sampling effective area.

[0012] In some embodiments, if the new sampling node accepts being the next-hop node of the current node, update the current node, and repeat the above steps to traverse the network graph to obtain the end point of the random walk sequence.

[0013] According to another aspect of the present application, there is provided a node search device in a network graph, including: a first acquisition module, configured to acquire a current node and a neighbor node list of the current node stored in advance in the network graph, where the neighbor node list includes at least one neighbor node, and the current node serves as the starting node of a random walk sequence; a determination module, configured to determine a sampling node of the current node from the neighbor node list based on the sampling strategy of the current node; an adjustment module, configured to adjust the sampling strategy if the sampling node rejects being the next-hop node of the current node; a first loop module, configured to determine a new sampling node of the current node from the neighbor node list based on the adjusted sampling strategy until the new sampling node accepts being the next-hop node of the current node; and a second loop module, configured to update the current node, repeat the above steps, and traverse the network graph to obtain the random walk sequence.

[0014] According to another aspect of the present application, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the node search method in the network graph of the present application.

[0015] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the node search method in the network diagram of the electronic device disclosed in the embodiments of the present application.

[0016] Thus, in the node search method in the network diagram of the embodiments of the present application, by adjusting the sampling strategy according to the sampled nodes that are rejected, the sampling probability of the rejected sampled nodes is increased, so as to increase the overall sampling probability of the neighbor nodes of the current node, thereby improving the sampling efficiency of the random walk sequence and reducing the sampling time.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings

[0018] The above and / or additional aspects and advantages of the present invention will become obvious and easily understood from the following description of the embodiments in conjunction with the drawings, where:

[0019] Figure 1 Schematic diagram of the application environment applicable to the node search method in the network diagram provided by the embodiments of the present application;

[0020] Figure 2 Schematic diagram of the principle of the alias sampling method provided by the embodiments of the present application;

[0021] Figure 3 Flowchart of the node search method in the network diagram according to an embodiment of the present application;

[0022] Figure 4 Schematic diagram of an adjustment method for the sampling effective area using the present application;

[0023] Figure 5 Flowchart of the node search method in the network diagram according to another embodiment of the present application;

[0024] Figure 6 Flowchart of the node search method in the network diagram according to yet another embodiment of the present application;

[0025] Figure 7 Flowchart of the node search method in the network diagram according to still another embodiment of the present application;

[0026] Figure 8 Flowchart of the node search method in the network diagram according to another embodiment of the present application;

[0027] Figure 9 Schematic diagram of another adjustment method for the sampling effective area using the present application;

[0028] Figure 10 A flowchart of a method for searching nodes in a network diagram according to another embodiment of the present application;

[0029] Figure 11 A flowchart of a method for searching nodes in a network diagram according to another embodiment of the present application;

[0030] Figure 12 A flowchart of a method for searching nodes in a network diagram according to another embodiment of the present application;

[0031] Figure 13 A flowchart of a method for searching nodes in a network diagram according to a specific embodiment of the present application;

[0032] Figure 14 A flowchart of a method for generating a random walk sequence according to another specific embodiment of the present application;

[0033] Figure 15 A block diagram of a device for a method for searching nodes in a network diagram according to an embodiment of the present application;

[0034] Figure 16 A block diagram of an electronic device for implementing the method for searching nodes in a network diagram according to an embodiment of the present application. Detailed implementation manners

[0035] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0036] Figure 1 A schematic diagram of an application environment applicable to the method for searching nodes in a network diagram provided by an embodiment of the present application. As Figure 1 shown, the method for searching network diagram nodes is applied to the search field. The application environment of this method may include a terminal device 1, a network 2, and a server 3. The search field includes, but is not limited to, news search, image search, music search, etc.

[0037] In some embodiments of the present application, the above-mentioned terminal device 1 may refer to an intelligent device with data calculation and processing functions, including but not limited to smart phones, personal digital assistants, tablet computers, etc. An operating system is installed on the terminal device 1, including but not limited to Android operating system, Symbian operating system, Windows mobile operating system, and Apple iPhone OS operating system, etc. Various application clients are installed on the terminal device 1, such as the application client searched by the user, etc. The terminal device 1 is used to receive the search conditions input by the user, and transmit the search conditions to the server 3 through the network 2 to execute a search related to the search conditions.

[0038] The network 2 may include a wired network and a wireless network. As Figure 1 shown, on the access network side, the terminal device 1 can access the network 2 in a wireless or wired manner, while on the core network side, the server 3 generally accesses the network 2 through a wired network. Of course, the above-mentioned terminal device 1 can also access the network 2 through a wired network, and the server 3 can also access the network 2 through a wireless network.

[0039] The server 3 can be a single server, or a server cluster composed of several servers, or the server 3 can include one or more virtualization platforms, or the server 3 can be a cloud computing service center. The server 3 can also be provided with a database 31, and network data related to the search conditions is stored in the database 31, and the network data can be stored in the form of a network diagram.

[0040] Embodiments of this application perform message search based on a message network graph. Among them, the message network graph is a graph model widely used in the field of message search, usually a second-order graph dynamic random walk, that is, the current node and the previous node are considered during the dynamic random walk process. Among them, the process of sampling the next node from the neighbors of a certain node according to a certain random walk strategy is called a single random walk, and repeating this process multiple times can generate a random walk sequence. After generating the random walk sequence, the messages corresponding to the random walk sequence can be used to generate a message recommendation list in the order of the random walk sequence to send recommendation information to the user. For example, when a user searches for picture information, the server 3 finds the closest target picture as the starting node picture of the random walk sequence according to the picture used by the user for searching or according to the keyword of the picture searched by the user, and then performs a random walk in the picture message network to generate a random walk sequence corresponding to the keyword and the target picture, and uses the picture messages in the random walk sequence to generate a picture message recommendation list and send it to the terminal device 1 to achieve the purpose of message recommendation to the user. Or, when a user searches for news, the server 3 obtains the most matching news information as the starting node news according to the keyword used by the user for searching, and then performs a random walk in the news message network to generate a random walk sequence corresponding to the keyword, and uses the news messages in the random walk sequence to generate a news message recommendation list and send it to the terminal device 1 to achieve the purpose of recommending news to the user.

[0041] Suppose c i represents the i-th node in the random walk sequence, then the distribution function is:

[0042]

[0043] Among them, Z represents the normalization constant, and π vx represents the normalized transition probability value from the previous-hop node v to the current node x, and π vx The calculation formula is π vx =α pq (v, x)·w vx , w vx represents the edge weight from the previous-hop node v to the current node x, and α pq (v, x) represents the dynamic component considering the previous-hop node v to the current node x, and the calculation formula is as follows:

[0044]

[0045] Among them, the nodes that the random walk tends to visit can be controlled to be farther away from the starting node or the visited nodes can be concentrated near the starting node by adjusting the parameters q and p. This application only considers the dynamic random walk in the scenario of an unweighted graph, that is, w vx =1.

[0046] That is to say, traverse each neighbor node t of the current node x, and determine whether the neighbor node t is a neighbor node of the previous hop node v. If the neighbor node t is a neighbor node of the previous hop node v, the transition probability value is 1. If the neighbor node t is the previous hop node v itself, the transition probability value is 1 / p. If the neighbor node t is neither a neighbor node of the previous hop node v nor the previous hop node v itself, the transition probability value is 1 / q.

[0047] The key problem of random walk is how to sample the next node from neighbor nodes, that is, the walking strategy. The dynamic random walk strategy needs to sample the next node by combining the information of the previous node and the current node that have been visited.

[0048] In the related art, the Alias sampling method is usually adopted.

[0049] Among them, for Alias sampling, the transition probability value can be calculated first according to the node connection relationship, and then the Alias Table can be constructed according to the transition probability value.

[0050] Specifically, for Alias sampling, the normalization operation can be first performed on all neighbor nodes of the current node x.

[0051] Since the Alisa method is an optimization method that exchanges space for time, therefore, according to the transition probability value obtained by normalization, the entire probability distribution can be compressed into a 1*N rectangle. For each event (the probability of sampling a certain node), it is converted into the area in the corresponding rectangle as As Figure 2 shown.

[0052] Through the above operations, it is easy to have the probability area of some events greater than 1 and the probability area of some events less than 1. The extra area of the events with the probability area greater than 1 is supplemented to the corresponding events with the area less than 1 to ensure that the area of each small square is 1. At the same time, ensure that each small square stores at most two events.

[0053] In the dynamic random walk scenario, assume that the average number of neighbor nodes of nodes in the message network graph is N. The time complexity required to calculate the transition probability value according to the node connection relationship is 0(N*N). Further, since it is necessary to construct the AliasTable for all edges, assuming the number of edges is E, the time complexity required to construct the Alias Table is 0(E*N*N), and the space complexity is 0(E*N).

[0054] It can be seen that in the related art, the alias sampling method has the problem of high spatio-temporal complexity. During preprocessing, it is necessary to enumerate all directed edges in the network, that is, the directed edges between the current node and all neighbor nodes, to calculate the AliasTable. The size of the network directly determines the size of the storage space. For example, a Twitter dataset contains 41.7M nodes and 2.93B network edges. Storing this network graph requires 22GB of space, while constructing an Alias Table using the alias sampling method requires 980TB of space.

[0055] Based on this, the present application proposes a method, apparatus, electronic device, and storage medium for node search in a network graph to solve the above problems.

[0056] The following describes the method, apparatus, electronic device, and storage medium for node search in a network graph according to embodiments of the present application with reference to the accompanying drawings.

[0057] Figure 3 It is a flowchart of a method for node search in a network graph according to an embodiment of the present application. It should be noted that the execution subject of the method for node search in a network graph in this embodiment is a node search device in the network graph. The node search device in the network graph can be implemented in software and / or hardware. The node search device in this embodiment can be configured in an electronic device or in a server for controlling the electronic device. The server communicates with the electronic device to control it.

[0058] As Figure 4 shown, the method for node search in a network graph may include:

[0059] Step 101, obtain a list of the current node and neighbor nodes of the current node in the pre-stored network graph. The neighbor node list includes at least one neighbor node, and the current node serves as the starting node of the random walk sequence.

[0060] It should be noted that the present application performs node search in a pre-stored network graph. For any node, the neighbor node list is the set of neighbor nodes of the node in the second-order unweighted graph. The neighbor node set may include the identification information of all neighbor nodes of the current node and the transition probability value corresponding to each neighbor node. All neighbor nodes include the previous-hop node of the current node.

[0061] For the current node, when selecting the next-hop node, it can be used as the current starting node of the random walk sequence, that is, randomly walk from the current node to the next-hop node.

[0062] Step 102, determine the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node.

[0063] It should be noted that the sampling strategy is to express the sampling probability histogram corresponding to each neighbor node of the current node. Among them, the abscissa of the histogram is the identification information of the neighbor node coordinates, and the ordinate is the transition probability value of each neighbor node.

[0064] As a feasible embodiment, the first sampling strategy after obtaining the neighbor node list of the current node occurs after adjusting the sampling probability of the previous hop node of the current node.

[0065] Step 103, if the sampling node refuses to be the next hop node of the current node, adjust the sampling strategy.

[0066] Step 104, based on the adjusted sampling strategy, determine a new sampling node of the current node from the neighbor node list until a new sampling node is obtained and accepted as the next hop node of the current node.

[0067] Step 105, update the current node, repeat the above steps, and traverse the network diagram to obtain a random walk sequence.

[0068] That is to say, when selecting the next hop node according to the current node, an initial sampling strategy can be generated according to the situation of the neighbor nodes of the current node, so as to perform the initial sampling according to the initial sampling strategy, and when the sampling node is rejected as the next hop node of the current node, adjust the sampling strategy, and then re - sample until the next hop node of the current node is selected, and then continue to use the next hop node as the current node to select the next hop node until a random walk sequence is obtained.

[0069] Thus, the node search method in the network diagram of the embodiment of the present application realizes the dynamic adjustment of the sampling probability of each neighbor node in the network diagram by adjusting the sampling strategy according to the sampling node whose sampling is rejected, increases the sampling probability of the rejected sampling node, so as to improve the sampling efficiency of the random walk sequence and reduce the sampling time.

[0070] As a feasible embodiment, as Figure 4 and Figure 5 shown, step 102, based on the sampling strategy of the current node, determine the sampling node of the current node from the neighbor node list, including:

[0071] Step 201, determine the sampling effective area of each neighbor node in the neighbor node list.

[0072] It should be noted that in the initial state, in the histogram corresponding to the neighbor node list, the sampling effective area of each neighbor node is the area covered by the intersection of the abscissa and ordinate. The abscissa of the histogram is the identification information of the neighbor node, and the ordinate is the transition probability value of each neighbor node. Therefore, as Figure 5As shown in (a), the multiple vertically rectangular shaded areas are the sampling effective areas of each neighbor node.

[0073] Step 202: When there is a sampling effective area to be cropped in the sampling effective area of the previous hop node of the current node that exceeds the preset threshold, determine the cropping area region of the sampling effective area to be cropped. The cropping area region is the area region in the sampling effective area to be cropped that exceeds the preset threshold.

[0074] It should be noted that the preset threshold can be the maximum value of the sampling effective areas formed by other nodes except the previous hop node among the neighbor nodes of the current node. That is to say, after determining the sampling effective areas of each neighbor node in the neighbor node list, excluding the sampling area of the previous hop node of the current node, identify the largest effective area from other neighbor nodes, and then compare the sampling area of the previous hop node of the current node with the largest effective area. If the sampling area of the previous hop node is greater than the largest effective area, further determine the cropping area region of the sampling effective area to be cropped. If the sampling area of the previous hop node is less than or equal to the largest effective area, there is no need to determine the cropping area region of the sampling effective area to be cropped.

[0075] Among them, the difference area obtained by subtracting the sampling area of the previous hop node from the largest effective area is the cropping area region.

[0076] Step 203: Adjust the maximum value of the dynamic transfer probability of the current node based on the cropping area region.

[0077] It should be noted that when there is a sampling effective area to be cropped in the sampling effective area of the previous hop node that exceeds the preset threshold, since there is only one previous hop node for the current node, therefore, the larger the area to be cropped, that is, the more the transfer probability value of the previous hop node exceeds other neighbor nodes, the larger the invalid area generated in the histogram. Therefore, in this application, after determining the cropping area region, the maximum value of the dynamic transfer probability of the current node is adjusted based on the cropping area region, so as to effectively reduce the invalid area caused by the cropping area region of the previous hop node, reduce the probability that the sampling probability falls into the invalid area and is rejected, and effectively improve the sampling rate. Further, as Figure 6 shown, Step 203: Adjust the maximum value of the dynamic transfer probability of the current node based on the cropping area region, including:

[0078] Step 301: Obtain the number of neighbor nodes included in the neighbor node list.

[0079] It should be understood that since there are no rejected neighbor nodes yet, the number of all neighbor nodes in the neighbor node list can be directly obtained.

[0080] Step 302: Adjust the maximum value of the dynamic transition probability of the current node according to the cropped area region and the number of neighbor nodes.

[0081] Specifically, the cropped area region can be evenly distributed according to the number of neighbor nodes to form a probability interval corresponding to the previous-hop node of the current node, and the probability interval is superimposed on the maximum value of the transition probability in the initial state. That is, the probability interval obtained according to the cropped area region is spliced onto the original list used to represent the neighbor nodes to form a new maximum value of the dynamic transition probability, as shown in (b) of Figure 5 as shown.

[0082] Thus, in this application, by converting the cropped area region into a probability interval corresponding to the cropped area region, the probability area in the extremely high vertical direction is converted into a probability area in the horizontal direction, thereby reducing the invalid area caused by the extremely high probability area, reducing the probability that the sampling probability falls into the invalid area and is rejected, and effectively improving the sampling rate.

[0083] As another feasible embodiment, as shown in Figure 7 Step 103: If the sampling node rejects being the next-hop node of the current node, adjust the sampling strategy, including:

[0084] Step 401: Obtain the sampling effective area corresponding to the sampling node.

[0085] Step 402: Based on the sampling effective area corresponding to the sampling node, adjust the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle.

[0086] That is, when the sampling node is rejected as the next-hop node of the current node, it means that the previous sampling probability falls into the region corresponding to the abscissa of the sampling node, but the corresponding probability value is greater than the transition probability value corresponding to the sampling node. That is, the sampling effective area corresponding to the sampling node is small and prone to generating a large invalid area. Therefore, it is necessary to adjust the sampling effective area corresponding to the sampling node to eliminate the invalid area corresponding to the sampling node, thereby reducing the total area of the rectangle corresponding to the sampling strategy (histogram), increasing the ratio of the sampling effective area in the total area of the rectangle, and effectively improving the sampling efficiency.

[0087] Further, in Step 402, based on the sampling effective area corresponding to the sampling node, adjusting the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle includes: adjusting the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node, and the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node being accepted or rejected as the next-hop node of the current node.

[0088] Specifically, as shown inFigure 8 and Figure 9 As shown in Figure 9 , in step 402, adjusting the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node includes:

[0089] Step 501, determining the sampling effective area corresponding to the sampling node.

[0090] Step 502, obtaining the number of neighbor nodes included in the neighbor node list.

[0091] Step 503, adjusting the maximum value of the dynamic transition probability of the current node based on the number of neighbor nodes and the sampling effective area.

[0092] That is to say, after the sampling node is rejected, obtain the effective area corresponding to the rejected sampling node, and then it is also necessary to obtain the number of neighbor nodes included in the neighbor node list. It should be understood that the neighbor nodes whose sampling is rejected will no longer belong to the neighbor node list because their effective areas need to be adjusted. That is, the number of neighbor nodes included in the currently obtained neighbor node list is the number of neighbor nodes that have not been sampled yet. Then, evenly distribute the sampling effective area according to the number of neighbor nodes that have not been sampled yet to obtain the probability interval of the neighbor nodes whose sampling is rejected, and superimpose the probability interval on the previous maximum value of the dynamic transition probability to obtain the adjusted maximum value of the dynamic transition probability of the current node.

[0093] It should be noted that each time the sampling effective area of the neighbor node whose sampling is rejected is adjusted, it is also necessary to adjust the effective area of the neighbor node whose sampling was rejected previously according to the number of neighbor nodes that have not been sampled yet, so as to ensure that the coverage range of the adjusted probability interval does not exceed the abscissa area covered by the neighbor nodes that have not been sampled yet.

[0094] Among them, Figure 10 after the sampling node e is rejected, adjust its corresponding sampling effective area and stack it above the maximum values of the dynamic transition probabilities of other neighbor nodes (the transition probability values corresponding to c and d), and then after the sampling node a is rejected, adjust its sampling effective area and stack it above the maximum value of the dynamic transition probability of the current other neighbor node (the maximum value of the dynamic transition probability corresponding to e).

[0095] That is, for the neighbor nodes whose sampling is rejected, the same processing method as that of the previous hop node for cropping the sampling effective area is adopted, that is, the effective area in the vertical direction is converted into the effective area in the horizontal direction, and then the effective area is spliced at the top of the abscissa region formed by the neighbor nodes that have not been sampled, forming the probability interval corresponding to the neighbor nodes whose sampling is rejected. Thus, the ineffective area corresponding to the neighbor nodes whose sampling is rejected in the effective sampling strategy is reduced, and the proportion of the effective sampling area in the sampling strategy is increased, thereby effectively improving the sampling efficiency.

[0096] As a feasible embodiment, as Figure 10 shown, step 104, based on the adjusted sampling strategy, determining the new sampling nodes of the current node from the neighbor node list includes:

[0097] Step 601, based on the adjusted sampling strategy, determining the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle.

[0098] It should be noted that the ratio of the sampling effective area corresponding to each neighbor node of the current node to the total area of the rectangle may include the ratio of the area value of the sampling effective area corresponding to the neighbor node in the total area of the rectangle, and also include the ratio of the position of the effective area in the total area of the rectangle.

[0099] Step 602, based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle, determining the new sampling nodes of the current node from the neighbor node list.

[0100] That is, since the sampling strategy (histogram) is adjusted after the sampling node is rejected, if sampling is performed according to the original sampling strategy, volume sampling will be outside the current histogram, that is, the corresponding sampling node cannot be obtained. Therefore, in order to better determine the new sampling nodes of the current node from the neighbor node list, it is necessary to determine the sampling effective area and the total area of the rectangle corresponding to each adjusted neighbor node to ensure that effective sampling nodes can be sampled.

[0101] Furthermore, as Figure 11 shown, based on the adjusted sampling strategy, determining the new sampling nodes of the current node from the neighbor node list includes:

[0102] Step 701, obtaining the maximum value of the dynamic transition probability of the current node, where the maximum value of the dynamic transition probability is used to determine the proportion of the probability of each neighbor node in the neighbor node list of the current node accepting or rejecting as the next hop node of the current node.

[0103] It should be noted that the maximum value of the dynamic transition probability of the current node may include the maximum value of the dynamic transition probability determined when determining the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node in step 102, and also includes the maximum value of the dynamic transition probability formed after adjusting the sampling strategy if the sampling node is rejected as the next-hop node of the current node in step 103.

[0104] Step 702: Determine a decision threshold based on the maximum value of the dynamic transition probability. The decision threshold is used to determine whether the sampling node is accepted or rejected as the next-hop node of the current node.

[0105] Among them, the decision threshold is the maximum value of the transition probability values of other neighbor nodes of the current node except the previous-hop node in the initial state.

[0106] Step 703: Randomly select a probability value within the sampling range determined by the maximum value of the dynamic transition probability.

[0107] For example, mark the maximum value of the dynamic transition probability as dub, and the randomly selected probability value prob satisfies [0, dub].

[0108] Step 704: When the difference between the probability value and the maximum value of the dynamic transition probability is greater than or equal to the decision threshold, it indicates that the sampling node is accepted as the next-hop node of the current node.

[0109] Step 705: When the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, then determine the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node.

[0110] It should be noted that since the maximum value of the dynamic transition probability of the current node may include the maximum value of the dynamic transition probability determined when determining the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node in step 102, and also includes the maximum value of the dynamic transition probability formed after adjusting the sampling strategy if the sampling node is rejected as the next-hop node of the current node in step 103, if the randomly selected probability value is greater than the maximum value of the dynamic transition probability, it means that the random probability must fall into the probability interval adjusted when determining the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node in step 102, or the probability interval formed by sampling the neighbor nodes that are rejected after adjusting the sampling strategy if the sampling node is rejected as the next-hop node of the current node in step 103.

[0111] Therefore, when the randomly selected probability value is greater than the maximum value of the dynamic transition probability, it can be determined that the sampled node corresponding to the randomly selected probability value can be accepted as the next-hop node of the current node. When the randomly selected probability value is less than the decision threshold, based on the sampling strategy of the current node, continue to determine the sampling node of the current node from the neighbor node list.

[0112] Further, if the new sampled node is accepted as the next-hop node of the current node, update the current node, and repeat the above steps to traverse the network graph to obtain the end point of the random walk sequence.

[0113] That is to say, when the sampled node is accepted as the next-hop node of the current node, use the sampled node as the new current node, and then select the next-hop node for the new current node until the end point of the random walk sequence is obtained by traversing the network graph.

[0114] It should be noted that traversing the network graph means traversing the network graph according to the specified number of walk times. When the number of random walks reaches the preset number, that is, when the number of nodes in the random walk sequence reaches the preset number, it is determined that the network graph has been traversed.

[0115] As a feasible embodiment, as Figure 12 shown, the node search method in the network graph further includes:

[0116] Step 801, when the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, randomly select any neighbor node identifier from the list of unsampled neighbor nodes as the sampled node.

[0117] Among them, the list of unsampled neighbor nodes includes all neighbor nodes in the neighbor node list of the current node x except the rejected neighbor nodes.

[0118] In the embodiment of the present application, assume that N is the neighbor node list of the current node x, V is the list of neighbor nodes that have been sampled currently, that is, the list of rejected neighbor nodes, Nu is the number of neighbor nodes that have not been sampled currently, Nv is the number of neighbor nodes that have not been sampled currently, that is, the number of nodes in N\V, prob is the probability value, and t is the sampled node t randomly selected from N\V, that is, the probability value corresponding to the sampled node t is prob.

[0119] Step 802, determine whether the probability value is less than the first preset transition probability.

[0120] Step 803, if the probability value is less than the first preset transition probability, determine that the sampled node corresponding to the probability value is the next-hop node.

[0121] Among them, the first preset transition probability is 1b, and 1b = min(1, 1 / p, 1 / q).

[0122] That is to say, if the probability value is small enough, less than the minimum value among the three transition probability values (1, 1 / p, 1 / q) corresponding to the current node x, then this probability will surely meet the acceptance probability under any condition. That is, regardless of whether the sampling node and the previous-hop node are neighbors, the sampling node can be accepted (sampled). Therefore, the sampling node x corresponding to this probability value prob is the next-hop node. In other words, if prob < 1b, then the sampling node x is the next-hop node.

[0123] Step 804: When the probability value is greater than or equal to the first preset transition probability, determine whether the sampling node is a neighbor node of the previous-hop node.

[0124] Step 805: If the sampling node is a neighbor node of the previous-hop node and the probability value is less than the second preset transition probability, determine that the sampling node is the next-hop node.

[0125] Wherein, the second preset transition probability is 1.

[0126] Step 806: If the neighbor node is not a neighbor node of the previous-hop node of the current node and the post-selection transition probability is less than the third preset transition probability, determine that the sampling node is the next-hop node.

[0127] Wherein, the third preset transition probability is 1 / q.

[0128] That is to say, when the probability value meets 1b < prob < ub, it is necessary to determine the acceptance probability according to the relationship between the sampling node x corresponding to the probability value prob and the previous-hop node x of the current node x. Among them, the acceptance probability is the maximum probability value at which the sampling node can be sampled. That is, when the probability value prob is greater than 0 and less than the acceptance probability, it is considered that the sampling node corresponding to this probability value is the next-hop node.

[0129] Specifically, when the sampling node x is a neighbor node of the previous-hop node t, the acceptance probability is 1. That is, if the probability value prob of the sampling node x < 1, then determine that the sampling node is the next-hop node. When the sampling node x is not a neighbor node of the previous-hop node t, the acceptance probability is 1 / q. That is, if the probability value prob of the sampling node x < 1 / q, then determine that the sampling node is the next-hop node.

[0130] It should be understood that if the difference between the probability value prob and the maximum value of the dynamic transition probability is less than the decision threshold and does not meet the condition of the acceptance probability, then the sampling node corresponding to this probability value is a rejected neighbor node, that is, a node that is not sampled, and it is necessary to adjust the transition probability value of this node to obtain the corresponding transition probability value interval.

[0131] As a feasible embodiment, the method for searching nodes in a network diagram further includes: adding the next-hop node of the current node to the node search set, identifying that the number of nodes in the node search set reaches a preset number, and using the node search set as a random walk sequence for search recommendation.

[0132] That is, for a second-order unweighted graph used for message search, the random walk between nodes can be stopped by setting a threshold for the number of walks. Specifically, the number of walks corresponding to the current node can be judged. If the number of walks is less than the second preset number threshold, the current node is continued to be used for random walk to select the next-hop node. If the number of walks is equal to the second preset number threshold, all nodes from the starting node to the current node are generated into a message search walk sequence according to the walk order. So as to generate a recommended message for search according to the message search walk sequence.

[0133] As a specific embodiment, as Figure 13 shown, the method for searching nodes in a network diagram includes the following steps:

[0134] Step 901, initialize auxiliary variables ub, 1b, theta, delta, let Nu = |N|, Nv = 0, ub = max(1, 1 / q), 1b = min(1, 1 / p, 1 / q), theta = min(1, 1 / q), delta = max(0, 1 / p - ub).

[0135] Wherein, N is the list of neighbor node information of the current node x, V is the list of sampled neighbor nodes, Nu is the number of currently unsampled neighbor nodes, Nv is the number of sampled neighbor nodes, v0 is the starting node of the message search walk sequence, S is the message search walk sequence, i is the number of walks, l is the second preset number threshold, that is, the maximum walk length, and v is the previous-hop node of the current node x.

[0136] Step 902, judge whether the number of walks i of the current node reaches the second preset number threshold l.

[0137] If so, output the message search walk sequence S; if not, execute step 503.

[0138] Step 903, judge whether the number of walks is 0.

[0139] If so, execute step 508; if not, execute step 504.

[0140] Step 904, calculate the current maximum transition probability value dub = ub + delta / Nu + Nv * theta / Nu, and generate a uniform random number prob from [0, dub].

[0141] If prob - ub ≥ delta / Nu, accept and return node V[(prob - ub - delta / Nu) * Nu / theta], that is, if prob is greater than the maximum value of the initial transition probability value, determine the identification information of the next-hop node according to the interval of the transition probability value where prob is located.

[0142] Specifically, accepting and returning means taking the candidate node corresponding to prob as the next-hop node, inputting the node identification into S, and at the same time, returning the entire scheme to step 801 to sample the next-hop node with the selected next-hop node as the current node.

[0143] If prob < ub, execute step 805; if prob - ub < delta / Nu, accept and return the previous-hop node, that is, the interval of the transition probability value corresponding to prob - ub < delta / Nu is the adjusted transition probability value interval of the previous-hop node with discrete transition probability.

[0144] Step 905, randomly select a point t from N / V. If prob < 1b, accept and return node x, otherwise, execute step 806.

[0145] Step 906, determine whether node t is a neighbor node of node v. If node t is a neighbor node of node v, record - prob = 1, and determine whether prob < 1. If so, accept and return node t. If not, execute step 807. If node t is not a neighbor node of node v, record - prob = 1 / q, and determine whether prob < 1 / q. If so, accept and return t. If not, execute step 807.

[0146] Step 907, if - prob = theta, add t to V, update Nu = Nu - 1, Nv = Nv + 1, and return to step 8802.

[0147] Step 908, randomly select a node from the neighbor nodes of V0.

[0148] As Figure 14 shown, the random walk sequence generation method includes the following steps:

[0149] Step 1001, configure the parameters related to dynamic random walk. Among them, the related parameters include the parameters p and q used to express the distance between the accessed node and the starting node, the maximum walk length l of the random walk sequence, and the starting node v0 of the random walk sequence.

[0150] Step 1002, initialize the random walk sequence S as empty, add the starting node v0 to the random walk sequence S, and initialize the iteration count i = 0.

[0151] Step 1003, determine whether i < l?

[0152] If yes, execute Step 1004; if no, output the random walk sequence S.

[0153] That is to say, when the number of walk iterations for node selection in the network graph reaches the maximum walk length l of the preset random walk sequence, the random walk sequence S is output.

[0154] Step 1004, determine whether i = 0?

[0155] If yes, execute Step 1005; if no, execute Step 1006.

[0156] Step 1005, randomly select a point v from the neighbor node set of the starting node v0, and execute Step 1007.

[0157] That is to say, for the nodes in the second hop, they can be directly randomly selected through the neighbor node set of the starting node v0, without the need for complex random probability calculations or adjustments.

[0158] Step 1006, combine the previous hop node and the neighbor nodes of the current node, and select a point v from the neighbor node list according to p and q in the relevant parameters, and execute Step 1007.

[0159] It should be noted that Step 1005 can adopt the node search method in the network graph proposed in this application, and the selected point v is the next hop node obtained based on the node search method in the network graph proposed in this application.

[0160] Step 1007, add the selected node v to the set Si = i + 1, vi = v.

[0161] Thus, by adopting the node search method in the network graph proposed in the embodiments of this application, the sampling speed of the next hop node can be accelerated, thereby effectively improving the acquisition speed of the random walk sequence, improving the speed of message recommendation to users while ensuring the reliability of message recommendation, and improving the user experience.

[0162] Furthermore, the applicant also verified the solution proposed in this application.

[0163] For example, fix p = 0.25, q = 1, the walk length l is 80, execute 50 epochs, and conduct experiments on three datasets, namely cora, wiki, and blogCatalog. The results are as follows in the table:

[0164] Table 1

[0165]

[0166] It can be seen that, when the data optimized and implemented according to the embodiments of the present application is compared with the data originally implemented by the original method, since the outlier area is cropped and evenly distributed to other nodes, the sampling hit rate can be significantly improved, and the running time can be reduced by about 48% - 66%.

[0167] For another example, when p = 1 is fixed, the walking length is 80, and 50 epochs are executed, experiments are carried out on three datasets of cora, wiki, and blogCatalog respectively, and the results are as follows in the table:

[0168] Table 2

[0169]

[0170] It can be seen that, when the data optimized and implemented according to the embodiments of the present application is compared with the data originally implemented by the original method, the method proposed in the present application generally requires fewer sampling times, so the running time is reduced by about 4% - 95%. Among them, the degree of reduction in the sampling times is related to the network structure information in the specific dataset.

[0171] To implement the above embodiments, the present invention also proposes a node search device in a network graph.

[0172] Figure 15 It is a block diagram of a node search device in a network graph provided for an embodiment of the present invention. As Figure 15 shown, the node search device 10 in the network graph includes:

[0173] A first acquisition module 11, configured to acquire a list of the current node and the neighbor nodes of the current node stored in advance, the neighbor node list includes at least one neighbor node, and the current node serves as the starting node of the random walk sequence.

[0174] A determination module 12, configured to determine a sampling node of the current node from the neighbor node list based on the sampling strategy of the current node.

[0175] An adjustment module 13, configured to adjust the sampling strategy if the sampling node refuses to be the next-hop node of the current node.

[0176] A first loop module 14, configured to determine a new sampling node of the current node from the neighbor node list based on the adjusted sampling strategy until a new sampling node is accepted as the next-hop node of the current node.

[0177] A second loop module 15, configured to update the current node, repeat the above steps, and traverse the network graph to obtain a random walk sequence.

[0178] Further, the first loop module 14 is further configured to determine, based on the adjusted sampling strategy, the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle; and determine new sampling nodes of the current node from the neighbor node list based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle.

[0179] Further, the adjustment module 13 is further configured to: obtain the maximum value of the dynamic transition probability of the current node, where the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node to accept or reject being the next-hop node of the current node; determine a decision threshold based on the maximum value of the dynamic transition probability, where the decision threshold is used to determine whether a sampling node accepts or rejects being the next-hop node of the current node; randomly select a probability value within the sampling range determined by the maximum value of the dynamic transition probability; when the difference between the probability value and the maximum value of the dynamic transition probability is greater than or equal to the decision threshold, it indicates that the sampling node accepts being the next-hop node of the current node; when the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, then determine the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node.

[0180] Further, the determination module 12 is further configured to: determine the sampling effective area of each neighbor node in the neighbor node list; when there is a sampling effective area to be cropped that exceeds a preset threshold in the sampling effective area of the previous-hop node of the current node, determine the cropping area region of the sampling effective area to be cropped, where the cropping area region is the area region that exceeds the preset threshold in the sampling effective area to be cropped; and adjust the maximum value of the dynamic transition probability of the current node based on the cropping area region.

[0181] Further, the determination module 12 is further configured to: obtain the number of neighbor nodes included in the neighbor node list; and adjust the maximum value of the dynamic transition probability of the current node according to the cropping area region and the number of neighbor nodes.

[0182] Further, the adjustment module 13 is further configured to: obtain the sampling effective area corresponding to the sampling node; and adjust the ratio between the sampling effective area of each neighbor node of the current node and the total area of the rectangle based on the sampling effective area corresponding to the sampling node.

[0183] Further, the adjustment module 13 is further configured to: adjust the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node, where the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node to accept or reject being the next-hop node of the current node.

[0184] Further, the adjustment module 13 is further configured to: determine the sampling effective area corresponding to the sampling node; obtain the number of neighbor nodes included in the neighbor node list; and adjust the maximum value of the dynamic transition probability of the current node based on the number of neighbor nodes and the sampling effective area.

[0185] Further, the second loop module 15 is further configured to: if the new sampling node accepts to be the next-hop node of the current node, update the current node, repeat the above steps, and traverse the network diagram to obtain the end point of the random walk sequence.

[0186] Thus, in the node search method for the network diagram in the embodiments of the present application, by adjusting the sampling strategy according to the sampling nodes whose sampling is rejected, the sampling probability of the rejected sampling nodes is increased, so as to increase the overall sampling probability of the neighbor nodes of the current node, improve the sampling efficiency of the random walk sequence, and reduce the sampling time.

[0187] It should be understood that the various units or modules described in the device correspond to the respective steps in the method described above. Therefore, the operation instructions and features described above for the method also apply to the device and the modules included therein, and will not be repeated here. The device can be pre-implemented in the browser or other security applications of the electronic device, or can be loaded into the browser or other security applications of the electronic device by means of downloading, etc. The corresponding modules in the device can cooperate with the modules in the electronic device to implement the solutions of the embodiments of the present application.

[0188] For the several modules or units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0189] Next, with reference to Figure 16 , Figure 16 FIG. shows a schematic structural diagram of a computer system of an electronic device or a server suitable for implementing the embodiments of the present application.

[0190] As Figure 16 shown, the computer system includes a central processing unit (CPU) 1501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1502 or the program loaded from the storage section 1508 into the random access memory (RAM) 1503. In the RAM 1503, various programs and data required for the operation instructions of the system are also stored. The CPU 1501, ROM 1502, and RAM 1503 are connected to each other through a bus 1504. The input / output (I / O) interface 1505 is also connected to the bus 1504.

[0191] The following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, etc.; an output section 1507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 1509 including a network interface card such as a LAN card, a modem, etc. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the I / O interface 1505 as needed. A removable medium 1511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1510 as needed so that a computer program read therefrom is installed into the storage section 1508 as needed.

[0192] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart Figure 2 can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1509, and / or installed from the removable medium 1511. When the computer program is executed by a central processing unit (CPU) 1501, the above-described functions defined in the system of the present application are executed.

[0193] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operation instructions of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two connected blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0195] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition module, a determination module, an adjustment module, a first loop module, and a second loop module. Among them, the names of these units or modules do not constitute a limitation on the units or modules themselves in some cases. For example, the first acquisition module can also be described as "acquiring a list of the current node and the neighbor nodes of the current node in a pre-stored network diagram, where the neighbor node list includes at least one neighbor node, and the current node serves as the starting node of a random walk sequence".

[0196] As another aspect, this application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are executed by one or more processors, they are used to perform the node search method in the network diagram described in this application.

[0197] The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for searching nodes in a network diagram, characterized in that, Perform message search based on a message network graph, where the network graph is a graph model applied in the field of message search, and the search field includes news search, image search, and music search. The method includes: Obtain the current node and the list of neighbor nodes of the current node in the pre-stored network graph. The list of neighbor nodes includes at least one neighbor node, and the current node serves as the starting node of the random walk sequence; Determine the sampling node of the current node from the list of neighbor nodes based on the sampling strategy of the current node; If the sampling node refuses to be the next-hop node of the current node, adjust the sampling strategy. The step of adjusting the sampling strategy when the sampling node refuses to be the next-hop node of the current node includes: obtaining the sampling effective area corresponding to the sampling node; based on the sampling effective area corresponding to the sampling node, adjust the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle; Based on the adjusted sampling strategy, determine the new sampling node of the current node from the list of neighbor nodes until the new sampling node accepts being the next-hop node of the current node; Update the current node, repeat the above steps, and traverse the network graph to obtain the random walk sequence; After generating the random walk sequence, generate a message recommendation list for the messages corresponding to the random walk sequence in the order of the random walk sequence to send recommendation information to the user.

2. The method for searching nodes in a network diagram according to claim 1, characterized in that The step of determining the new sampling node of the current node from the list of neighbor nodes based on the adjusted sampling strategy includes: Based on the adjusted sampling strategy, determine the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle; Based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle, determine the new sampling node of the current node from the list of neighbor nodes.

3. The method for searching nodes in a network diagram according to claim 2, characterized in that, The step of determining the new sampling node of the current node from the list of neighbor nodes based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle includes: Obtain the maximum value of the dynamic transition probability of the current node, where the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the list of neighbor nodes of the current node accepting or rejecting being the next-hop node of the current node; Determine a decision threshold based on the maximum value of the dynamic transition probability, where the decision threshold is used to determine whether the sampling node accepts or rejects being the next-hop node of the current node; Randomly select a probability value within the sampling range determined by the maximum value of the dynamic transition probability; When the difference between the probability value and the maximum value of the dynamic transition probability is greater than or equal to the decision threshold, it indicates that the sampling node accepts being the next-hop node of the current node; When the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, the sampling node of the current node is determined from the neighbor node list based on the sampling strategy of the current node.

4. The method for searching nodes in a network diagram according to claim 1, wherein Determining the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node includes: Determining the sampling effective area of each neighbor node in the neighbor node list; When there is a to-be-clipped sampling effective area in the sampling effective area of the previous-hop node of the current node that exceeds a preset threshold, determining a clipped area region of the to-be-clipped sampling effective area, where the clipped area region is the area region in the to-be-clipped sampling effective area that exceeds the preset threshold; Adjusting the maximum value of the dynamic transition probability of the current node based on the clipped area region.

5. The method for searching nodes in a network diagram according to claim 4, wherein Adjusting the maximum value of the dynamic transition probability of the current node based on the clipped area region includes: Obtaining the number of neighbor nodes included in the neighbor node list; Adjusting the maximum value of the dynamic transition probability of the current node according to the clipped area region and the number of neighbor nodes.

6. The method for searching nodes in a network diagram according to claim 1, characterized in that, Adjusting the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle based on the sampling effective area corresponding to the sampling node includes: Adjusting the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node, where the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities for each neighbor node in the neighbor node list of the current node to be accepted or rejected as the next-hop node of the current node.

7. The method for searching nodes in a network diagram according to claim 6, characterized in that, Adjusting the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampling node includes: Determining the sampling effective area corresponding to the sampling node; Obtaining the number of neighbor nodes included in the neighbor node list; Adjusting the maximum value of the dynamic transition probability of the current node based on the number of neighbor nodes and the sampling effective area.

8. The method for searching nodes in a network diagram according to claim 1, wherein If the new sampling node accepts to be the next-hop node of the current node, update the current node, and repeat the above steps to traverse the network graph to obtain the end point of the random walk sequence.

9. A node search device in a network diagram, characterized in that, Performing message search based on a message network graph, where the network graph is a graph model applied in the field of message search, and the search fields include news search, image search, and music search. The device includes: A first acquisition module, configured to acquire the current node and the neighbor node list of the current node in a pre-stored network graph, where the neighbor node list includes at least one neighbor node, and the current node serves as the starting node of a random walk sequence; A determination module, configured to determine the sampling node of the current node from the neighbor node list based on the sampling strategy of the current node; An adjustment module, configured to adjust the sampling strategy if the sampling node rejects to be the next-hop node of the current node; the adjustment module is further configured to: obtain the sampling effective area corresponding to the sampling node; adjust the ratio between the sampling effective areas of each neighbor node of the current node and the total area of the rectangle based on the sampling effective area corresponding to the sampling node; The first loop module is used to determine new sampled nodes of the current node from the neighbor node list based on the adjusted sampling strategy until the new sampled nodes are accepted as the next-hop nodes of the current node; The second loop module is used to update the current node, repeat the above steps, traverse the network graph to obtain the random walk sequence; after generating the random walk sequence, generate a message recommendation list for the messages corresponding to the random walk sequence in the order of the random walk sequence to send recommendation information to the user.

10. The device according to claim 9, characterized in that, The first loop module is further used to determine the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle based on the adjusted sampling strategy; determine new sampled nodes of the current node from the neighbor node list based on the ratio of the sampling effective area corresponding to each neighbor node of the adjusted current node to the total area of the rectangle.

11. The device according to claim 10, wherein, The first loop module is further used to: obtain the maximum value of the dynamic transition probability of the current node, and the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node being accepted or rejected as the next-hop node of the current node; Determine a decision threshold based on the maximum value of the dynamic transition probability, and the decision threshold is used to determine whether the sampled node is accepted or rejected as the next-hop node of the current node; randomly select a probability value within the sampling range determined by the maximum value of the dynamic transition probability; When the difference between the probability value and the maximum value of the dynamic transition probability is greater than or equal to the decision threshold, it means that the sampled node is accepted as the next-hop node of the current node; When the difference between the probability value and the maximum value of the dynamic transition probability is less than the decision threshold, then determine the sampled node of the current node from the neighbor node list based on the sampling strategy of the current node.

12. The device according to claim 9, characterized in that, The determination module is further used to: determine the sampling effective area of each neighbor node in the neighbor node list; when there is a to-be-clipped sampling effective area of the previous-hop node of the current node that exceeds the preset threshold, determine the clipped area region of the to-be-clipped sampling effective area, and the clipped area region is the area region that exceeds the preset threshold in the to-be-clipped sampling effective area; adjust the maximum value of the dynamic transition probability of the current node based on the clipped area region.

13. The device according to claim 12, characterized in that, The determination module is further used to: obtain the number of neighbor nodes included in the neighbor node list; adjust the maximum value of the dynamic transition probability of the current node according to the clipped area region and the number of neighbor nodes.

14. The device according to claim 9, wherein The adjustment module is further used to: adjust the maximum value of the dynamic transition probability of the current node based on the sampling effective area corresponding to the sampled node, and the maximum value of the dynamic transition probability is used to determine the ratio of the probabilities of each neighbor node in the neighbor node list of the current node being accepted or rejected as the next-hop node of the current node.

15. The device according to claim 14, characterized in that, The adjustment module is further used to: determine the sampling effective area corresponding to the sampled node; obtain the number of neighbor nodes included in the neighbor node list; adjust the maximum value of the dynamic transition probability of the current node based on the number of neighbor nodes and the sampling effective area.

16. The device according to claim 9, wherein, The second loop module is further configured to: if a new sampling node is accepted as the next-hop node of the current node, update the current node, repeat the above steps, and traverse the network diagram to obtain the end point of the random walk sequence.

17. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the node search method in the network diagram according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the node search method in the network diagram according to any one of claims 1-8.

Citation Information

Patent Citations

  • Degree offset sampling method and system for social network

    CN110717107A