Dual-cost driven network communication cascade intelligent perception and traceability method and device

Through the dual-cost-driven intelligent perception and traceability method of network propagation cascade, the monitoring set is optimized by efficient hybrid algorithms and network structure strategies, the problem of high cost in network propagation source inference is solved, and low-cost traceability of communication source and network public opinion perception is realized.

CN119182582BActive Publication Date: 2025-08-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411249461.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-08-22
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The prior art has high cost and unstable problems in network propagation source inference, especially in power law distribution networks, and traditional methods are difficult to effectively minimize the cost of positioning propagation source.

Method used

Using a dual-cost-driven network propagation cascade intelligent perception and traceability method, the optimal target monitor set Vo is determined through the efficient hybrid algorithm EHA, and combined with the network structure-based policy NSS and greedy detection strategy GDS, the monitor set is dynamically optimized to minimize the detection cost of the propagation source.

Benefits of technology

It realizes stable and efficient dynamic tracing of the source of dissemination at lower costs, and can quickly perceive and efficiently trace the network public opinion information. It is suitable for fighting false information, monitoring network structure infrastructure, analyzing the robustness of complex systems and controlling information or disease spread.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119182582B_ABST
    Figure CN119182582B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-cost driven network communication cascade intelligent perception and traceability method and device, comprising: step 1: determining the optimal target monitor set V according to the topological structure information of the basic network; o Step 2: According to the historical propagation set ζ information, intelligently optimize the monitor V o Based on the configuration of the outbreak, determine the efficient transmission outbreak sensing set #imgabs0#. Step 3: Based on the outbreak sensing set #imgabs1# and combined with real-time information obtained from the specific outbreak #imgabs2#, dynamically determine the initial transmission source that caused the outbreak #imgabs3#. The method provided by this invention can achieve stable and efficient dynamic transmission source tracing at a low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information processing technology, and specifically relates to a dual-cost driven network communication cascade intelligent perception and tracing method and device. Background Art

[0002] Contagious processes occurring on networks can be used to model a variety of real-world scenarios, such as information dissemination, disease, and social behavior. By leveraging fundamental epidemic models, researchers can gain insights into real-world transmission dynamics and develop strategies for tackling complex tasks. In this context, one of the most critical aspects is inferring the source of transmission, whose solutions have potential applications in areas such as combating disinformation and suppressing epidemics.

[0003] The problem of inferring the source of an outbreak aims to optimize a collection of monitors in order to minimize the cost of locating the source that triggered the outbreak. Traditional approaches assume that the states of some or all nodes are observable and propose developing efficient estimators to assess the likelihood of each node being the source. In this context, two main strategies are commonly employed to find the true source.

[0004] The first strategy focuses on the node with the highest probability and gradually expands the search area by considering its first-level neighbors, second-level neighbors, and so on, until the true source is found. In this approach, the error distance is often used to measure the effectiveness of different methods; if the inferred source is close to the true source in terms of hop distance, the estimator is considered to be better. While this optimization objective is effective for general random networks, it can be unstable for networks with power-law degree distributions.

[0005] In contrast, the second type of strategy first relies on the likelihood provided by the estimator to rank the nodes, and then checks each node in turn until the true source is found. In this case, all nodes with a likelihood greater than zero are considered potential sources, so if the estimator results in a smaller number of nodes with a likelihood greater than zero, the estimator is considered better. In other words, the optimization goal is to minimize the number of nodes with a likelihood greater than zero, which is usually more stable than the error distance because it can be achieved by mining the network structure without knowing the underlying diffusion pattern or model. This type of strategy is more applicable to real-world scenarios than the previous one, but most existing work on this type of strategy infers the source of each outbreak based on all the information collected by the monitor set, which sometimes results in very high costs. Summary of the Invention

[0006] In order to overcome the shortcomings of the prior art, the present invention provides a dual-cost driven network propagation cascade intelligent perception and traceability method and device, including: Step 1: Determine the optimal target monitor set V according to the topological structure information of the basic network oStep 2: According to the historical propagation set ζ information, intelligently optimize the monitor V o Configuration, determine the efficient perception set of the propagation burst Step 3: Based on the outbreak perception set Combined with the specific outbreak The real-time information obtained from the The initial source of the outbreak. The method provided by the present invention can achieve stable and efficient dynamic tracing of the source of the outbreak at a relatively low cost.

[0007] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0008] Step 1: Determine the optimal target monitor set V based on the topological structure information of the basic network o ;

[0009] Step 2: Based on the information of the historical propagation set ζ, intelligently optimize the optimal target monitor set V o Configuration, determine the efficient perception set of the propagation burst

[0010] Step 3: Efficiently perceive the set based on the burst Combined with the specific outbreak The real-time information obtained from the The initial source of transmission of the outbreak.

[0011] Furthermore, the step 1 is specifically as follows:

[0012] For a specific network G(V,E), V and E represent the vertex set and edge set in the network respectively; the size of the given monitor set n o and hyperparameters x, using the efficient hybrid algorithm EHA for the optimal target monitor set V in the first stage TSS-I o The solution of

[0013] The efficient hybrid algorithm EHA includes: o BPD algorithm of initial solution and monitor set V o EPP algorithm for effective improvement of initial solution.

[0014] Furthermore, the acquisition of the monitor set V o The BPD algorithm for the initial solution is as follows:

[0015] Step 1-1: Create an empty monitor set V o , and assign an initial position to the node sequence S; initialize a counter to track the order in which the nodes are added; define G′(V o ) is the remaining network after removing the monitor set;

[0016] Step 1-2: When G′(V o ), that is, when there is a cycle in the remaining network G′=(V′,E′): first calculate the removal probability of each node Select the node u with the highest probability of removal from the remaining network, that is: Add node u to the monitor set V o And record it in the sequence S(i), update the counter, and add 1 to i; V′ and E′ respectively represent the vertex set and edge set in the remaining network after removing the monitor set;

[0017] Step 1-3: When i≤n: select the remaining network G′(V o )’s largest connected component LCC has the largest disassembly effect node u, add node u to the monitor set V o And record it in the sequence S(i), update the counter, i plus 1;

[0018] Step 1-4: Repeat steps 1-2 and 1-3 until the remaining network G′ has a cycle and the loop ends, obtaining the monitor set V o Initial solution.

[0019] Furthermore, the monitor set V o The EPP algorithm for effectively improving the initial solution is as follows:

[0020] Step 1-5: According to the node sequence S, randomly select a subinterval S(l1:l2) of the node sequence; l1 and l2 are two randomly generated parameters for selecting the node sequence subinterval, 1≤l1 <l2≤n;

[0021] Step 1-6: Optimize the node subinterval S(l1:l2) according to the RR relationship strategy to reduce the size of the largest connected component LCC and optimize the order of node deletion;

[0022] Step 1-7: Get the improved monitor set V o .

[0023] Furthermore, the step 2 is specifically as follows:

[0024] For a specific network G(V,E), the monitor set V o And node sequence S, given the outbreak set ζ, introduce the network structure-based strategy NSS and the greedy-based strategy GDS to determine the monitor set V o The order of checking nodes in the middle, intelligently constructing the burst-aware set described in the second stage TSS-II

[0025] Furthermore, the network structure-based strategy NSS is as follows:

[0026] Step 2-1: Initialize index i to 1 and initialize the outbreak perception set is an empty set;

[0027] Step 2-2: Until an outbreak is detected Or before i exceeds the size of the monitor set, execute the following loop: select node u from the node sequence S, where u = S(i), and then add node u to the outbreak-aware set In the above example, update index i plus 1;

[0028] Step 2-3: End the loop and get the burst perception set

[0029] Furthermore, the greedy-based strategy GDS is as follows:

[0030] Step 2-4: Initialize the outbreak awareness set is an empty set;

[0031] Step 2-5: Before the desired detection rate is reached, perform the following operations: o Select a node u from the set so that u is added to the outbreak-aware set The objective function can be minimized And add node u to the outbreak-aware set middle;

[0032] Step 2-6: End the loop and add to the burst perception set The order of the nodes in , update the node sequence S;

[0033] Step 2-7: Get the burst perception set

[0034] Furthermore, the step 3 is specifically as follows:

[0035] For a specific network G(V,E), the perception set For the third stage TSS-III, the greedy strategy PGS based on percolation and CIR algorithm are used to intelligently optimize the candidate set V by dynamically checking the nodes. c and minimize the set of dynamic inference monitors

[0036] Furthermore, the CIR algorithm includes:

[0037] Step 3-1: Use burst-aware aggregation Initializes the dynamic inference monitor collection

[0038] Step 3-2: When there is a current set of monitors that has not been checked Or the current candidate set The requirements have not been met yet, and the following loop is executed: i) Based on the dynamic inference monitor set Get Subnetwork Update the current monitor set Use the greedy strategy PGS based on percolation to extract the current monitor set Select the next node u to be checked d iv) Select the node u d Join the dynamic inference monitor collection

[0039] Step 3-3: End the loop and update the candidate set

[0040] Step 3-4: Get the candidate set V c .

[0041] Furthermore, the greedy strategy PGS based on percolation is as follows:

[0042] Steps 3-5: When potential monitors to be checked are collected When the number of nodes in is greater than 1, execute the following loop: i) Select a node u such that: From a given subnetwork Remove node u; α′(v) represents the subnetwork G t Remove The connected components of the remaining network after all nodes are quantized;

[0043] Step 3-6: When the loop ends, the last remaining node v is the next node u to be checked d ;

[0044] Step 3-7: Get the next node u to be checked d .

[0045] A dual-cost driven network communication cascade intelligent perception and traceability device, comprising:

[0046] The acquisition unit obtains the optimal target monitor set V according to the topological structure information of the basic network o ;

[0047] The perception unit intelligently optimizes the monitor V according to the information of the historical propagation set ζ o Configuration, determine the efficient perception set of the propagation burst

[0048] The tracing unit, based on the outbreak perception set Combined with the specific outbreak The real-time information obtained from the The initial source of transmission of the outbreak.

[0049] The beneficial effects of the present invention are as follows:

[0050] By optimizing the detection order of network monitor nodes, the method of the present invention can achieve intelligent and rapid perception and efficient tracing of network public opinion information at a relatively low cost, and can thus be widely used in various real-world scenarios, such as combating false information, monitoring the status of network structure infrastructure, analyzing the robustness and resilience of complex systems, and controlling the spread of information or diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flow chart of the device of the present invention;

[0052] Figure 2 is a schematic diagram of the network decomposition strategy used in the present invention;

[0053] Figure 3 is the proportion of dynamic monitor set nodes of TSS and the four comparison methods (a, e, i) HD, (b, f, j) CI, (c, g, k) MSRG and (d, h, l) FINDER) θ φ Probability density distribution diagram, where φ = 1. S in each sub-diagram represents the corresponding method. (a)-(d), (e)-(h), and (i)-(l) represent the experimental results on the Airline network, Gowalla network, and web-Google network, respectively.

[0054] Figure 4 is the success rate of TSS with HD, CI, MSRG and FINDER φ and the average proportion of nodes in the dynamic monitor collection Comparison results, where S is used in each sub-figure to represent the corresponding method, where (a) (e), (b) (f), (c) (g) and (d) (h) represent the experimental results on four large-scale networks: Power network, Gowalla network, Twitter network and web-Google network respectively;

[0055] Figure 5 Is the average proportion of dynamic monitor set nodes The graph shows the variation of the monitor set ratio φ, where the solid curve and the dashed curve represent the CIR strategy with and without PGS, respectively; (a), (b), (c), and (d) represent the experimental results on four large-scale networks: Power Network, Gowalla Network, Twitter Network, and Web-Google Network, respectively;

[0056] Figure 6 is the average proportion of the dynamic monitor set by the TSS device and the other four methods Comparison diagram, (a)(e), (b)(f), (c)(g) and (d)(h) are the experimental results of CI, MSGR, FINDER and TSS respectively. The x-axis and y-axis represent the GDS strategy and NSS strategy used in the second stage NSS-II respectively; (a)-(d) and (e)-(h) represent the experimental results in Airline network and Power network respectively;

[0057] Figure 7 is the cost of TSS, CI and FINDER devices Comparison of cost performance on four large network datasets as a function of the monitor set node ratio φ, i.e., C1, where (a), (b), (c), and (d) represent the experimental results on four large-scale networks: Power Network, Gowalla Network, Twitter Network, and Web-Google Network, respectively. DETAILED DESCRIPTION

[0058] The present invention will be further described below with reference to the accompanying drawings and examples.

[0059] This invention aims to address the relatively high cost of detecting and tracing each outbreak in the existing network tracing process. To overcome the shortcomings of existing communication source inference technology and minimize the cost of detecting each network public opinion, this invention provides a dual-cost driven network communication cascade intelligent perception and tracing method. The specific implementation scheme is as follows:

[0060] For the TSS-I stage, the present invention proposes an efficient hybrid algorithm (EHA) by combining belief propagation guided sparsification (BPD) and explosive penetration based pruning (EPP) methods to minimize the static monitor set (one-time cost), thereby ensuring a limited candidate set even in the worst case. For the TSS-II stage, the present invention studies the network structure-based strategy (NSS) and greedy detection strategy (GDS) to detect the epidemic in a timely manner, and based on this, execute the subsequent TSS-III as early as possible. Once the epidemic is detected, the TSS-III stage is entered, in which a check, inference and exclusion (CIR) algorithm based on the penetration-based greedy search (PGS) strategy is proposed to dynamically minimize the dynamic cost of inferring the source.

[0061] The present invention can achieve rapid perception and efficient tracing of online public opinion information at a low cost by optimizing the detection order of network monitor nodes. It can be widely used in various real-world scenarios, such as combating false information, monitoring the status of network structure infrastructure, analyzing the robustness and resilience of complex systems, and controlling the spread of information or diseases.

[0062] like Figure 1 As shown, the present invention provides a dual-cost driven efficient dynamic three-stage tracing device for the source of transmission. It mainly uses the efficient hybrid device (EHA) combining the belief propagation guided sparsification solver BPD proposed in the TSS-I stage with the explosive percolation based pruning strategy EPP to minimize the static monitor set (one-time cost), the network structure based device (NSS) and greedy detection device (GDS) proposed in the TSS-II stage to detect the epidemic in a timely manner, and the inspection, inference and exclusion device (CIR) based on the percolation based greedy search device (PGS) proposed in the TSS-III stage to intelligently minimize the dynamic cost of the inferred source, thereby solving the dual-cost source inference problem. Specifically, it includes the following steps:

[0063] Step 1: Assume the diffusion being studied It is transmitted by an unknown source u s ∈V, propagates on a contact network G(V,E) containing n=|V| nodes and m=|E| edges, where V and E are the node set and edge set respectively. In order to facilitate the description of the proposed model and method, it is assumed that Following the discrete susceptible-infected-recovered (SIR) transmission dynamics model, in this case, each node v in the network G can be in one of three states, namely susceptible (S), infected (I) or recovered (R), that is, σ v ∈{S,I,R}. Let t denote the timestamp of the dynamic evolution, starting from t=0. At each time unit, i.e. t→t+1, if the edge e uv In the SI state, an infected node v infects its susceptible neighbor u with an infection probability β. Meanwhile, the infected node v recovers with a recovery probability γ.

[0064] Step 2: For the target epidemic The main goal of this invention is to find an estimator Ψ(·) that returns a candidate set V c =Ψ(V o ), where the true source u can be determined with high probability s Belong to this collection. This process is based on the monitor (monitoring) collection Based on the information collected by the nodes in the network, the estimator first evaluates the probability that each node may be the source, denoted as This probability is conditioned on Vo Information Collected Then based on Assemble candidate set V c ,in V o In this case, since the states of all nodes are available, the infection timestamp of a node can be indirectly and partially inferred from the states of its surrounding neighbors. That is, if only one of a node's neighbors is in an infected or recovered state, it can be concluded with high probability that the node was the last to be infected. The states of the remaining nodes can then be further obtained through recursive inference.

[0065] Step 3: Assume that under a certain strategy, the candidate set V c Add more nodes in V c Probability Will follow n c The increase gradually increases, where n c =|V c | is the candidate set V c The size of n o =|V o |For monitor set V o For some fixed n c and n o , May vary with different V o configuration varies, assuming V c By having the maximum probability n c nodes. That is, by optimizing V o To maximize Therefore, the diffusion source inference (DSI) optimization problem can be written as:

[0066]

[0067] That is, the DSI problem is equivalent to the monitor set V o Optimization, assuming that V c A deeper search is performed on the source u and a certain cost is applied to each node to accurately locate the source u s , when n→∞, n c / n→0Search the entire candidate set V c The cost will approach zero, that is, V c is finite. Therefore, in n c / n→0, this paper mainly studies the most interesting scenario, namely That is, by searching for V c The location of the source of transmission can be determined.

[0068] Step 4: Next, consider determining the candidate set V c Make Specifically, for a given network G(V,E) and an outbreak event make Let Ω be the set of all connected components in the residual network G′(V′,E′) that are present after removing all V o The nodes in V′ are obtained by o And E′=(V′×V′)∩E. Now, for node u∈V o , defining Γc(u)=Γ(u)∩V′, the component coverage α(u) can be expressed as:

[0069]

[0070] where ci(v)∈Ω is the component in G′ that contains the node v. Let V o ′ indicates that the o The key monitor set extracted from the , that is, the target monitor set that detects the infection first:

[0071]

[0072] Given G(V,E), V o 、 And with V o Related infection status and timestamp information, define the candidate set V c :

[0073]

[0074] Then the candidate set constructed according to this method must contain the propagation source u s .

[0075] like Figure 2 As shown, if V o =u,v, then the original network is decomposed into a subgraph G′=c consisting of five connected pieces i , i=[1,5], the connected cover and candidate set composition examples are:

[0076] Diffusion events occurring in component c2 can only spread to c1 through u2 and u1. According to the previous conditions, u2 must be infected before u1. Therefore, by checking the infection timestamps t(u) of u1 and u2, if it is found that t(u2) < t(u1), the source of propagation must belong to c2. If u1, u2, u3, and u4 are assumed to be monitors, then the size of the candidate set to which the source must belong should be limited to max|c1|, |c2|, |c3|, |c4|+|c5|; if only u1, u2, and u4 are monitors, then the upper limit should be max|c1|+c3+1, |c2|, |c3|+|c4|+|c5|+1. Therefore, the process of determining the minimum number of nodes required in the monitor set incurs a one-time cost C1. If a diffusion event occurs in c2, then only u1 and u2 need to be checked to determine; however, if the diffusion occurs in c1, then u1, u2, u3, and u4 need to be checked to finally confirm. Therefore, determining which component the source belongs to, that is, the minimum number of nodes to be checked in the monitor set, incurs a dynamic cost C2.

[0077] Step 5: Based on Figure 2 the inference results, Equation (1) can be further rewritten as:

[0078]

[0079] where the weight factor 0 ≤ ∈ ≤ 1 is used to control the contributions of these two parts of the costs to the total cost C t = ∈C1+(1 - ∈)C2.

[0080] Finally, the bi-cost propagation source inference (BDSI) problem is defined as follows: Given a graph G(V, E), a diffusion set ζ, a tolerance δ, and a weight factor ∈, the bi-cost propagation source inference problem requires finding a monitor set to minimize E(C t ), such that for most there is and n c / n ≤ δ, where n c represents the size of the candidate set V c .

[0081] Step 6: For the optimization of the one-time cost C1 proposed in Step 5, in the TSS-I stage, the present invention proposes an efficient hybrid apparatus (EHA) by combining belief propagation-guided sparsification (BPD) and explosion percolation-based pruning (EPP) to minimize the static monitor set, that is, the one-time cost, and ensure a finite candidate set even in the worst case.

[0082] Specifically, obtain the BPD algorithm for the initial solution of the monitor set V o :

[0083] First, when G′(V o ), that is, when there is a cycle in the remaining network G′=(V′,E′): calculate the probability of removing the node Select the node u with the highest probability of removal, that is: Add node u to the monitor set V o And record it in the sequence S(i), let i=i+1. When i≤n: select the remaining network G′(V o )’s largest connected component LCC has the largest disassembly effect node u, add node u to the monitor set V o And record it into the sequence S(i), let i=i+1. After the cycle ends, the monitor set V is finally obtained o Initial solution.

[0084] Specifically, for the monitor set V o EPP algorithm for effective improvement of the initial solution: First, randomly select a node sequence subinterval S(l1:l2). Optimize the node subinterval S(l1:l2) according to the RR relationship strategy to reduce the size of the largest connected component LCC. Repeat multiple cycles to obtain the improved monitor set V o .

[0085] Specifically, the EHA device first obtains the monitor set V through the BPD algorithm o Initial solution, and then use the EPP algorithm to analyze the monitor set V o The initial solution is improved and the one-time cost C1 is minimized by using an efficient mixing apparatus (EHA).

[0086] Step 7: Optimize the dynamic cost C2 proposed in step 5. Specifically, first use the network structure-based device (NSS) and greedy detection device (GDS) proposed in the TSS-II stage to timely detect public opinion, and then execute the subsequent TSS-III as soon as possible based on this.

[0087] Specifically, the network structure-based device NSS: initialize index i and outbreak perception set Execute the following loop: select node u from the node sequence S, where u = S(i), and then add node u to the burst-aware set In the example, let i=i+1. After the loop ends, the burst perception set can be obtained.

[0088] Similarly, specifically, based on the greedy strategy GDS: Initialize the outbreak-aware set is an empty set. Before the expected detection rate is reached, the following loop is executed: oSelect a node u from the set so that u is added to the outbreak-aware set The objective function can be minimized Add node u to the burst-aware set End the loop and add it to the burst perception set The order of the nodes in the update node sequence S, and finally get the burst perception set

[0089] Step 8: Once an epidemic is detected, the TSS-III stage is entered, in which a detection, inference and exclusion device (CIR) based on a percolation-based greedy search (PGS) device is proposed to intelligently minimize the dynamic cost C2 of the inferred source.

[0090] Specifically, the greedy device PGS based on percolation: When the number of nodes in is greater than 1, execute the following loop: i) Select a node u such that: From a given subnetwork Remove node u. When the loop ends, the last remaining node v is the next node u to be checked. d .

[0091] Specifically, the CIR device: initializes a dynamic inference monitor set When there is a current monitor set that has not been checked Or the current candidate set The requirements have not been met yet, and the following loop is executed: i) Based on the dynamic inference monitor set Get Subnetwork Update the current monitor set Use the greedy strategy PGS based on percolation to extract the current monitor set Select the next node u to be checked d iv) Select the node u d Join the dynamic inference monitor collection End the loop and update the candidate set Finally, the CIR device obtains a certain number of propagation sources u under the condition of minimizing the dynamic cost C2 of the inferred source. s The candidate set V c .

[0092] Table 1 Experimental network dataset

[0093] Dataset describe Number of nodes Number of sides web-Google Google Web Network 875713 4322051 Airline Air transport network 3146 18164 Power Power system network 4941 6594 Gowalla Location-based social networking 196591 950327 Crime Criminal Network 829 1473 Twitter Twitter social network 532325 694606

[0094] To verify the effectiveness of the present invention, experiments were conducted on a real social network. The experimental network data are shown in Table 1. The experimental parameters are set as follows: (1) The dynamic propagation dataset is generated by the SIR propagation model on the real network. The node infection probability and recovery probability are random values ​​in the interval (0, 1]; (2) An outbreak is considered valid only when the proportion of infected nodes exceeds 5% of the total population. Each evaluation outbreak set ζ contains 103 valid outbreaks. In addition, τ p =τ o = 50 experiments. (3) The BPD of EHA is set to x = 12.0, and the top 1% nodes with the highest degrees are directly removed before calculation, that is, they are directly used as monitors. (4) The EPP of EHA terminates when the solution S is stable (determined by running the EPP device on S 10 times in a row and not achieving any improvement in R(l1,l2)). (5) The GDS of TSS-II is set to d r =1.0, PGS is set to nt o = 10, CIR every time check Node, that is

[0095] The experimental comparison methods include: Random method (Random) High Degree (HD) method, which selects nodes with larger degrees in the network as observation points; Collective Influence (CI), Min-sum and Reverse-greedy (MSRG) strategy, Finding key nodes in the network through deep reinforcement learning (FINDER), Generalized Network Demolition (GND) method, which is developed based on graph partitioning, Cycle Ratio (CR) method, which is similar to HD but ranks each node according to its participation in the shortest cycle in the network, Greedy Strategy (GS).

[0096] Figure 3 The probability density of θ1 for TSS compared with HD, CI, MSRG and FINDER is given, ignoring the constraints Instead, it is constructed by any node in V That is θ φ Probability density where φ = 1.

[0097] Depend on Figure 3 It can be seen that the TSS method obtains a smaller set of dynamic inference monitors in more than 98% of cases compared with HD, 93% compared with CI, 99% compared with MSRG, and 98% compared with FINDER, indicating the effectiveness of the proposed method.

[0098] Figure 4The success rates of TSS, HD, CI, MSRG and FINDER are given. φ and the average proportion of nodes in the dynamic monitor collection For each method, the present invention calculates ψ in the range of φ = [0.01, 0.40] with an equal interval of 0.01. φ and These results are based on a set of 103 outbreaks in the SIR dynamics. Similarly, the TSS-II process here uses the NSS strategy, and CIR is performed without considering PGS.

[0099] Depend on Figure 4 You know, small (TSS) is accompanied by a large ψ φ (TSS), indicating that the proposed TSS converges quickly with the increase of φ (see Figure 4 a-4d with ψ φ (TSS) increases the sparsity of data points. For example, on the web-Google network, when φ = 0.06, the TSS method achieves ψ φ =93.5% and The FINDER method has ψ φ =7.3% and It should be noted that the larger the monitor set, the more likely it is to satisfy the constraint n c / n≤δ, the greater the possibility.

[0100] Figure 5 Gives the average proportion of nodes in the dynamic monitor set The graph changes with the ratio φ of the monitor set. The solid and dashed curves represent the CIR strategies with and without PGS, respectively. The dataset follows the SIR dynamics, and NSS is used in TSS-II.

[0101] Depend on Figure 5 It can be seen that As a function of φ, the CIR with and without PGS is compared. In general, Figure 5 As shown, the PGS strategy can effectively reduce the set of dynamic inference monitors and improve the performance of all studied baseline methods, in this way ensuring n c / n≤δ monitor set V o becomes crucial. Figure 5Another conclusion that can be observed is that TSS with PGS performs worse than TSS without PGS. The reason may be that both EPP and PGS of EHA are developed during the percolation process, so the expansion operation of PGS cannot provide a better solution than EHA. However, a smaller TSS can be obtained by simply reducing the number of nodes checked each time or optimizing R(l1,l2). These results also further confirm the effectiveness of the algorithm used in TSS-I.

[0102] Figure 6 The average value of the proportion of dynamic monitor sets occupied by TSS device and other four methods is given. The x-axis and y-axis represent the GDS strategy and NSS strategy used in the second stage NSS-II, respectively. In each method of each network, the ratio of the monitor set φ = [0.01, 0.40] is taken. value, with an interval of 0.01, τ p =τ o ∈[2,11], with an interval of 1. Each graph contains 400 types of φ and τ o configuration, and the results are also based on 103 burst sets ζ following the SIR dynamics.

[0103] Depend on Figure 6 It can be seen that GDS performs better than NSS in most cases, which shows that the early detection through GDS strategy can effectively reduce the dynamic inference monitor set. More specifically, on average over all baseline methods, in more than 85% of cases, GDS can obtain a smaller On the power grid network, it is 94%. In addition, considering these ratios on the proposed method, we obtained results of about 80% and 78% respectively, which is slightly lower than the average. Based on this and Figure 5 Based on the results reported in

[15] , we can conclude that, on the one hand, early detection of an outbreak can effectively facilitate the subsequent inference process in TSS-III in most cases; on the other hand, such early detection can also worsen the outcomes. Therefore, more effective strategies may exist to better balance early detection and effective inference in order to further minimize the set of dynamic inference monitors. It is important to note that GDS can only be used in conjunction with PGS.

[0104] Figure 7 The costs of TSS, CI and FINDER devices are given Comparison of cost performance on four large network datasets as a function of the monitor set node ratio φ, i.e., C1. Similarly, the datasets follow the SIR dynamics, and NSS is used in TSS-II.

[0105] Depend on Figure 7 It can be seen that the cost C on the four large networks t As a function of φ (i.e. C1), the present invention specifically considers the comparison of TSS-IIS without PGS with NSS and CIR to fairly measure the effectiveness of the solution S obtained in TSS-1. Obviously, for each method, the cost C t will reach a minimum value at a certain φ, denoted as And for all four networks, TSS can obtain the minimum This shows the effectiveness of the method proposed in this invention.

[0106] Based on network communication data drive and network structure, the present invention has developed a dual-cost driven efficient dynamic three-stage tracing device for the source of communication. By intelligently optimizing the inspection order of observation points in the network, the present invention can achieve rapid perception and efficient tracing of network public opinion information at a lower cost. It can be widely used in various real-life scenarios, such as combating false information, monitoring the status of network structure infrastructure, analyzing the robustness and resilience of complex systems, and controlling the spread of information or diseases.

[0107] In summary, the device of the present invention realizes efficient dual-cost propagation source inference. The dynamic three-stage method proposed therein can effectively minimize the one-time cost required for deploying monitoring in the process of information tracing and the dynamic cost of tracing information from monitoring to determine the source of public opinion, thereby realizing low-cost and efficient information tracing, showing effectiveness, efficiency and robustness, and is suitable for the public opinion control problem of social networks.

[0108] The dual-cost driven network communication cascade intelligent perception and traceability method provided by the present invention includes the following contents:

[0109] Step 1: Determine the optimal target monitor set V based on the topological structure information of the basic network o ;

[0110] Step 2: Intelligently optimize the monitor V based on the information of the historical propagation set ζ o Configuration, determine the efficient perception set of the propagation burst

[0111] Step 3: Based on the outbreak perception set Combined with the specific outbreak The real-time information obtained from the The initial source of transmission of the outbreak.

[0112] The method provided by the present invention can achieve efficient inference of dual-cost propagation sources at a relatively low cost.

[0113] Specifically, it is assumed that the diffusion under study It is transmitted by an unknown source u s ∈V, and spreads on a contact network G(V,E) containing n = |V| nodes and m = |E| edges, where V and E are the node set and edge set respectively. Following the discrete susceptible-infected-recovered (SIR) transmission dynamics model, in this case, each node v in the network G can be in one of three states, namely susceptible (S), infected (I) or recovered (R), that is, σ v ∈{S,I,R}.

[0114] For target public opinion The main goal of this invention is to find an estimator Ψ(·) that returns a candidate set V c =Ψ(V o ), where the true source u can be determined with high probability s Belong to this collection. This process is based on the monitor (monitoring) collection Based on the information collected by the nodes in the network, the estimator first evaluates the probability that each node may be the source, denoted as This probability is conditioned on V o Information Collected Then based on Assemble candidate set V c ,in V o Information Collected.

[0115] Assume that under a certain strategy, c Add more nodes in V c Probability Will follow n c As the value of n increases, for some fixed n c and n o , Will vary with different V o configuration varies, assuming V c By having the maximum probability n c nodes. That is, by optimizing V o To maximize

[0116] Therefore, the DSI problem is equivalent to the monitor set V o Optimization, assuming that V c A deeper search is performed on the source u and a certain cost is applied to each node to accurately locate the source u s , when n→∞, n c / n→0Search the entire candidate set V cThe cost will approach zero, that is, V c Finally, by introducing the candidate set model, the candidate set constructed by the present invention must contain the propagation source u s .

[0117] Specifically, for the optimization of the one-time cost C1, in the TSS-I stage, the present invention proposes an efficient hybrid arrangement (EHA) by combining belief propagation guided sparsification (BPD) and explosive percolation based pruning (EPP) to minimize the static monitor set, i.e., the one-time cost, ensuring a limited candidate set even in the worst case.

[0118] Specifically, obtain the monitor set V o BPD algorithm for initial solution:

[0119] First, when G′(V o ), that is, when there is a cycle in the remaining network G′=(V′,E′): calculate the probability of removing the node Select the node u with the highest probability of removal, that is: Add node u to the monitor set V o And record it in the sequence S(i), let i=i+1. When i≤n: select the remaining network G′(V o )’s largest connected component LCC has the largest disassembly effect node u, add node u to the monitor set V o And record it into the sequence S(i), let i=i+1. After the cycle ends, the monitor set V is finally obtained o Initial solution.

[0120] Specifically, for the monitor set V o EPP algorithm for effective improvement of the initial solution: First, randomly select a node sequence subinterval S(l1:l2). Optimize the node subinterval S(l1:l2) according to the RR relationship strategy to reduce the size of the largest connected component LCC. Repeat multiple cycles to obtain the improved monitor set V o .

[0121] Specifically, the EHA device first obtains the monitor set V through the BPD algorithm o Initial solution, and then use the EPP algorithm to analyze the monitor set V o The initial solution is improved and the one-time cost C1 is minimized by using an efficient mixing apparatus (EHA).

[0122] Furthermore, for the optimization of the dynamic cost C2, specifically, the network structure-based device (NSS) and greedy detection device (GDS) proposed in the TSS-II stage are first used to timely detect public opinion, and based on this, the subsequent TSS-III is executed as soon as possible.

[0123] Specifically, the network structure-based device NSS: initialize index i and outbreak perception set Execute the following loop: select node u from the node sequence S, where u = S(i), and then add node u to the burst-aware set In the example, let i=i+1. After the loop ends, the burst perception set can be obtained.

[0124] Specifically, based on the greedy strategy GDS: Initialize the outbreak perception set is an empty set. Before the expected detection rate is reached, the following loop is executed: o Select a node u from the set so that u is added to the outbreak-aware set The objective function can be minimized Add node u to the burst-aware set End the loop and add it to the burst perception set The order of the nodes in the update node sequence S, and finally get the burst perception set

[0125] Furthermore, once an epidemic is detected, the TSS-III stage is entered, in which a detection, inference and exclusion device (CIR) based on a percolation-based greedy search (PGS) device is proposed to minimize the dynamic cost C2 of the inferred source.

[0126] Specifically, the greedy device PGS based on percolation: When the number of nodes in is greater than 1, execute the following loop: i) Select a node u such that: From a given subnetwork Remove node u. When the loop ends, the last remaining node v is the next node u to be checked. d .

[0127] Specifically, the CIR device: initializes a dynamic inference monitor set When there is a current monitor set that has not been checked Or the current candidate set The requirements have not been met yet, and the following loop is executed: i) Based on the dynamic inference monitor set Get Subnetwork Update the current monitor set Use the greedy strategy PGS based on percolation to extract the current monitor set Select the next node u to be checked d iv) Select the node u d Join the dynamic inference monitor collection End the loop and update the candidate set Finally, the CIR device obtains a certain number of propagation sources u under the condition of minimizing the dynamic cost C2 of the inferred source. s The candidate set V c , ultimately achieving efficient inference of dual-cost propagation sources.

[0128] The present invention also provides a dual-cost driven network propagation cascade intelligent perception and traceability device, comprising:

[0129] The acquisition unit obtains the optimal target monitor set V according to the topological structure information of the basic network o ;

[0130] The perception unit intelligently optimizes the monitor V according to the information of the historical propagation set ζ o Configuration, determine the efficient perception set of the propagation burst

[0131] The tracing unit, based on the outbreak perception set , and combined with the The real-time information obtained from the The initial source of transmission of the outbreak.

Claims

1. A dual-cost driven network communication cascade intelligent perception and traceability method, characterized by: The steps include: Step 1: Determine the optimal target monitor set V based on the topological structure information of the basic network o ; For a specific network G(V,E), V and W represent the vertex set and edge set in the network respectively; the size of the given monitor set n o and hyperparameters x, using the efficient hybrid algorithm EHA for the optimal target monitor set V in the first stage TSS-I o The solution of The efficient hybrid algorithm EHA includes: o BPD algorithm of initial solution and monitor set V o EPP algorithm for effective improvement of initial solution; Step 2: Based on the information of the historical propagation set ζ, intelligently optimize the optimal target monitor set V o Configuration, determine the propagation burst perception set For a specific network G(V,E), the monitor set V o And node sequence S, given the outbreak set ζ, introduce the network structure-based strategy NSS and the greedy-based strategy GDS to determine the monitor set V o The order of checking nodes in the middle, intelligently constructing the burst-aware set described in the second stage TSS-II Step 3: Based on the burst perception collection Combined with specific outbreaks The real-time information obtained from the The initial source of transmission of the outbreak; For a specific network G(V,E), the perception set For the third stage TSS-III, the greedy strategy PGS based on percolation and CIR algorithm are used to intelligently optimize the candidate set V by dynamically checking the nodes. c and minimize the set of dynamic inference monitors 2. A dual-cost driven network communication cascade intelligent perception and traceability method according to claim 1, characterized in that: The acquisition monitor set V o The BPD algorithm for the initial solution is as follows: Step 1-1: Create an empty monitor set V o , and assign an initial position to the node sequence S; initialize a counter to track the order in which the nodes are added; define G′(V o ) is the remaining network after removing the monitor set; Step 1-2: When G′(V o ), that is, when there is a cycle in the remaining network G′=(V′,E′): first calculate the removal probability of each node Select the node u with the highest probability of removal from the remaining network, that is: Add node u to the monitor set V o And record it in the sequence S(i), update the counter, and add 1 to i; V′ and E′ respectively represent the vertex set and edge set in the remaining network after removing the monitor set; Step 1-3: When i≤n: select the remaining network G′(V o )’s largest connected component LCC has the largest disassembly effect node u, add node u to the monitor set V o And record it in the sequence S(i), update the counter, i plus 1; Step 1-4: Repeat steps 1-2 and 1-3 until the remaining network G′ has a cycle and the loop ends, obtaining the monitor set V o Initial solution.

3. The dual-cost driven network communication cascade intelligent perception and traceability method according to claim 2 is characterized in that: The monitor set V o The EPP algorithm for effectively improving the initial solution is as follows: Step 1-5: According to the node sequence S, randomly select a subinterval S(l1:l2) of the node sequence; l1 and l2 are two randomly generated parameters for selecting the node sequence subinterval, 1≤l1 <l2≤n; Step 1-6: Optimize the node subinterval S(l1:l2) according to the RR relationship strategy to reduce the size of the largest connected component LCC and optimize the order of node deletion; Step 1-7: Get the improved monitor set V o .

4. The dual-cost driven network communication cascade intelligent perception and traceability method according to claim 3 is characterized in that: The network structure-based strategy NSS is as follows: Step 2-1: Initialize index i to 1 and initialize the outbreak perception set is an empty set; Step 2-2: Until an outbreak is detected Or before i exceeds the size of the monitor set, execute the following loop: select node u from the node sequence S, where u = S(i), and then add node u to the outbreak-aware set In the above example, update index i plus 1; Step 2-3: End the loop and get the burst perception set The greedy-based strategy GDS is as follows: Step 2-4: Initialize the outbreak awareness set is an empty set; Step 2-5: Before the desired detection rate is reached, perform the following operations: o Select a node u from the set so that u is added to the outbreak-aware set The objective function can be minimized And add node u to the outbreak-aware set middle; Step 2-6: End the loop and add to the burst perception set The order of the nodes in , update the node sequence S; Step 2-7: Get the burst perception set 5. The dual-cost driven network communication cascade intelligent perception and traceability method according to claim 4 is characterized in that: The CIR algorithm includes: Step 3-1: Use burst-aware aggregation Initializes the dynamic inference monitor collection Step 3-2: When there is a current set of monitors that has not been checked Or the current candidate set The requirements have not been met yet, and the following loop is executed: i) Based on the dynamic inference monitor set Get Subnetwork ii) Update the current monitor set iii) Use the greedy strategy PGS based on percolation to extract the current monitor set Select the next node u to be checked d iv) Select the node u d Join the dynamic inference monitor collection Step 3-3: End the loop and update the candidate set Step 3-4: Get the candidate set V c .

6. The dual-cost driven network communication cascade intelligent perception and traceability method according to claim 5 is characterized in that: The greedy strategy PGS based on percolation is as follows: Step 3-5: When the potential monitor set V to be checked o t′ When the number of nodes in is greater than 1, execute the following loop: i) Select a node u such that: ii) From a given subnetwork V o t′ Remove node u; α′(v) represents the subnetwork G t Remove V o t′ The connected components of the remaining network after all nodes are quantized; Step 3-6: When the loop ends, the last remaining node v is the next node u to be checked d ; Step 3-7: Get the next node u to be checked d .

7. A device using the network communication cascade intelligent perception and traceability method according to claim 1, characterized in that: include: The acquisition unit obtains the optimal target monitor set V according to the topological structure information of the basic network o ; The perception unit intelligently optimizes the monitor V according to the information of the historical propagation set ζ o Configuration, determine the propagation burst perception set The tracing unit, based on the outbreak perception set Combined with the specific outbreak The real-time information obtained from the The initial source of transmission of the outbreak.

Citation Information

Patent Citations

  • Belief State Determination for Real-Time Decision-Making

    US20220382279A1

  • Method and system for determining information dissemination source, electronic device, and storage medium

    WO2024108913A1