Parameter estimation method and system based on unbiased contrast divergence index random graph model

Through the parameter estimation method of the unbiased contrast divergence index random graph model, the coupled sampling and difference compensation technology of the main chain and the secondary chain are used to solve the calculation complexity and estimation deviation problems in large-scale complex network analysis, and efficient and accurate parameter estimation is achieved, which is suitable for social networks, financial transaction networks, recommendation system networks and biomolecular networks.

CN120387259APending Publication Date: 2025-07-29HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468120.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing index random graph models have problems with high computational complexity, estimation accuracy and deviation in large-scale complex network analysis, which is difficult to meet the needs of real-time or high concurrency, especially in the fields of social networks, information dissemination, etc.

Method used

The unbiased contrast divergence index random graph model is adopted, and the main chain and the secondary chain are run in parallel for coupled sampling, combined with finite step difference compensation and log-likelihood function center-of-gravity processing, an unbiased gradient is constructed for parameter iterative optimization to ensure the accuracy and efficiency of the estimation results.

Benefits of technology

It significantly shortens the parameter estimation time, ensures the unbiasedness and accuracy of the estimation results, adapts to the real-time analysis needs of large-scale complex networks, and improves computing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387259A_ABST
    Figure CN120387259A_ABST
Patent Text Reader

Abstract

The invention provides a parameter estimation method and system based on an unbiased contrast divergence index random graph model, and the method and system are used for analyzing a complex relation network, and the method comprises the following steps: constructing an index random graph model for an observation network, and setting the initial parameters of the model; under the current parameters, simultaneously operating two Markov chains of a main chain and an auxiliary chain to perform parallel sampling and share a random proposal, and when the two chains reach the same network state in a limited step, recording the step number as a coupling moment and cutting off auxiliary chain sampling; at the coupling moment, finite step difference value compensation is executed; comparing the observation network statistics with model expectation statistics, and constructing a log-likelihood function of a current parameter through a differential form relative to an initial parameter; performing iterative tuning on the model parameters according to the unbiased gradient of the log-likelihood function; and repeatedly executing until the model parameters meet the convergence condition, and outputting a final parameter estimation value. The operation time can be remarkably shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of social network analysis, and particularly relates to a parameter estimation method and system based on an unbiased contrast divergence exponential random graph model. Background Art

[0002] In today's digital age, the scale and complexity of social networks continue to climb, and in-depth analysis and understanding of them are becoming increasingly crucial in many fields, such as social media, organizational relationships, information dissemination, education management, etc. The exponential random graph model (ERGM) is a statistical model that jointly models the node attributes and structural attributes of complex networks. Due to its interpretability and flexibility, it has received extensive attention from academia and industry. However, existing ERGM parameter estimation methods still face relatively prominent challenges when dealing with large-scale complex networks:

[0003] 1. High computational complexity

[0004] Some algorithms that can theoretically provide unbiased estimation results (such as Monte Carlo maximum likelihood estimation, MCMLE) often have an overly heavy computational burden in practical applications. When the number of network nodes and the scale of edges increase, the operation time of MCMLE may increase exponentially, making it difficult to meet application scenarios with strict requirements for real-time or high concurrency. For example, in the analysis of large-scale social networks involving tens of thousands or hundreds of thousands of nodes, MCMLE usually requires long sampling iterations, consuming a large amount of computing power, resulting in limited application value.

[0005] 2. Estimation accuracy and bias issues

[0006] To reduce the operation overhead, researchers have introduced approximate methods such as contrastive divergence (CD) and maximum pseudolikelihood estimation (MPLE), which can complete parameter estimation in a shorter time. However, these methods all have a certain degree of estimation bias. In scenarios with high accuracy requirements (such as fine-grained social research, organizational structure optimization, mining of important information dissemination chains, etc.), the bias may lead to incorrect conclusions; in academic research that requires rigorous quantitative analysis, it is more likely to cause unpredictable impacts. In addition, in management application scenarios, biased estimation is also difficult to achieve the expected intervention or decision-making effects.

[0007] 3. Insufficient adaptability to large networks

[0008] Although some improved algorithms or distributed implementations have been attempted to accelerate the parameter estimation of ERGM, in the case of continuous expansion of network scale and superimposed multi-dimensional attribute features, traditional methods still face bottlenecks in terms of time complexity and resource requirements. If extremely simplified or accuracy-weakened methods are adopted to meet real-time requirements, the estimation results often lose reliability; if precise but time-consuming schemes are adhered to, it is difficult to meet urgent needs such as management decision-making, online analysis, or real-time monitoring.

[0009] 4. Growing demand for typical applications

[0010] Taking the mental health of college students as an example, the academic performance and behaviors of students are often deeply influenced by the peer network. To accurately and efficiently analyze the "peer relationship network of college students", existing tools either cannot obtain analysis results in time due to excessive time consumption, or affect the subsequent decision-making effect due to too large deviations. Similar problems also exist in a wider field of complex network analysis, such as real-time monitoring of social media public opinion, optimization of enterprise organizational structure, and analysis of biomolecular interaction networks.

[0011] In summary, although the exponential random graph model has theoretical advantages and interpretability in characterizing network structure and node attributes, how to significantly shorten the running time while ensuring unbiased estimation has always been a key technical problem not properly solved by existing technologies. Especially in large-scale, complex multi-attribute networks, it is necessary to balance accuracy and meet real-time or near-real-time analysis requirements, which poses a more stringent test to existing algorithms, and a new idea and method are urgently needed to break through. Summary of the invention

[0012] The purpose of the present invention is to solve the deficiencies in the above-mentioned background technology, and provide a parameter estimation method and system based on an unbiased contrast divergence exponential random graph model, which can significantly shorten the running time on the premise of ensuring unbiased parameter estimation.

[0013] The technical solution adopted by the present invention is: a parameter estimation method based on an unbiased contrast divergence exponential random graph model for analyzing complex relationship networks, including the following steps:

[0014] 1) Obtain observed network data including network node attributes and edge information;

[0015] 2) Construct an exponential random graph model for the observed network and set the initial parameters of the model;

[0016] 3) Under the current parameters, simultaneously run two Markov chains, the main chain and the sub-chain, for parallel sampling and share random proposals. When the two chains reach the same network state within a finite number of steps, record this number of steps as the coupling moment and truncate the sub-chain sampling;

[0017] 4) At the coupling moment, finite-step difference compensation is performed: a one-time difference accumulation is performed on the sampling states of the main chain and the secondary chain in each step before the coupling to correct the deviation caused by the finite-step truncation, thereby obtaining an unbiased expectation of the network statistics;

[0018] 5) comparing the observed network statistics with the model expected statistics and constructing the log-likelihood function of the current parameters by taking the difference form relative to the initial parameters;

[0019] 6) iteratively tuning the model parameters based on the unbiased gradient of the log-likelihood function;

[0020] 7) Repeat steps 3) to 6) until the model parameters meet the convergence conditions and output the final parameter estimates;

[0021] 8) Using the final parameter estimation value of the exponential random graph model to perform structural or attribute analysis on the complex relationship network, and output corresponding analysis results.

[0022] In the above technical solution, the log-likelihood function adopts a centroid method, which regards the log-likelihood under the initial parameters as the benchmark value, and increases or decreases the log-likelihood of the current parameter relative to the benchmark value as the new log-likelihood function of the current parameter, so as to reduce irrelevant constant terms.

[0023] In the above technical solution, the new log-likelihood function is expressed in the exponential random graph model as "the difference between the observable network statistics under the current parameters and the normalization constant" minus "the corresponding value of the initial parameters"; wherein the observed network statistics reflect the structural indicators of the network and the influencing factors of node attributes, and the normalization constant represents the normalization term of the model probability distribution.

[0024] In the above technical solution, in the finite-step difference compensation, the preliminary statistics obtained by sampling the main chain are used as a benchmark, and the state differences between the main chain and the auxiliary chain at each step before the coupling moment are gradually added and subtracted and summarized at one time; after detecting that the main chain and the auxiliary chain reach the same network state for the first time, the accumulation of new differences is stopped, so that the estimation of the network statistics remains unbiased under the finite-step truncation condition.

[0025] In the above technical solution, the unbiased gradient of the log-likelihood function is obtained by subtracting the finite-step difference compensation from the observed network statistic to obtain the expected value of the model statistic.

[0026] In the above technical solution, the main chain-side chain coupling sampling adopts a maximum coupling strategy in each sampling step to increase the probability that the two Markov chains reach the same network state within a finite number of steps; this strategy includes:

[0027] Jointly generate a single random proposal, which is then accepted by both the main chain and the secondary chain.

[0028] Calculate the acceptance rate thresholds separately and generate an identical random number to compare the acceptance rates of the main chain and the secondary chain respectively;

[0029] When the random number is lower than the acceptance rate of a certain chain, that chain enters the proposal state; otherwise, it remains in its original state;

[0030] Through the above synchronous proposal and single random number mechanism, the main chain and the secondary chain have a greater chance of reaching the same network state in the same step, thereby triggering the coupling moment and truncating the secondary chain sampling.

[0031] In the above technical solution, after each completion of the main chain-secondary chain coupling sampling, finite-step difference compensation, and model parameter tuning, the main chain and the secondary chain are re-run based on the updated parameters for the next round of coupling sampling, forming multiple iterations; in each iteration, the Markov chain is initialized according to the current latest parameters and sampling and compensation are performed until the model parameters meet the convergence conditions.

[0032] In the above technical solution, the observed network data includes records of node attributes and edge relationships; corresponding network structure feature statistics and node attribute effect statistics are defined in the exponential random graph model to meet the analysis requirements of different scenarios.

[0033] In the above technical solution, the complex relationship network includes social networks, financial transaction networks, recommendation system networks, and biomolecular networks.

[0034] The present invention provides a parameter estimation system based on an unbiased contrast divergence exponential random graph model for implementing the parameter estimation method based on the unbiased contrast divergence exponential random graph model described in the above technical solution, including:

[0035] An input data interface for receiving the node, edge, and attribute information of the complex relationship network;

[0036] A Markov chain sampling module configured to run the main chain and the secondary chain in parallel in each iteration, and reach the same network state within a finite number of steps through a coupling mechanism of shared random proposals and truncate the secondary chain sampling;

[0037] A compensation and log-likelihood function processing module for performing a one-time correction on the network statistics based on the state difference between the main chain and the secondary chain before coupling to obtain an unbiased expectation estimate, and constructing a centered log-likelihood function by comparing the expectation estimate with the observed statistics;

[0038] A parameter update module for iteratively updating the parameters of the exponential random graph model according to the unbiased gradient of the log-likelihood function until the convergence conditions are met;

[0039] An output module for outputting the obtained model parameter estimates, and using the parameters to perform structural or attribute analysis on the complex relationship network, and generating and presenting the analysis results.

[0040] The beneficial effects of the present invention are as follows: The present invention significantly shortens the ERGM parameter estimation time, eliminates the need for long-time sampling by realizing the coupling of the main chain and the secondary chain within a limited number of steps and performing difference compensation, and can obtain stable results faster than the traditional Monte Carlo maximum likelihood parameter estimation method (hereinafter referred to as MCMLE). In the "finite-step difference compensation" link, the residuals generated by truncated sampling are corrected to ensure that the estimation remains unbiased, thus avoiding the biased conclusions generated by contrast divergence and maximum pseudolikelihood. The numerical stability is enhanced by centering the log-likelihood function, effectively reducing the numerical instability or additional overhead during the log-likelihood calculation of large-scale networks. The present invention compares and updates the parameters according to the observed statistics and the model expected statistics after each sampling, forming an iterative and efficient method, and finally converging to an accurate description of the observable network. The present invention can be applied to various complex relationship network analysis tasks such as social networks, financial networks, recommendation systems, and biological networks.

[0041] Further, the present invention avoids performing large-scale absolute value operations every time the log-likelihood is calculated by means of the difference relative to the initial parameters. After the value of the log-likelihood is "centered", the algorithm is more likely to converge when comparing the increments and realizes more reliable iterative updates, improving the numerical stability and efficiency. The expected statistics after the finite-step difference compensation correction cooperate with the centered log-likelihood, which can make the calculation of the gradient more direct and reduce the risk of numerical overflow or instability.

[0042] Further, the present invention defines a log-likelihood splitting method, which is convenient for quickly understanding the constant terms to be retained or eliminated when implementing the algorithm. The description of the functions of the statistics and the normalization constants helps those skilled in the art to successfully reproduce the mathematical / programming implementation of this method and enhances the feasibility. The present invention divides the large constant terms in the log-likelihood and cancels the same parts, which is beneficial to saving resources during repeated calculations in multiple rounds of iteration.

[0043] Further, in the finite-step difference compensation of the present invention, based on the preliminary statistics obtained by sampling the main chain, the differences between the main chain and the secondary chain at each step before the coupling moment are gradually calculated and summarized at one time, so that the finite-step truncation still remains unbiased. By summarizing the differences at one time at the coupling moment, no complex operations are added subsequently, ensuring the simplicity and high efficiency of the method. Using the main chain as a reference can reduce the influence of initial bias or random noise, avoid the drawback of obtaining an unbiased estimate only through long sampling, and improve the estimation accuracy.

[0044] Furthermore, the present invention incorporates the unbiased expectation into the calculation of the logarithmic likelihood gradient; this enables the parameter update direction to accurately reflect the true gap between the observations and the model, uses the "finite-step difference compensation late expectation" instead of simply the truncated sampling expectation, and solves the bias problem caused by CD and MPLE. Adopting the simple form of "gradient = observation - compensated late expectation" facilitates the implementation of algorithms such as gradient ascent / descent, Newton's method, and stochastic gradient, improving compatibility.

[0045] Furthermore, the truncation rule under the coupling mechanism constructed by the present invention innovatively combines multiple judgment conditions. On the one hand, a fixed upper limit of the number of iterations is set, based on the basic consideration of computing resources and time costs, to prevent the exhaustion of computing resources due to unlimited iterations and the inability to obtain results for a long time; on the other hand, a dynamic convergence index is introduced to monitor the state changes of the two Markov chains in real time. This combination of static and dynamic truncation methods greatly optimizes the running efficiency while ensuring the accuracy of the results, effectively shortening the running time of parameter estimation. By jointly generating random proposals and calculating the acceptance rates respectively to accelerate the arrival of the coupling moment, maximum coupling can significantly increase the probability that the two chains become the same in the same step, thus achieving "finite-step truncation" faster. By comparing the acceptance rates with a single random number, the double-chain state synchronization mechanism is greatly simplified, enabling a high probability of jumping to the same state at each step without having to wait for burn-in for a long time. Since the earlier the coupling occurs, the fewer the sampling steps required, it can especially accelerate the algorithm execution in large-scale networks.

[0046] Furthermore, the present invention improves the estimation accuracy through multiple rounds of iteration. After each iteration, the latest parameters can be used to re-couple and compensate, enabling the unbiased gradient to be calculated under the condition of "current parameters" and converging more precisely to the optimal value. In practical applications, it can be detected at any time whether convergence is achieved according to the timeliness requirements. If the network environment changes, a new round of iteration can also be restarted, which has flexibility.

[0047] Furthermore, the present invention is applicable to multi-dimensional networks with node attributes: not only limited to simple node connections, but also able to incorporate node characteristics (such as age, occupation, risk rating, etc.) into the model, and estimate them together with the structural parameters and attribute effects. Using statistics such as homophily effect, activity, and reciprocity, statistics can be flexibly defined for different fields, compatible with various analysis requirements.

[0048] Furthermore, the present invention encapsulates functions such as coupled sampling, difference compensation, logarithmic likelihood processing, and parameter update into system modules, which can be deployed and implemented in specific devices, facilitating engineering implementation. This system can be integrated into enterprise service platforms, cloud analysis tools, or real-time monitoring systems, bringing actual economic value. Brief Description of the Drawings

[0049] Figure 1Schematic flow chart of the method of the present invention;

[0050] Figure 2 Schematic network structure diagram of the present invention;

[0051] Figure 3 Schematic diagram of common network structure features of the present invention;

[0052] Figure 4 Algorithm flow chart of Embodiment 1;

[0053] Figure 5 For the Markov chain s(y i ) and s(y' i ) coupling process schematic diagram;

[0054] Figure 6 Schematic diagram of the time consumption of each method under the small-scale network with random parameters in Embodiment 1;

[0055] Figure 7 Schematic diagram of the time consumption of each method under the large-scale network with random parameters in Embodiment 1;

[0056] Figure 8 Schematic diagram of the time consumption of each method under the large-scale network with fixed parameters in Embodiment 1. Detailed implementation manners

[0057] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, which are convenient for clearly understanding the present invention, but they do not constitute a limitation to the present invention.

[0058] Embodiment 1

[0059] As Figure 1 shown, the present invention provides a parameter estimation method based on an unbiased contrast divergence index random graph model for analyzing complex relationship networks, which is characterized by including the following steps:

[0060] 1) Obtain complex relationship network data including node attribute and edge relationship information;

[0061] 2) Construct an exponential random graph model for the complex relationship network and set the initial parameters of the model;

[0062] 3) Run two Markov chains, the main chain and the sub-chain, for coupled sampling simultaneously, where:

[0063] The main chain performs Markov chain sampling in the network state space according to the current parameters of the model;

[0064] The sub-chain starts from the observed network or a specific initialization state and evolves with a random number generation process shared with the main chain;

[0065] When it is detected that the main chain and the sub-chain first reach the same network state within a finite number of steps, record this number of steps as the coupling moment and truncate the sub-chain sampling;

[0066] 4) At the coupling moment, calculate the finite-step difference compensation based on the state differences of the main chain and the sub-chain in each step before the coupling moment, which is used to correct the expected value of the network statistic of the model under the finite-step truncated sampling to obtain an unbiased estimate;

[0067] 5) Compare the expected value of the model statistic after the finite-step difference compensation with the observed statistic of the complex relationship network, and take the difference of the logarithmic likelihood functions of the two as the logarithmic likelihood function of the current parameter, which characterizes the deviation of the current model parameter relative to the observed network data;

[0068] 6) Construct an unbiased gradient according to the logarithmic likelihood function, and perform iterative update on the model parameter; repeat steps 3)-6) until the parameter meets the convergence condition to obtain the final parameter estimate value of the model;

[0069] 7) Use the final parameter estimate value of the exponential random graph model to analyze the structure or attributes of the complex relationship network, and output the corresponding analysis results.

[0070] Specifically, in step 1), the complex relationship network includes social networks, financial transaction networks, recommendation system networks, and biomolecular networks. The observed network data contains records of node attributes and edge relationships.

[0071] The node and attribute information can be selected based on the application scenario: user attributes (gender, age, interests, etc.) in social networks, account types (personal / enterprise) and risk ratings in financial networks, user activity or commodity categories in recommendation systems, protein function labels in biomolecular networks, etc.

[0072] When structuring the data, categorical attributes can be encoded in the format of labels or dummy variables, and continuous attributes (such as transaction amounts, expression levels) can be stored as numerical values, and standardized or discretized if necessary.

[0073] The edge relationship information between nodes corresponds to Figure 2 、 Figure 3 the shown connections or arrows: representing social connections, transaction behaviors, purchase interactions, or molecular interactions between nodes.

[0074] If the network is directed or weighted, it is necessary to record the direction (such as (i→j)) and weight (such as transfer amount, scoring intensity) in the data. A timestamp can also be added to characterize the evolution of the dynamic network.

[0075] Structurally output the observed network data: form the node and edge information into an input format (such as a node list, an edge list, an attribute table, etc.) that can be read by the model to interface with subsequent ERGM modeling tools. Figure 3 Example relationships such as (a), (b), (c), etc. in it can all be regarded as the observed network substructures extracted in this stage.

[0076] Specifically, in step 2), define corresponding network structure feature statistics and node attribute effect statistics in the exponential random graph model to meet the analysis requirements of different scenarios.

[0077] After obtaining the network data, in step 2), construct an exponential random graph model for the complex relationship network and set the initial parameters of the model. Several structure statistics (such as the number of edges, triangle closure, reciprocity) and attribute effects (such as homophily, sender / receiver effect) can be selected as the sufficient statistics of the model according to the actual analysis requirements. Specific examples are as Figure 2 , Figure 3 shown:

[0078] According to the network type and analysis requirements, select appropriate statistics from the structure indicators exemplified in Figure 2 and Figure 3 as the sufficient statistics of the ERGM model. For example, density edges ( Figure 3 all the connected edges shown in (a) in it, used to characterize the overall edge connection tendency of the network), reciprocity mutual ( Figure 3 the bidirectional edges shown in (b) in it, used to measure the two-way connection of paired nodes in a directed network), triangle closure or k-star structure ( Figure 3 (c), (d), (e) in it, representing the clustering structure of "friends of friends also become friends" in the network).

[0079] When the nodes contain exogenous information such as interests, risk levels, functional categories, etc., homophily or main effect statistics can be defined. For example, Figure 2 "Attribute n" of nodes X1, X2, X3, X4 in it means that matching detection or main effect estimation can be performed on this attribute;

[0080] The homophily parameter reflects the additional connection preference between nodes with the same attributes. For example, students of similar ages are more likely to be friends with each other;

[0081] The main effect (sender / receiver) can reveal that a certain attribute value is more likely to initiate / receive a connection. For example, whether high-credit-score accounts in a transaction network are more often transferred funds.

[0082] Incorporate the selected statistics into the model and define their corresponding parameters, such as Figure 2Middle nodes and edges can respectively carry multi-dimensional attributes. For directed networks, a reciprocity statistic item can be additionally added; for weighted or temporal networks, generalized ERGM or TERGM extensions need to be adopted accordingly.

[0083] During implementation, the initial values of all parameters (such as density, reciprocity, homophily, node main effects, etc.) can be set to 0 or small random numbers, or heuristic values can be given based on prior experience. Through the "centered log-likelihood" process, the log-likelihood calculation can be made relatively stable, avoiding numerical instability caused by too large initial values. Figure 2 The directions of the nodes X1 to X4 and edges shown can be used as a reference for the initial state, and then the parameters are continuously iteratively corrected during subsequent coupled sampling and difference compensation processes.

[0084] Through the operations of the above steps 1) and 2), the method of the present invention can be applied to multiple scenarios such as social networks, financial transaction networks, recommendation system networks, or biological networks in subsequent steps (such as main-chain - sub-chain coupled sampling, finite-step difference compensation, and parameter update), helping to shorten the estimation time and ensure unbiased results, such as Figure 3 All different local structure examples (such as one-way / two-way / multi-party centers, etc.) can be incorporated into the framework of the parameter estimation method based on the unbiased contrastive divergence exponential random graph model proposed by the present invention (hereinafter referred to as UCD-ERGM, and the full name is Unbiased Contrastive Divergence for Exponential Random Graph Models). Combining with node attribute statistics, complex influencing factors such as homophily and node behavior differences in the network can also be characterized, effectively supporting the structural interpretation and decision-making applications of large-scale networks.

[0085] Specifically, in step 3), as Figure 4 shown, the main-chain - sub-chain coupled sampling adopts the maximum coupling strategy in each sampling step to increase the probability that the two Markov chains reach the same network state within a finite number of steps; this strategy includes:

[0086] Generating a random proposal for both the main chain and the sub-chain to update the current network state;

[0087] Calculating the upper limit of the acceptance rate for the main chain and the sub-chain respectively, and generating a common random number U from the range [0, 1];

[0088] If U is less than the upper limit of the acceptance rate corresponding to the main chain, the main chain accepts the proposal, otherwise it maintains the original state;

[0089] If U is less than the upper limit of the acceptance rate corresponding to the sub-chain, the sub-chain accepts the proposal, otherwise it maintains the state of the sub-chain in the previous step;

[0090] Once it is detected that the network states of the main chain and the sub-chain become the same in the same step, it is regarded as successful coupling. Record this moment as the coupling moment and truncate the sub-chain sampling.

[0091] In the sampling of each step before coupling, the state differences between the main chain and the sub-chain are accumulated to unbiasedly correct the model expectation value after truncation in a finite number of steps.

[0092] Specifically, in step 4), in the finite-step difference compensation, based on the preliminary statistic obtained from the main-chain sampling, addition and subtraction operations are gradually performed for the state differences between the main chain and the sub-chain in each step before the coupling moment and summarized at once; after detecting that the main chain and the sub-chain first reach the same network state, the accumulation of new differences is stopped, so that the estimation of the network statistic in the case of finite-step truncation remains unbiased.

[0093] The unbiased expectation expression of the parameter θ is as follows:

[0094]

[0095] where s() represents the statistic corresponding to the network state generated by the Markov chain, y k is the network state of the main chain at the k-th step (the step corresponding to the coupling moment), y i is the network state of the main chain at the i-th step, y i ' -1 is the network state of the sub-chain at the (i - 1)-th step, and no new difference is generated when the states of the two chains are the same after the coupling moment τ.

[0096] Specifically, in step 5), the log-likelihood function adopts a centering method, regards the log-likelihood under the initial parameter as the reference value, and takes the increase or decrease of the log-likelihood of the current parameter relative to this reference value as the new log-likelihood function of the current parameter to reduce irrelevant constant terms.

[0097] The new log-likelihood function in the exponential random graph model is expressed as "the difference between the observable network statistic under the current parameter and the normalization constant" minus "the corresponding value of the initial parameter"; where the observable network statistic reflects the structural indicators of the network and the influencing factors of node attributes, and the normalization constant represents the normalization term of the model probability distribution.

[0098] The expression of the original log-likelihood function is as follows:

[0099] L(θ) = {θ T s(y obs ) - logK(θ)}

[0100] The new log-likelihood function uses the difference between the log-likelihood functions of the current parameter and the initial parameter as the log-likelihood function of the current parameter, and its expression is as follows:

[0101]

[0102] Among them, θ0 represents the initial parameter, where s(y obs ) is the vector of statistics of the observation network, and K() is the normalization constant.

[0103] The expression of the new log-likelihood function can be further written as:

[0104]

[0105] Specifically, in step 6), the unbiased gradient of the log-likelihood function is obtained from the expected value of the model statistics obtained by subtracting the finite-step difference compensation from the statistics of the observation network.

[0106] Taking the partial derivative of the log-likelihood function, the gradient of the log-likelihood function with respect to the parameter θ is defined as:

[0107]

[0108] Among them, E θ s(Y) represents the unbiased expectation obtained by compensating the parameter θ through the finite-step difference.

[0109] Specifically, in step 6), after each completion of the main-chain - sub-chain coupling sampling, finite-step difference compensation, and model parameter tuning, the main chain and the sub-chain are re-run based on the updated parameters for the next round of coupling sampling, forming multiple iterations; in each iteration, the Markov chain is initialized according to the current latest parameters and sampling and compensation are performed until the model parameters meet the convergence condition.

[0110] Specifically, in step 7), in the social network analysis application, if the parameters related to "activity" or "popularity" are significantly positive, it indicates that some nodes have a strong tendency in initiating / receiving relationships, and potential "KOLs" or "opinion leaders" can be located by calculating the out-degree / in-degree of these nodes and comparing with the average level. The analysis results can include an influence ranking list or a list of key nodes, which is convenient for platform operators to formulate targeted marketing or community operation strategies.

[0111] If the estimated value of the homophily parameter corresponding to a certain attribute (such as interest tag, region) is positive and significant, it indicates that users with similar attributes in this social network are more likely to connect to each other, and community automatic division suggestions or homophily degree rankings can be output to provide support for "attribute-based friend recommendation" or "circle operation" for the social platform.

[0112] If the transitivity parameter (triangles or gwesp term) is significantly positive, it indicates a strong clustering effect of "friends of friends also become friends" in the social network. A visualization graph of the community structure or a clustering coefficient report can be output to assist network managers in understanding community cohesion.

[0113] In the analysis of financial transaction networks, if parameters related to "reciprocity" (mutual) or "triangle closure" are significantly positive, it indicates the existence of frequent two-way transactions or multi-node closed-loop cycles. Output these nodes, edges, and their parameter values as key monitoring objects to assist the risk department in investigating suspicious fund flows, such as anti-money laundering or abnormal trading groups.

[0114] If the main effect parameter corresponding to a certain attribute (such as enterprise account / personal account) is positive, it means that these nodes are more inclined to initiate or receive funds; if the significance is high, the statistical analysis results of the attribute and transaction activity can be output, such as "enterprise accounts tend to have higher in-degree and out-degree". This helps financial institutions adopt differentiated risk control measures for different account types.

[0115] In the application of recommendation system network analysis, interpret items such as attribute matching and interest homogeneity in the final parameters. If it shows a strong positive parameter between a certain category of products and a certain user attribute, customized recommendation rules can be output (for example, "young users significantly prefer specific types of trendy products").

[0116] If the parameters related to user-user similarity items or item-item co-purchased items are large, it indicates a strong collaborative filtering phenomenon, which can assist the recommendation system in generating a correlation matrix; in actual deployment, similar products can be automatically pushed to users with similar attributes or behaviors to improve the recommendation accuracy.

[0117] In the application of biological network function analysis, if the homogeneity parameter between co-functional proteins or similar genes is high, a molecular list of "possibly involved in the same pathway" can be output to guide experimenters for further verification.

[0118] If the estimated "activity" parameter of a certain molecular node is high and the connection objects are distributed across multiple functional modules, it can be marked as a potential "Hub" protein / gene when outputting the analysis results, indicating that it may have multiple functions or a key regulatory position.

[0119] If the triangle closure and clustering parameters are significantly positive, it can show that "there are clear functional complexes or modules in this biological network", and a molecular list of functional modules can be automatically output for key attention by biological researchers.

[0120] In any field, the above parameter estimates and significance indicators can be output to managers or researchers through reports and visualization tools. For example, generating centrality heatmaps, ranking the intensity of attribute homogeneity, visualizing high-reciprocity or closed-loop sub-networks, etc., enables users to clearly identify potential risk points or opportunity points in the network structure at a glance.

[0121] In practical systems such as social platforms, financial platforms, or scientific research platforms, a real-time monitoring or dynamic early warning mechanism can be constructed in combination with the results of the UCD-ERGM algorithm described in this application: Once new data flows in, possible abnormal nodes or emerging groups can be predicted based on the updated parameters, enabling more timely and accurate decision-making interventions.

[0122] In summary, by performing structural and attribute analysis on the final parameter estimates in step 7), the abstract UCD-ERGM output can be transformed into actionable insights and specific decision-making guidance, enabling various applications such as business operations, risk control early warning, and scientific research in a wide range of complex relationship network scenarios, and truly forming a closed-loop from data modeling to the implementation of application value for the method described in the present invention.

[0123] To further illustrate the principle of the present invention, the mathematical formula of the present invention is derived and demonstrated below.

[0124] In the exponential random graph model, the probability of an event is expressed as:

[0125]

[0126] where Y represents the random state of the network, y is its specific instance, representing the specific network state of a main chain; y* represents the random state. θ is the model parameter, representing the effect size of each network feature. s(y) is the vector of network feature statistics, such as node degree, triangle closure structure, and reciprocity. K(θ) is the normalization constant to ensure the validity of the probability distribution. Υ is the set of all possible networks.

[0127] The log-likelihood function L(θ) of the exponential random graph model is expressed as: L(θ) = θ T s(y obs ) - logK(θ), where y obs represents the observed network. In a directed network with n nodes, K(θ) contains 2 n(n-1) possible networks, making the calculation complex.

[0128] This embodiment proposes a UCD-ERGM parameter estimation method. UCD-ERGM can effectively integrate the advantages of CD and MCMLE, thereby shortening the running time while maintaining unbiased parameter estimation. The UCD-ERGM parameter estimation method, like contrastive divergence (CD), is also optimized based on the stochastic gradient descent algorithm. This embodiment introduces a second Markov chain, enables the two Markov chains to run collaboratively, and constructs a coupling mechanism to terminate the Markov chain iteration.

[0129] Two assumptions are proposed for the Markov chain:

[0130] Hypothesis 1 The Markov chain s(y i ) satisfies that when i → ∞, E(s(y i )) → E θ s(Y);

[0131] Among them, E(s(y i )) represents the expected value of s(y i );

[0132] E θ s(Y) represents the expected value under the distribution of the Markov chain s(Y) of the random state at theta.

[0133] Hypothesis 2 Define the expected difference EΔ i = E(s(y i )) - E(s(y i-1 )) to satisfy

[0134] Among them The infinite sum of the absolute value of the above difference is constrained to be a finite value, aiming to illustrate its convergence.

[0135] Two assumptions are proposed for the second Markov chain s(y i '):

[0136] Hypothesis 3 s(y i ) and s(y i ) have the same marginal distribution;

[0137] Hypothesis 4 After a certain random moment τ, it satisfies s(y i ) = s(y i ' -1 ) for all i ≥ τ.

[0138] Furthermore, a new log-likelihood function is constructed:

[0139]

[0140] In the formula, θ0 represents the initial value, that is

[0141] By resolving, the problem that K() is too complex to calculate is cleverly solved.

[0142] For the log-likelihood function Take the partial derivative to obtain the gradient of the log-likelihood, expressed as:

[0143]

[0144] Through Assumption 1 and Assumption 2, the sample mean can be made equal to the expectation, and thus unbiased estimation of the parameters can be achieved.

[0145] Specifically, through Assumption 1, we get

[0146]

[0147] That is,

[0148] From Assumption 2, we know that and then the expectation E θ of s(Y) is obtained as an unbiased estimate, expressed as:

[0149]

[0150] Through Assumption 3 and Assumption 4, a Markov chain coupling mechanism can be constructed to terminate the Markov chain iteration, thereby shortening the running time of parameter estimation.

[0151] Specifically, through Assumption 3, the expected value expressions of two Markov chains with respect to θ are obtained, expressed as:

[0152]

[0153] where s(y i ) represents the first Markov chain, and s(y i ' -1 ) represents the second Markov chain.

[0154] According to Assumption 4, when i ≥ τ, we have and then the expected value of θ is expressed as:

[0155]

[0156] Unbiasedness is achieved by introducing the second Markov chain.

[0157] The final gradient of the log-likelihood function is obtained:

[0158] Use the above formula to iteratively update the estimated parameters: Furthermore, the optimal solution is obtained.

[0159] In this embodiment, a method comparison experiment is designed to verify the performance of the parameter estimation method based on the unbiased contrast divergence index random graph model proposed in the present invention. Considering the actual application of parameter estimation analysis, networks usually have different characteristics, and these characteristics may have different impacts on the performance of the UCD-ERGM parameter estimation algorithm.

[0160] Specifically, the quality of parameter estimation is evaluated by using bias, variation range, and overall accuracy as indicators: the absolute relative bias (ARB) is used to record the bias; the standard deviation (SE) of the parameter estimation result is used to measure the inconsistency of the parameter estimation result; and the root mean square error (RMSE) is used to measure the overall accuracy of the parameter estimation.

[0161] ARB represents the ratio of the absolute deviation of the parameter estimation result to the true value. The calculation formula of ARB is:

[0162] In the formula, m represents the number of experiments, represents the parameter estimation result of the i-th experiment, and θ represents the true value of the parameter estimation. The smaller the ARB, the smaller the bias degree of the parameter estimation result.

[0163] The calculation formula of SE is:

[0164] The smaller the SE, the more efficient and accurate the estimation method.

[0165] The calculation formula of RMSE is:

[0166] The smaller the RMSE, the more accurate the parameter estimation result is on the overall level.

[0167] In this experiment, the algorithm proposed in the present invention is compared with other ERGM parameter estimation algorithms, including common ERGM parameter estimation algorithms: MCMLE, CD, and MPLE. Keeping the initially input network unchanged, different algorithms are used to estimate the parameters of this network, and the experimental results and experimental running times are compared.

[0168] The process of the method comparison experiment is as follows: for a comprehensive performance comparison, the comparison experiment is divided into two categories. The networks involved in the experiment all contain 5 network statistical characteristics, namely edges, mutual, transitive, sender, and receiver. The experimental simulation networks are all simulated and generated by Pnet software.

[0169] The first type is a random parameter network that follows an exponential distribution. The network scale includes small networks and large networks. The number of nodes in the small network is 285, and the number of nodes in the large network is 2000. Set the density of the random network to 0.03 to ensure that the network conforms to the characteristics of sparse social networks. First, use Pnet software to generate a simulated random parameter network, use different ERGM parameter estimation algorithms to estimate the parameters of the same network, record the generated parameter estimation values and the algorithm running time, and finally conduct a parameter performance comparison, as shown in Tables 1-4.

[0170] Table 1 Mean values of each parameter of each method under the small random parameter network.

[0171]

[0172] Table 2 ARB, RMSE, and SE of each parameter of each method under the small random parameter network.

[0173]

[0174] Table 3 Mean values of each parameter of each method under the large random parameter network

[0175]

[0176] Table 4 ARB, RMSE, and SE of each parameter of each method under the large random parameter network

[0177]

[0178] Figure 6 Shows the time consumption of each method under the small random parameter network. As Figure 6 shown, the parameter estimation running time of UCD-ERGM is much less than that of MCMLE parameter estimation, and greater than that of MPLE parameter estimation, which is in line with expectations.

[0179] Figure 7 Shows the time consumption of each method under the large random parameter network. As Figure 7 shown, the parameter estimation running time of UCD-ERGM is much less than that of MCMLE parameter estimation, and greater than that of MPLE parameter estimation, which is in line with expectations.

[0180] The second type is a fixed parameter network that follows an exponential distribution. The network scale is a large network with 2000 nodes. Set the density of the random network to 0.02 to ensure that the network conforms to the characteristics of sparse social networks. First, use Pnet software to generate a simulated fixed parameter network, use different ERGM parameter estimation algorithms to estimate the parameters of this network, record the generated parameter estimation values and the algorithm running time, and finally conduct a parameter performance comparison, as shown in Tables 5-6.

[0181] Table 5 Mean values of each parameter of each method under a large network of random parameters.

[0182]

[0183] Table 6 ARB, RMSE, and SE of each parameter of each method under a large network of random parameters.

[0184]

[0185] Figure 8 The time consumption of each method under a large network of fixed parameters is shown. As Figure 8 shown, the running time of parameter estimation of UCD-ERGM is much less than that of MCMLE parameter estimation and greater than that of MPLE parameter estimation, which is in line with expectations.

[0186] In summary, compared with the traditional ERGM parameter estimation method, the method proposed in the present invention can shorten the model running time while ensuring unbiased parameter estimation, thereby providing theoretical guidance for decision-makers to make scientific decisions.

[0187] The following gives an implementation example of the present invention in combination with the data analysis task of the college student peer relationship network. This example completely covers the processes of observed data acquisition, model construction, parameter estimation, and structural attribute interpretation, reflecting the application of the method of the present invention in real social network analysis.

[0188] Taking the students of the 2022 grade in a certain university as the research object, the campus card swiping data and questionnaire data are collected, which are used for peer relationship recognition and node attribute acquisition respectively. The peer network edge set is constructed through spatio-temporal co-occurrence analysis and social nomination survey, and node attributes such as gender, only child, class, dormitory information, and scores of each dimension of the Big Five Personality are obtained. After desensitization processing, data cleaning, and format unification, structured inputs are generated: including a node list (with attributes) and an edge list (with direction information), constituting a directed attribute network for modeling.

[0189] An ERGM model is constructed on the above network to describe the structural mechanism of network generation. The model statistics include:

[0190] Endogenous structure variables: the number of edges (edges), reciprocity (mutual), geometrically weighted shared partners (GWESP), popularity concentration (popularity), activity concentration (activity), which are used to characterize density, reciprocal relationships, closed structures, and degree distributions.

[0191] Exogenous attribute variables: including gender, only child, room number, class, and scores of each dimension of the Big Five Personality, which are used as main effect variables and homophily variables respectively, reflecting the influence of node characteristics and attribute similarity on edge formation.

[0192] The UCD-ERGM parameter estimation method of the present invention generally includes the following core steps:

[0193] Main-chain and sub-chain coupled sampling: Construct two Markov chains with different initial states (main chain and sub-chain) and evolve them in parallel under the current parameters. The maximum coupling strategy is adopted during the sampling process to synchronize the two chains within a finite number of steps as soon as possible. Once the states are detected to be the same, record this step as the coupling moment and stop the sampling of the sub-chain.

[0194] Finite-step difference compensation: Before coupling, record the differences in the network statistics of each step of the main chain and the sub-chain, and sum them up at once after coupling. Together with the statistics corresponding to the coupling point state, they form an unbiased estimate of the model expectation value.

[0195] Centered log-likelihood construction: Construct a loss function (such as in the form of L2 norm) based on the difference between the observed statistics and the above unbiased expectation as the model optimization objective, avoiding direct calculation of the normalization constant.

[0196] Unbiased gradient iterative update: Adjust the parameters according to the gradient direction, and repeat the sampling, compensation, and update processes until the model parameters converge to obtain the final estimated value.

[0197] In this embodiment, MCMLE, CD, and MPLE are also used to estimate the parameters. Finally, the parameter estimation results obtained by the four parameter estimation methods are shown in Table 7.

[0198] Table 7 Parameter Estimation Results Table a

[0199]

[0200] Judging from the parameter estimation results, the UCD-ERGM parameter estimation method is closer to MCMLE in estimating most model parameters. The two parameter estimation methods of CD and MPLE have relatively large deviations in parameter estimation such as mutual and only-child sender effects. The ARB index can be used to more accurately evaluate the deviation degree of each parameter estimation method compared with MCMLE.

[0201] Table 8 shows the ARB metrics of several parameter estimation methods. From the ARB results, it can be seen that compared with MCMLE, the bias of UCD-ERGM is generally smaller than that of CD and MPLE. However, in the estimation of the GWESP parameter, it is greater than CD and significantly smaller than MPLE, and in the estimation of the edges parameter, it is greater than CD and MPLE. In addition, the overall parameter estimation results of UCD-ERGM are better than those of CD and MPLE. It is worth mentioning that in terms of the openness of homophily, although UCD-ERGM is much better than CD and MPLE, numerically, there is still a large gap between the parameter estimation results of the three methods and those obtained by MCMLE. However, in the experiment estimated by MCMLE, the parameter estimation results of this variable have large fluctuations and are ultimately not significant. Therefore, the results once again show that on the premise that the running time is much shorter than MCMLE, UCD-ERGM generally has a lower parameter estimation bias.

[0202] Table 8 Parameter Estimation Results Table b

[0203]

[0204] Next, through the UCD-ERGM parameter estimation results shown in Table 9, the collected college student peer relationship network data is analyzed from two aspects: exogenous attribute variables and endogenous structure variables.

[0205] Table 9 UCD-ERGM Parameter Estimation Results Table

[0206]

[0207] Note: In the table, ***, **, *, · indicate that the parameter is significant at the statistical levels of 0.1%, 1%, 5%, and 10% respectively.

[0208] Combined with the model estimation results proposed in the present invention, a significance analysis and result interpretation are carried out for each parameter. Structural statistics such as mutual and GWESP are significantly positive, indicating that the peer network of this group has strong reciprocity and triangle closure characteristics; attribute statistics such as gender and dormitory homophily parameters are significantly positive, indicating that gender consistency and shared living environment promote relationship establishment; the main effects of extraversion and neuroticism scores in the personality dimension on initiating connections are positive, reflecting the driving role of personality traits in peer formation. Finally, a structure influence factor, an attribute effect table, and a node behavior explanation are output to support managers in identifying key nodes and optimizing student social support strategies.

[0209] Example 2

[0210] The present invention provides a parameter estimation system based on an unbiased contrast divergence exponential random graph model for implementing the parameter estimation method based on the unbiased contrast divergence exponential random graph model described in the above technical solution, including:

[0211] An input data interface for receiving node, edge, and attribute information of the complex relationship network;

[0212] A Markov chain sampling module configured to run the main chain and the secondary chain in parallel at each iteration, and reach the same network state within a finite number of steps through a coupling mechanism of sharing random proposals and truncate the sampling of the secondary chain;

[0213] A compensation and log-likelihood function processing module for performing a one-time correction on network statistics based on the state difference between the main chain and the secondary chain before coupling to obtain an unbiased expectation estimate, and constructing a centered log-likelihood function by comparing the expectation estimate with the observed statistics;

[0214] A parameter update module for iteratively updating the parameters of the exponential random graph model according to the unbiased gradient of the log-likelihood function until the convergence condition is satisfied;

[0215] An output module for outputting the obtained model parameter estimation value, and using the parameter to perform structural or attribute analysis on the complex relationship network, generating and presenting the analysis result.

[0216] Embodiment 3

[0217] The present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the parameter estimation method based on the unbiased contrast divergence exponential random graph model described in the above technical solution.

[0218] Embodiment 4

[0219] The present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the parameter estimation method based on the unbiased contrast divergence exponential random graph model described in the above technical solution.

[0220] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0221] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0222] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0223] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0224] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these fall within the protection scope of the present invention

[0225] The content not detailedly described in this specification belongs to the known prior art of those skilled in the art.

Claims

1. A parameter estimation method based on an unbiased contrast divergence exponential random graph model for analyzing complex relationship networks, characterized in that, Including the following steps: 1) Obtain the observed network data including network node attributes and edge information; 2) Construct an exponential random graph model for the observed network and set the initial parameters of the model; 3) Under the current parameters, run two Markov chains, the main chain and the sub-chain, in parallel for sampling and share random proposals. When the two chains reach the same network state within a finite number of steps, record this number of steps as the coupling moment and truncate the sub-chain sampling; 4) At the coupling moment, perform a finite-step difference compensation: accumulate the differences between the main-chain and sub-chain sampling states for each step before coupling at one time to correct the bias caused by the finite-step truncation, so as to obtain an unbiased expectation of the network statistic; 5) Compare the observed network statistic with the model expected statistic, and construct the log-likelihood function of the current parameters in the form of a difference relative to the initial parameters; 6) Iteratively optimize the model parameters according to the unbiased gradient of the log-likelihood function; 7) Repeat steps 3) to 6) until the model parameters meet the convergence condition, and output the final parameter estimate value; 8) Use the final parameter estimate value of the exponential random graph model to analyze the structure or attributes of the complex relationship network, and output the corresponding analysis results.

2. The method according to claim 1, wherein: The log-likelihood function adopts a centering method, regards the log-likelihood under the initial parameters as the reference value, and takes the increase or decrease of the log-likelihood of the current parameters relative to this reference value as the new log-likelihood function of the current parameters to reduce the irrelevant constant terms.

3. The method according to claim 2, wherein The new log-likelihood function in the exponential random graph model is expressed as "the difference between the observable network statistic under the current parameters and the normalization constant" minus "the corresponding value of the initial parameters"; where the observable network statistic reflects the structural indicators of the network and the influencing factors of node attributes, and the normalization constant represents the normalization term of the model probability distribution.

4. The method according to claim 1, characterized in that In the finite-step difference compensation, based on the preliminary statistic obtained from the main-chain sampling, perform addition and subtraction operations step by step for the state differences between the main chain and the sub-chain before the coupling moment and summarize them at one time; after detecting that the main chain and the sub-chain first reach the same network state, stop accumulating new differences, so that the estimation of the network statistic under the finite-step truncation remains unbiased.

5. The method according to claim 1, wherein The unbiased gradient of the log-likelihood function is obtained by subtracting the expected value of the model statistic obtained after the finite-step difference compensation from the observable network statistic.

6. The method according to claim 1, characterized in that The main-chain-sub-chain coupled sampling adopts the maximum coupling strategy in each sampling step to increase the probability that the two Markov chains reach the same network state within a finite number of steps; this strategy includes: Generate a single random proposal jointly, which is used by both the main chain and the sub-chain to judge whether to accept; Calculate the acceptance rate threshold respectively, and generate the same random number to compare the acceptance rates of the main chain and the sub-chain; When the random number is lower than the acceptance rate of a certain chain, this chain enters the proposal state, otherwise it remains in the original state; Through the above synchronous proposal and single random number mechanism, the main chain and the sub-chain have a greater chance of reaching the same network state in the same step, thus triggering the coupling moment and truncating the sub-chain sampling.

7. The method according to claim 1, wherein After each completion of the main-chain - sub-chain coupling sampling, finite-step difference compensation, and model parameter tuning, the main chain and the sub-chain are re-run based on the updated parameters for the next round of coupling sampling, forming multiple iterations; in each iteration, the Markov chain is initialized according to the current latest parameters and sampling and compensation are performed until the model parameters meet the convergence conditions.

8. The method according to claim 1, characterized in that, The observed network data contains records of node attributes and edge relationships; corresponding network structure feature statistics and node attribute effect statistics are defined in the exponential random graph model to meet the analysis requirements of different scenarios.

9. The method according to claim 1, wherein The complex relationship network includes social networks, financial transaction networks, recommendation system networks, and biomolecular networks.

10. A parameter estimation system based on an unbiased contrast divergence exponential random graph model, characterized in that, Used to implement the parameter estimation method based on the unbiased contrast divergence exponential random graph model according to any one of claims 1 - 9, including: An input data interface for receiving node, edge, and attribute information of the complex relationship network; A Markov chain sampling module configured to run the main chain and the sub-chain in parallel during each iteration, and reach the same network state within a finite number of steps through a coupling mechanism of sharing random proposals and truncate the sub-chain sampling; A compensation and log-likelihood function processing module for performing a one-time correction on the network statistics based on the state difference between the main chain and the sub-chain before coupling to obtain an unbiased expectation estimate, and constructing a centered log-likelihood function by comparing this expectation estimate with the observed statistics; A parameter update module for iteratively updating the parameters of the exponential random graph model according to the unbiased gradient of the log-likelihood function until the convergence conditions are met; An output module for outputting the obtained model parameter estimates, and using these parameters to perform structural or attribute analysis on the complex relationship network, generating and presenting the analysis results.