Construction method of multi-layer network link prediction model considering multiple correlation characteristics

By building a multi-layer network link prediction model, extracting node, inter-layer and community correlation characteristics, multiple problems of complex network link prediction in the existing technology are solved, the prediction accuracy and applicability are improved, and the developer recommendation effect in open source collaborative networks is improved.

CN120045779APending Publication Date: 2025-05-27HUBEI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510041632.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When dealing with complex networks, existing link prediction methods have limitations in single-layer network methods, insufficient utilization of community structure information, limited mining of inter-layer correlation characteristics, insufficient generalization capabilities of model, and insufficient accuracy and coverage of developer recommendations in open source collaborative networks.

Method used

A multi-layer network link prediction model construction method that takes into account multiple correlation features is proposed, including extracting node correlation features, inter-layer correlation features and community correlation features, and constructing a multi-attribute decision matrix to generate link prediction scores through shared connection index, improved Fluid Community algorithm and Community Association Likelihood Estimation algorithm.

Benefits of technology

It improves the accuracy and wide applicability of link prediction, can more accurately analyze the collaborative relationships between developers, improve the accuracy and efficiency of developer recommendations, and promote the collaboration efficiency and community activity of open source projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045779A_ABST
    Figure CN120045779A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing a multi-layer network link prediction model considering multiple association features, and the method comprises the steps: obtaining the similarity features between network layers through proposing a common connection index, and combining with local association features extracted based on node similarity measurement and quasi-local association features based on a community detection algorithm (FCM). And topological structure information of the multi-layer network is comprehensively mined. And fusing the above characteristics through a multi-attribute decision analysis method, constructing a comprehensive decision matrix, generating a link prediction score, and realizing accurate prediction of potential links in the multilayer network. According to the model, local, global and quasi-local feature information in a multi-layer network is fully utilized, rich community features are formed by developer recommendation services applied to an open source collaborative network and collaborative behaviors and contribution modes among different developers, active communities and key developers can be effectively identified through community detection and associated feature extraction, and the development efficiency is improved. Therefore, accurate developer recommendation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of complex network analysis, and relates to a method for constructing a multi-layer network link prediction model taking into account multiple correlation features. Background Art

[0002] Link prediction is a core research topic in complex network analysis, and its purpose is to predict potential links in the network that have not yet been observed or that may be formed in the future. Link prediction has important theoretical and practical significance in scenarios such as social networks, biological networks, and scientific collaboration networks.

[0003] The rapid development of open source collaborative networks provides new application scenarios for link prediction technology. In open source collaborative networks, how to effectively identify and recommend potential developers to promote efficient collaboration and sustainable development of projects is a research problem that needs to be solved urgently. Through link prediction technology, the collaborative relationship between developers can be more accurately analyzed, cross-project cooperation opportunities can be explored, and the accuracy and efficiency of developer recommendations can be improved. This not only helps to maintain the activeness of the open source community, but also promotes the sustainable and healthy development of open source projects. Therefore, one of the main objectives of the present invention is to apply link prediction technology to open source collaborative networks to achieve accurate recommendations for potential developers and optimize community collaborative relationships, thereby providing strong support for the sustainable development of the open source ecosystem.

[0004] Complex network data usually has the following characteristics: a large number of nodes, complex edge relationships, highly complex network topology, and may contain multi-level structures. However, existing link prediction methods have the following problems and challenges when dealing with these characteristics: (1) Limitations of the single-layer network approach Traditional link prediction methods (such as common neighbors, resource allocation, and Jaccard coefficient) usually rely on the topological characteristics of a single-layer network for calculation. Although these methods are computationally simple and efficient, they cannot consider the complexity of inter-layer interactions in multi-layer networks. For example, in a social network, the behavior of users on different social platforms may affect link formation at the same time. This cross-platform inter-layer relationship is ignored in the single-layer method, resulting in a decrease in prediction accuracy.

[0005] (2) Insufficient use of community structure information Community structure is an important characteristic of complex networks, which is usually manifested as the phenomenon that some nodes in the network are closely connected while other nodes are sparsely connected. However, existing methods for utilizing community structure mostly focus on node relationships at a single level or within a community, while ignoring potential connections between nodes across communities. This deficiency makes the prediction results weak in capturing complex interactions between communities.

[0006] (3) Insufficient mining of inter-layer correlation features The correlation features between different layers in a multi-layer network are an important source of information for link prediction. However, most existing studies use simple similarity measures, such as Pearson correlation coefficient or cosine similarity, when mining inter-layer features. Although these measures can describe some inter-layer correlations, it is difficult to capture more complex inter-layer interaction characteristics, which limits the applicability of the methods.

[0007] (4) Insufficient model generalization ability Most current link prediction models are optimized for specific networks or specific scenarios and lack the ability to widely adapt to various complex networks. This lack of generalization ability makes it difficult for the models to be applied to networks with different characteristics. Especially when facing dynamically evolving networks, it is even more difficult to ensure the stability and efficiency of the models.

[0008] (5) The demand for developer recommendation in open-source collaboration networks In open-source collaboration networks, how to accurately identify and recommend potential developers is an important challenge. Traditional link prediction methods have achieved good results in single-layer networks and with limited features. However, when facing the complex multi-layer structure and diverse relationships in open-source communities, the accuracy and coverage of recommendations are often insufficient. Open-source projects usually involve dynamic collaborations across multiple platforms, multiple tasks, and multiple developers. The interactions between different layers and the collaboration opportunities across projects need to be fully explored and utilized.

[0009] Regarding the above problems and deficiencies, the main challenges faced by link prediction research are: (1) How to comprehensively utilize multi-dimensional feature information in multi-layer networks In multi-layer networks, the potential links between nodes not only depend on local neighborhood relationships but are also significantly affected by inter-layer interactions and community structures. How to effectively integrate these features becomes a key issue.

[0010] (2) How to characterize the influence of complex interactions between communities In multi-layer networks, the interaction characteristics between different communities (such as the distribution of weak connections and the sparsity of connections between communities) have a significant impact on link prediction. However, existing studies still lack sufficient analysis of these complex interactions.

[0011] (3) How to improve the generalization ability of the method Current link prediction methods perform well on specific datasets, but their performance may drop significantly when applied to scenarios with different network structures. How to improve the generality of the algorithm is an important topic.

[0012] (4) How to accurately mine and utilize multi-level and multi-dimensional correlation features in open-source collaboration networks These characteristics make it difficult for traditional single - feature link prediction methods to be competent, and effectively integrating node, inter - layer, and community features has become an important direction for improving the accuracy and applicability of developer recommendations. At the same time, the model needs to have strong generalization ability to adapt to the needs of dynamic participation of members and frequent changes in project status in the open - source community. Summary of the Invention

[0013] The technical problem to be solved by the present invention is to provide a method for constructing a multi - layer network link prediction model considering multiple associated features, to solve the problems in the prior art such as insufficient utilization of community information, limited mining of inter - layer associated features, and lack of generalization ability; especially focusing on the complex multi - layer network of open - source collaboration networks, by studying the multi - layer network link prediction model, to provide more accurate and effective technical support for developer recommendation and cross - project collaboration in open - source collaboration networks, thereby improving the collaboration efficiency of open - source projects and the activity of the community.

[0014] To solve the above - mentioned technical problems, the technical solution adopted by the present invention is: a method for constructing a multi - layer network link prediction model considering multiple associated features, including the following steps: Step 1, extract node - associated features, including similarity measurement methods such as Common Neighbors (CN), Resource Allocation (RA), and Adamic / Adar (AA). Step 2, extract inter - layer associated features, and calculate the cross - layer structural similarity using the Common Connectivity Index (CCI). Step 3, extract community - associated features, and detect communities using the improved Fluid Communities algorithm (FCM). Step 4, construct the Community Association Likelihood Estimation (GALE) algorithm for extracting community - associated features to extract the community - associated scores of node pairs. Step 5, construct a multi - attribute decision matrix, and generate link prediction scores by integrating multiple features.

[0015] In Step 1, the node - associated features include the Common Neighbors (CN) algorithm, Preferential Attachment (PA) algorithm, Resource Allocation (RA) algorithm, and Jaccard Coefficient.

[0016] In Step 2, use multiple inter - layer similarity measurement methods to extract inter - layer associated features ; This type of feature fuses the hierarchical structure topology information of the multi-layer network, considering the influence of the interaction between different sub-network layers in the multi-layer network on link prediction; In subsequent experiments, the influence of different methods on the link prediction results of the multi-layer network is compared and analyzed; The inter-layer correlation features are obtained through the following methods: Step 2-1. First is the Tanimoto coefficient. This algorithm is based on the Jaccard similarity algorithm and takes into account the unique elements of the two target sets. The difference between the union and the intersection is added to the algorithm to emphasize the influence of the unique elements of the two sets on similarity; For the sets 、 , the Tanimoto similarity coefficient is defined as: ; Step 2-2. Secondly is the Pearson correlation coefficient. The Pearson correlation coefficient, as an index to measure the linear correlation degree of two variables, is also used as a measure of set similarity; This algorithm converts the set into a variable representation. While calculating the size of the variable itself, a quantitative representation of the change trend relative to the average value is added, considering the deviation of the variable value from its mean; The Pearson correlation coefficient of the node sets Γ(A) and Γ(B) is defined as: ; Step 2-3. Then is the cosine similarity CSL. The algorithm is a relatively common method to measure the similarity of two vectors. Similar to the Pearson correlation coefficient algorithm, it also converts the set into a vector. The difference is that the CSL algorithm ignores the size of the vector and only measures the correlation through the direction relationship between the vectors; For the sets 、 , the vector representations of the sets are respectively 、 , and the calculation method of the cosine similarity of the two sets is as follows: ; Step 2-4. Finally is the Common Connection Index (CCI). In a multi-layer network, the number of common connections between different layers reflects the inter-layer synergy and interdependence relationship, and reflects the tightness between layers; Based on the "birds of a feather flock together" principle in social network theory, two network layers with more common connections are more similar in structure or function; For the network layers 、 , the connection number of the node in is , the connection number in is , the number of common connections is , and the inter-layer similarity index of the two-layer network is defined as: .

[0017] In step 3, the community association features are extracted through the following steps: Step 3-1, Detect communities. Use community detection algorithms to partition the multi-layer network and generate a set of community structures: , where the community label of each node a is: ; The algorithm uses a variety of community detection algorithms to partition the multi-layer network; The DER algorithm simulates the community structure as a diffusion process using the concept of entropy; Step 3-2, After community partitioning, construct the community association feature extraction algorithm GALE (Community Association Likelihood Estimation); Extract the community association scores of node pairs through the GALE algorithm. First, propose the concept of the strong influence area of a node. For the original network , define the node cluster whose distance from the node in the network is less than or equal to 3 as , representing the strong influence area of node . The nodes in this area have a strong connection with node ; The specific definition method is as follows: ; Step 3-3, Then define the set of nodes in the network whose distance from node is greater than 3 and less than 6 as , that is, the weak influence area of node ; There is a weak connection between the nodes in this area and node ; In this way, the distribution and scope of influence in the network are more carefully divided, enriching the complexity and depth of network analysis; The specific definition method is as follows: ; For the node pair , define the intersection of the strong influence areas of node and node as the common strong influence domain , and the intersection of the weak influence areas as the common weak influence domain ; The specific definition method is as follows: ; ; Step 3-4, Finally, calculate the community association score based on the node distribution in the common strong influence domain and weak influence domain: ; ; .

[0018] In step 3-1, the core of its algorithm is that during this diffusion process, the total entropy of the network decreases, showing high stability; in the tightly connected areas forming communities, this entropy decrease is more significant; the core concept of the SURP algorithm is the measure of "Surprise", which is used to evaluate the unexpectedness or unusualness of community division in the network; this algorithm is based on statistical significance and evaluates the difference between the observed community division in the network and the random division; the core theory of the SIGNI algorithm is statistical significance testing, which attempts to identify groups of nodes whose internal link density is significantly higher than the random expectation by comparing the difference between the observed network structure and the expected structure of the random network model.

[0019] In step 3-1, the core theory of the EIGEN algorithm is eigenvector centrality, which measures the importance of a node in the network according to the number of its neighbors and the importance of the neighbors; the core theory of the Girvan-Newman algorithm is to remove the edges connecting different communities to detect hidden communities or modules; its algorithm identifies the edges connecting different communities by calculating the betweenness of edges in the network, and reveals the hidden communities in the network based on these edges; after gradually removing these edges, the network is gradually decomposed into multiple internally tightly connected modules or communities.

[0020] In step 3-1, the Fluid Communities algorithm is based on the concept of a dynamic fluid model to simulate the interaction of fluids in the environment; in this algorithm, each community is regarded as a kind of "fluid", flowing and interacting in the network, and this concept of flow and interaction simulates the mutual relationship between community members in the network; finally, based on this algorithm, the Fluid Communities community detection algorithm for multiplex networks (Fluid Communities for Multiplex, FCM) is proposed.

[0021] The FCM algorithm first divides the multiplex network into multiple sub-networks according to the hierarchical structure, and each network layer uses the Fluid Communities algorithm for community detection separately to obtain the community detection results of different network layers; a hyperparameter is introduced in the FCM algorithm, which stipulates that a cluster of nodes with the same community label in at least sub-networks can form the final community in the multiplex network; during the actual calculation process of the FCM algorithm, is set according to the hierarchical structure of the original multiplex network, the larger the value, the sparser the distribution between communities and the tighter the internal connection within the community.

[0022] In step 3-2, the community association feature extraction algorithm GALE is based on the community label distribution of the results of the community detection algorithm: the community structure divided by the community detection algorithm (CDM) provides basic information for community association feature extraction; Weight adjustment: The influence of a node in a community is inversely proportional to the community size. This normalized weight adjustment ensures that the contribution of a single node in a large community will not be overestimated due to its large size; when the total number of nodes in the community where node a is located is large, its contribution to the community association score of the target node pair is adjusted according to the following rules: ; This normalization method reduces the calculation bias caused by differences in community size; Multi-dimensional evaluation of the community influence range: By dividing the strongly influential area and the weakly influential area, the GALE method not only focuses on the direct community connections between nodes, but also comprehensively considers the distribution characteristics of cross-community nodes to more comprehensively evaluate the possibility of nodes forming potential links; Adaptability to dynamic networks: The GALE method processes dynamic network changes through incremental calculation. When new nodes or edges affect the community structure of the network, it quickly adjusts the strongly and weakly influential areas of nodes and the community association scores, so as to maintain the real-time and accuracy of the link prediction results.

[0023] Within the common strongly influential area and the common weakly influential area, the node distribution and its community labels are important factors for evaluating the community association score of node pairs; if there are multiple nodes with different community labels in the common strongly influential area of the target node pair, and the communities to which these nodes belong are large in size, it indicates that the community structure in this area is relatively complex and has high heterogeneity; this complexity reduces the possibility of node pairs forming potential links in this area, resulting in a lower community association score; on the contrary, when the community label distribution in the common strongly influential area of the target node pair is relatively single or mainly concentrated in the same community, the probability of link formation is higher, and the community association score is also correspondingly higher.

[0024] The main beneficial effects of the present invention are as follows: Improve the accuracy and wide applicability of link prediction. By comprehensively using various feature information in multi-layer networks, the prediction performance can be improved, providing more accurate technical support for complex network analysis.

[0025] Promote cross-domain application research. The improvement of the link prediction method can further promote its applications in fields such as social networks, bioinformatics, and scientific cooperation, providing support for data-driven scientific research and business decisions.

[0026] Achieve adaptation to diverse network scenarios. By using the multi-dimensional analysis method of community information, the application requirements of networks with different scales and structural complexities can be better met.

[0027] By studying the multi-layer network link prediction model, more accurate and effective technical support is provided for developers in the open-source collaboration network for recommendation and cross-project collaboration, thereby improving the collaboration efficiency of open-source projects and the activity of the community. Brief Description of the Drawings

[0028] The present invention will be further described below in conjunction with the drawings and embodiments.

[0029] Figure 1 It is a schematic diagram of the link prediction task provided by an embodiment of the present invention.

[0030] Figure 2 It is a flowchart of the TOPSIS framework provided by an embodiment of the present invention.

[0031] Figure 3 It is a flowchart of the ILNC-LP algorithm provided by an embodiment of the present invention.

[0032] Figure 4 It is a flowchart of the MCFMN-LP algorithm provided by an embodiment of the present invention.

[0033] Figure 5 It is a schematic diagram of the FCM process of the community detection algorithm provided by an embodiment of the present invention.

[0034] Figure 6 It is a bar chart of the AUC results of the benchmark experiment of the INLC-LP algorithm provided by an embodiment of the present invention.

[0035] Figure 7 It is a bar chart of the Precision results of the benchmark experiment of the INLC-LP algorithm provided by an embodiment of the present invention.

[0036] Figure 8 It is a bar chart of the AUC results of the benchmark experiment of the MCFMN-LP algorithm provided by an embodiment of the present invention.

[0037] Figure 9 It is a bar chart of the Precision results of the benchmark experiment of the MCFMN-LP algorithm provided by an embodiment of the present invention. Detailed Embodiments

[0038] As Figures 1 to 6 In, a method for constructing a multi-layer network link prediction model considering multiple association features, characterized by including the following steps: Step 1, extract node association features, including similarity measurement methods such as common neighbor (CN), resource allocation (RA), Adamic / Adar (AA), etc.; Step 2, extract inter-layer association features, and calculate the cross-layer structure similarity using the common connection index (CCI); Step 3: Extract community association features and detect communities using the improved Fluid Communities algorithm (FCM). Step 4: Construct the Community Association Likelihood Estimation (GALE) algorithm for community association extraction to obtain the community association scores of nodes. Step 5: Construct a multi-attribute decision matrix and generate link prediction scores by integrating multiple features.

[0039] Example 1 When developing a link prediction algorithm applicable to multiplex networks, the local, global, and quasi-local information in multiplex networks is comprehensively extracted and fused to improve the accuracy and applicability of link prediction. The specific research contents are as follows: (1) A link prediction method combining node association features and inter-layer association features (Inter-Layerand Node Characteristic based Link Prediction, ILNC-LP) is proposed. Inter-layer association features are obtained through various inter-layer similarity measurement techniques to explore the influence of the hierarchical structure on link formation in multiplex networks, and the Common Connection Index (CCI) is innovatively introduced as a supplement to inter-layer similarity measurement. This method integrates these features through a multi-attribute decision analysis framework to introduce hierarchical structure information and improve the accuracy of link prediction in multiplex networks. Through experimental verification on multiple datasets, the proposed algorithm significantly outperforms traditional node similarity-based link prediction methods in terms of prediction performance.

[0040] (2) A link prediction method for multiplex networks combining multiple association features (Mult-CorrelationFeatures based Link Prediction in Multiplex Networks, MCFMN-LP) is proposed. By fusing inter-layer, node, and community association features, the influence of community structure on potential links between nodes in multiplex networks is deeply analyzed, thereby improving the accuracy of link prediction. At the same time, an FCM community detection algorithm improved from the Fluid Communities algorithm is proposed to adapt to the special structural characteristics of multiplex networks. Finally, through the comparative analysis of AUC and Precision evaluation metrics, it is verified that the link prediction method integrating multiple association features is superior to single-feature methods in the multiplex network environment and shows higher performance.

[0041] A variety of inter-layer similarity measurement methods are proposed to extract inter-layer correlation features and describe the influence of sub-networks at each layer on the formation of links between nodes. At the same time, the Common Connection Index (CCI) is innovatively introduced as a new method for inter-layer similarity measurement. In addition, a link prediction method (Inter-Layer and Node Characteristic based Link Prediction, ILNC-LP) that fuses inter-layer correlation features and node correlation features is proposed. Based on the multi-attribute decision analysis framework, this method organically combines node correlation features and inter-layer correlation features, introduces hierarchical structure information on the basis of local features, and improves the link prediction accuracy of multi-layer networks.

[0042] In the application scenario of the open-source collaboration network, the cross-project collaboration relationships among developers and the dynamic interaction characteristics among multiple platforms make the link prediction model have great potential in developer recommendation. By better capturing the complex correlation features in multi-layer networks, the recommendation accuracy for potential developers can be improved, thereby promoting the collaboration efficiency and community activity of open-source projects. Through experiments on multiple datasets, evaluation metrics such as AUC and Precision are used to verify the prediction results. The results show that the proposed algorithm exhibits better prediction performance compared with traditional node similarity-based methods when solving the multi-layer network link prediction problem.

[0043] The node correlation feature extraction module obtains the correlation features between nodes through a variety of node similarity measurement methods. The feature extraction method based on node similarity has high computational performance, wide applicability, and stable prediction performance. Many classic node similarity measurement methods, such as the Jaccard coefficient and the common neighbor coefficient, have relatively simple calculation processes and do not involve complex mathematical operations. Therefore, they can maintain low computational complexity and high operation efficiency when dealing with large-scale networks. These methods only rely on network topology structure information and do not require additional attributes or metadata of nodes, which makes them particularly suitable for large complex networks with limited node attribute information.

[0044] The hierarchical correlation feature extraction module extracts inter-layer correlation features through a variety of inter-layer similarity measurement methods, and comprehensively utilizes the hierarchical structure topology information of multi-layer networks. This module fully considers the complex interactions between different sub-network layers and their influence on the link prediction accuracy. In subsequent experiments, multiple methods will be used for comparative analysis to evaluate the influence of these features on the link prediction performance of multi-layer networks, so as to verify the effectiveness and applicability of the proposed method.

[0045] Adopt a framework based on multi-attribute decision analysis to systematically examine the edge characteristics in a multi-layer network environment. In this framework, the attributes of each edge in the network are regarded as key factors affecting the decision-making process, unconnected node pairs are regarded as potential alternative solutions, and the association characteristics of these node pairs in the sub-network are used as the attribute values of the corresponding attributes. In this way, the inter-layer association characteristics and the association characteristics between nodes are comprehensively integrated, and all edge attributes are used to evaluate the missing links or new links that may be formed in the future.

[0046] The multi-attribute decision analysis method (Multiple-Attribute Decision-Making, MADM) is an analytical method that systematically synthesizes multiple decision attributes or criteria to provide an optimal solution. When facing complex decision-making problems, MADM provides an effective analytical framework. For a decision-making problem, first determine all possible alternative solutions, then identify all key attributes, and select or construct an appropriate evaluation algorithm to assign weights to each attribute, so as to reflect the importance of each attribute in the decision-making process. By combining the scores of all attributes of each alternative solution, the optimal decision-making solution can be found.

[0047] The multi-layer network link prediction method proposed in the present invention provides a new solution for developer recommendation in an open-source collaboration network by integrating node association characteristics, inter-layer association characteristics, and community association characteristics. Aiming at the complex multi-level interaction relationships among developers in the open-source collaboration network, the following application solutions are designed: (1) Data collection and network construction Collect multi-source data from open-source communities (such as GitHub, GitLab), including information on developers' participation in projects, code submission records, writing behaviors, following relationships, and technology stack tags, etc. Construct a multi-layer network based on the collected data, and each layer of the network represents different types of developer relationships.

[0048] (2) Feature extraction and association analysis Extract node association characteristics, inter-layer association characteristics, and community association characteristics in the multi-layer network to identify potential connections between developers. Among them, the community detection module is responsible for grouping developers, mining the potential of community collaboration, and analyzing the implicit cooperation relationships of developers in combination with the interaction information of the multi-layer network.

[0049] (3) Construction of developer recommendation model Using the MCFMN-LP algorithm, integrate the feature data into the multi-attribute decision analysis framework to generate a predictive ranking of potential collaboration relationships. By comprehensively analyzing the association information in the multi-layer network, screen the potential developers that best match the requirements of the target project.

[0050] (4) Dynamic update and adaptive optimization As the open-source collaboration network dynamically changes (such as developers joining new projects, technology stack updates, etc.), regularly update the multi-layer network structure and associated features to maintain the real-time performance and adaptability of the model. And through a feedback mechanism, adjust the model parameters according to the actual recommendation effect to improve the accuracy and coverage of the recommendation.

[0051] (5) Recommendation Results and Application Output Generate a prioritized list of potential developer recommendations for the target project, including the associated features, technology matching degree, and community influence information of each recommended developer. The recommendation results can be directly pushed to the project maintainer through the tool interface of the open-source collaboration platform to support more efficient team formation and collaboration optimization.

[0052] Example 2, As Figure 1 shown, the goal of the link prediction task is to predict the links at a future time or those not yet observed in the dataset, that is, to find the links that may exist but are missing in the current observation. . The common steps of link prediction methods usually first select appropriate similarity metrics or construct likelihood evaluation models based on the existing information in the network. With the development of research, similarity metrics are gradually used in combination to increase the accuracy of link prediction.

[0053] Existing similarity methods are difficult to effectively capture the complex multi-level interaction relationships in the open-source collaboration network, resulting in relatively low accuracy of prediction results. In addition, many similarity-based link prediction methods tend to focus on local or global information while ignoring the comprehensive integration and mining of other topological information. This limitation is particularly prominent when dealing with multi-layer networks such as open-source collaboration networks with diverse types and complex structures. Therefore, developing link prediction methods that can effectively integrate different-level features and adapt to complex network environments has become a key direction for improving the prediction performance of multi-layer networks. The main challenges are as follows: (1) How to capture the structural characteristics of multi-layer networks as a supplement to network features in link prediction methods.

[0054] (2) How to obtain relevant topological information in some densely connected regions of the network to enrich the feature content.

[0055] (3) How to effectively integrate multiple associated features to improve the accuracy of link prediction for multi-layer networks.

[0056] In view of the existing main problems and challenges, a multi - layer network link prediction method is proposed, comprehensively considering multiple association features. First, to capture the structural characteristics of the multi - layer network, it starts from three dimensions: local, quasi - local, and global, and constructs a node association feature extraction module, an inter - layer association feature extraction module, and a community association feature extraction module respectively to capture topological features at different levels. Among them, the community association feature extraction module effectively solves the key problem of how to obtain the topological information of some densely connected regions in the network. The following several node similarity measurement methods are selected to obtain node association features: (1) Common Neighbor algorithm CN ; (2) Preferential Attachment algorithm PA ; (3) Resource Allocation algorithm RA ; (4) Jaccard similarity coefficient ; Extract inter - layer association features using multiple inter - layer similarity measurement methods , and adopt a multiple inter - layer similarity algorithm: (1) Tanimoto coefficient ; (2) Pearson correlation coefficient ; Cosine similarity CSL ; (3) Common Connection Index CCI A new inter - layer similarity measurement method is proposed, namely the Common Connection Index CCI. Traditional inter - layer similarity measurement methods often emphasize the correlation of node degrees in different layers in the data analysis of nodes in different network layers, rather than the actual common connections. For node pairs , if there are connections in both network layers and , it is said that there is a common connection between node and in and , and the number of common connections is counted as 1.

[0057] In a multi-layer network, the number of common connections between different layers reflects the inter-layer cooperation and interdependence, and embodies the closeness between layers. Based on the "homophily" principle in social network theory, two network layers with more common connections are more similar in structure or function. For network layers , , the number of connections of node in is , and the number of connections in is . The number of common connections is . The inter-layer similarity index of the two-layer network is defined as: ; Among them, the first part of the numerator represents the total number of common connections in the two layers and . For each node , represents the number of connections that are simultaneously formed with other nodes in the two layers, and calculates the total number of common neighbors of all nodes in the two network layers. reflects the similarity of direct connections between the two layers. If the two layers have the same connections on many nodes, then will be larger. The second part of the numerator is the total number of connections of all nodes in the two layers, which is the sum of the connection numbers of all nodes in each layer, representing the total activity or total connection density of the two layers. The larger this value, the more connections there are in the two layers in general, regardless of whether these connections match. The denominator represents twice the product of the connection numbers of the two layers. This product form is used to normalize the total number of common connections in the numerator, so that becomes a dimensionless index, keeping the range between 0 and 1.

[0058] This formula is used to measure the structural similarity between two network layers, and the core lies in the normalized ratio calculation. The numerator part reflects the direct structural correspondence between the two layers, while the denominator serves as a normalization factor, considering the connection numbers of each layer to eliminate the deviation caused by the difference in the size of the network layers. Through this normalization process, the formula not only evaluates the absolute number of similarity or the direct number of common connections, but also combines the scale and density of the layers to ensure the fairness of the index, enabling it to make reasonable comparisons between network layers of different sizes and densities. The common connection index measures the structural and functional similarity of network layers by evaluating the common connections between layers, and is particularly applicable to multi-layer network scenarios with direct interactions or functional associations.

[0059] In the open-source collaboration network, there are usually complex multi-level interaction relationships between different projects and platforms. The design of CCI can effectively capture the structural characteristics between these layers and provide more accurate similarity assessment. This is crucial for applications such as developer recommendation, helping to improve the accuracy of prediction and collaboration efficiency, thus supporting the sustainable development of the open-source ecosystem.

[0060] Example 3, The TOPSIS method is adopted to fuse multiple associated features and provide a systematic link prediction scheme. The detailed process of the TOPSIS method is as Figure 2 shown. By mapping the performance of each pair of nodes in different feature dimensions into points in a multi-dimensional space and calculating the distances between these points and the positive ideal solution (i.e., the combination of attributes most likely to form a link) and the negative ideal solution, the possibility of link formation between node pairs is evaluated. This method fully considers the mutual relationships between various features and is especially suitable for scenarios where the inter-layer interactions in a multi-layer network are more significant. Therefore, it is selected as the basic framework for fusing multiple associated features.

[0061] As a multi-attribute decision-making method, TOPSIS can integrate multiple link prediction features and provide a flexible and effective analysis tool for link prediction. Its advantages lie in not only improving the accuracy of prediction but also enhancing the robustness of the prediction model by integrating information at different levels and dimensions. In addition, the TOPSIS method also has significant advantages in analyzing the open-source collaboration network. The open-source collaboration network is usually composed of multi-level and multi-type interaction relationships, and the relationships between nodes are complex and dynamic. TOPSIS can effectively integrate various feature information, evaluate the potential collaboration possibilities between developers, and thus improve the understanding and prediction of developer recommendation and community dynamic evolution, providing strong support for the efficient management and development of the open-source ecosystem.

[0062] Example 4, The architecture diagram of the ILNC-LP method is as Figure 3 shown. First, obtain the inter-layer association features between the target layer and the supplementary layer , and construct an attribute evaluation algorithm to calculate the weights assigned to all attributes. In the multi-attribute decision-making analysis framework, the classic method for determining attribute weights is the information entropy method. The information entropy of an attribute can be used to measure the amount of information provided by this attribute. The key to weight assignment is to ensure that attributes with more information are assigned larger weights. Based on the idea of the information entropy method, according to the inter-layer association features , the weight calculation formulas for the target layer and the supplementary layer are respectively: ; ; Assume that at the target layer In, there is For unconnected nodes, the entire multi-layer network exists The edge attributes are transferred to the multi-attribute decision analysis method, and the total number of alternative decision solutions is , the total number of attributes that affect the decision , so the decision matrix is ​​expressed as: ; In the matrix middle, The row represents the target layer The total number of Alternative decision solutions or unconnected node pairs, Columns represent the total number of multilayer networks Layer network, i.e. There are edge attributes, where the first column of attributes is the target layer The edge attributes of . As an alternative or node pair In Properties The attribute value on The attribute value on the layer edge attribute is the node pair Node-related features on this layer ,exist The attribute value on the layer edge attribute is based on the node pair Whether there is a link in this layer is determined. If the node pair has a link in the supplementary layer network, the attribute value is defined as 1. If there is no link, the attribute value is defined as 0. The specific expressions are defined as: ; ; Standardized decision matrix It is expressed as: ; The classic multi-attribute decision-making scheme TOPSIS (Technique for Order of Preference by Similarity to Ideal Solution) proposes a concept that the optimal decision solution should be closest to the positive ideal solution and farthest from the negative ideal solution. The hypothetical solution with all attribute values ​​being the maximum is the positive ideal solution, and the hypothetical solution with all attribute values ​​being the minimum is the negative ideal solution. This scheme considers the comprehensive distance between the alternative decision solutions and the positive ideal solution and the negative ideal solution, and evaluates the pros and cons of the alternative solutions from the overall effect. For the standardized decision matrix , decision-making scheme about attributes The positive ideal solution of With the negative ideal solution They are respectively expressed as: ; ; Calculate the standardized attribute values using the Gaussian kernel function and and The proximity scores. The higher the proximity score, the closer the distance. In the link prediction problem for multi-layer networks, the correlation between layers is mostly non-linear. The characteristic of the Gaussian kernel function that maps data to a high-dimensional space is more effective for processing non-linear data. Based on the Gaussian kernel function, the standardized attribute values The proximity scores regarding the positive ideal solution and the negative ideal solution are respectively expressed as: ; ; According to the theoretical basis of the TOPSIS method, comprehensively considering the influence of the two proximity scores on the alternative solutions, for the edge attribute decision matrix , in the matrix, the alternative solution Regarding the attribute The ideal score , the calculation formula is as follows: ; Multiply the weights of the target layer and the supplementary layer respectively with the ideal scores of the alternative solution under the corresponding attributes and sum them up to obtain the comprehensive score of the alternative solution , the calculation formula is as follows: ; Sort the comprehensive scores of all alternative solutions from high to low as the link prediction output result. The alternative solution with the highest ranking is the node pair most likely to have missing links or generate new links.

[0063] The above describes the technical details of the ILNC-LP algorithm that fuses node association features and inter-layer association features. On this basis, a link prediction method considering multiple association features (MCFMN-LP) is further proposed, and its specific algorithm process is as Figure 4 shown. Compared with ILNC-LP, MCFMN-LP introduces a community detection algorithm and a community association feature extraction module to more comprehensively describe the multiple association features in the network.

[0064] In this method, the original network is divided into multiple communities by various community detection algorithms such as DER (Diffusion Entropy Reducer), SURP (SurpriseCommunities), SIGNI (Significant Communities), EIGEN (Eigenvector), Girvan - Newman, Fluid Communities, and FCM (Fluid Communities for Multiplex), and corresponding community labels are assigned to each node. Based on these community detection results, the community association feature extraction module extracts community association features, and then incorporates this feature as one of the key attributes into the decision matrix.

[0065] Finally, by fusing node association features, inter - layer association features, and community association features, the comprehensive scores of each alternative are calculated and sorted based on the scores to obtain the final link prediction result. This multi - dimensional feature fusion method can more accurately capture the complex structural relationships in the network, significantly improving the accuracy of link prediction and the robustness of the model.

[0066] Example 5 An improved multi - layer network community detection method is proposed, namely the FCM (Fluid Communities for Multiplex) algorithm based on the Fluid Communities algorithm, and its specific process is as Figure 5 shown. The traditional FluidCommunities algorithm can effectively identify hidden community structures in a single - layer network. However, in the study of multi - layer networks, community detection algorithms are usually applied by simplifying the multi - layer network into a single - layer weighted network, which to some extent weakens the hierarchical structure information unique to multi - layer networks. To address this limitation, an improved version of the FluidCommunities algorithm for multi - layer networks (Fluid Communities for Multiplex, FCM) is proposed to better retain and utilize the hierarchical characteristics of multi - layer networks, thereby improving the accuracy and effectiveness of community detection.

[0067] The FCM algorithm first divides the multi - layer network into multi - layer sub - networks according to the hierarchical structure, and the Fluid Communities algorithm is used separately for community detection in each layer of the network to obtain the community detection results of different network layers. A hyperparameter is introduced in the FCM algorithm, stipulating that a node cluster with the same community label in at least layers of sub - networks can form the final community in the multi - layer network. As Figure 5 shown, network layer Among them, the community is composed of a set of nodes ; the community is composed of a set of nodes ; in the network layer among them, the community and the community are respectively composed of a set of nodes and . Assume that the value is 2. For the multi-layer network composed of the network layer and , in the final community detection result, the community is the set of nodes , and the community is the set of nodes . The calculation method is as follows: ; ; In the actual calculation process of the FCM algorithm, the value of is set according to the hierarchical structure of the original multi-layer network. The larger the

[0068] value, the sparser the distribution among communities and the closer the connections within communities. Build a calculation method for community association scores, CALE (Community Association Likelihood Estimation), based on the likelihood estimation algorithm to calculate community association scores. The feature extraction module first cleans the community detection results. To facilitate subsequent calculations, it is necessary to ensure that all nodes have community labels. For discrete node clusters without assigned node labels or single nodes screened out in the multi-layer network , the algorithm will assign new communities to such nodes in order, even if the community has only one member node. Assume that the total number of communities after the final partition is , the community detection algorithm used is represented as , and the community structure of the network can be represented as . For the node , the assigned community label is represented as: ; ; The total number of nodes in the community where the node is located is represented as .

[0069] Based on the three-degree influence theory, the concept of the strong influence area of nodes is proposed. For the original network , the nodes within the network Node clusters with a distance less than or equal to 3 are defined as , representing the nodes 's strong influence area, where the relationship between nodes within this area and nodes is strongly connected. The specific definition method is as follows: ; Combining the "Three - degree Influence" theory and the "Six - degree Separation" theory, the concept of the weak influence area of nodes is proposed. Specifically, the set of nodes in the network that are at a distance greater than 3 and less than 6 from node is defined as , that is, the weak influence area of node . In this way, the distribution and scope of influence in the network are divided more meticulously, enriching the complexity and depth of network analysis. The specific definition method is as follows: ; For the node pair , the intersection of the strong influence areas of node and node is defined as the common strong influence domain , and the intersection of the weak influence areas is defined as the common weak influence domain . The specific definition method is as follows: ; ; Under the framework of network - divided community structure, considering the characteristics that nodes within a community are more closely connected and nodes between communities are more sparsely connected, a likelihood score estimation method is proposed to calculate the community association score between two nodes. If there are multiple points with different community labels from these two nodes in the common influence area of the target node pair, it means that the community structure within this area is relatively complex, thus reducing the probability of link formation based on the community structure. At the same time, the number of members of the communities to which the nodes with different community labels in the common influence area belong reflects the influence and stability of different communities in the network. Based on these two factors, when the community structure diversity in the common influence area is high and the designed communities have a large number of members, the connection formation probability of the target node pair is relatively low. For node , the specific likelihood function calculation method of the community association score is as follows: ; ; ; The proposed method for calculating community association scores estimates the impact of community structure on link formation by analyzing the distribution of community labels within the common influence area of target node pairs and the differences between these community labels and the community labels of the target node pairs. Essentially, it is based on the core idea of likelihood scoring. In the case of an existing community structure, it evaluates the likelihood of link formation, that is, the data generation process and the associated conditional probabilities. Based on the ILNC-LP algorithm, community association features are integrated to obtain a multi-layer network link prediction model that takes into account multiple association features. The following are the specific implementation steps for integrating community association features.

[0070] Based on the framework of the multi-attribute decision analysis method, the community association features are used as the attributes of the decision matrix, and the community association scores of unconnected node pairs are the corresponding attribute values. Assume that the unconnected node pair The corresponding alternative decision-making scheme is The community association features are added as the last column of attributes to the edge attribute decision matrix in the ILNC-LP algorithm. The specific calculation method of the attribute values is as follows: ; The edge attribute decision matrix After integrating the community association feature attributes, it becomes , which is expressed as follows: ; In the matrix , The submatrix formed by the th row and the previous columns is the edge attribute decision matrix in the ILNC-LP algorithm, and the th column represents the community association feature attributes. According to the ILNC-LP algorithm, calculate the ideal score of the alternative with respect to the edge attribute , and multiply it by the edge attribute weight as the attribute value corresponding to the edge attribute of the alternative in the multiple association feature decision matrix . The specific calculation method of the attribute values in the matrix is as follows: ; Construct the multiple association feature decision matrix according to the attribute values: ; Construct the standardized multiple association feature decision matrix according to the standardized attribute values: ; Finally, the same multi-attribute matrix processing method as the ILNC-LP algorithm is adopted to obtain the final ideal score and comprehensive score: ; ; Sort the comprehensive scores of all alternative solutions as the multi-layer network link prediction result that fuses multiple associated features.

[0071] The specific design and implementation process of applying the multi-layer network link prediction algorithm to the open-source collaboration network developer recommendation model is as follows: (1) Data collection and preprocessing Based on the open-source collaboration data and analysis capabilities provided by the open-source toolset OpenDigger, an efficient and flexible data collection process is invented and designed to comprehensively capture the development behaviors and project relationships in the open-source collaboration network. OpenDigger is an open-source analysis report project for all open-source data initiated by X-lab, which aims to combine the wisdom of global developers to jointly analyze and gain insights into open-source related data. First, define the data source content, and use the GitHub data obtained by the OpenDigger toolset as the core data source, including multi-dimensional information such as developer participation, collaboration behaviors, and project dynamics. The specific data content includes developer activities (such as code commits, Pull Requests, Issue comments, etc.), project characteristics (such as Stars, number of technology branches, etc.), and developer network relationships. Then, batch collect the required data through the standardized data interfaces provided by OpenDigger (such as openrank.json, activity.json, etc.). At the same time, ensure the complete access to the static data set, and use the OpenDigger static data URL to download historical metrics and project data to ensure the breadth and integrity of the collection.

[0072] After determining the data collection method, the model conducts specific multi-dimensional data collection. The data types mainly include developer behavior data, project relationship data, and collaboration network data. Among the developer behavior data, developer OpenRank is used to measure the comprehensive influence of developers, Activity is used to capture the daily activity and contribution frequency of developers, and technical feature indicators are used to analyze the developer's technology stack (programming languages, tool tags, etc.) and their professional fields. Among the project relationship data, project OpenRank and Stars reflect the popularity and technical influence of projects, and participant feature indicators are used to obtain the composition and participation types of project developers. Among the collaboration network data, use interactions such as Pull Requests and Issue comments to construct the collaboration network between developers, and generate the project association network between projects based on the multiple projects participated by developers.

[0073] After the data collection and acquisition process is completed, the data preprocessing module is entered. To improve data quality and provide reliable input for subsequent multi-layer network construction, a multi-level data preprocessing strategy is designed to ensure the accuracy, integrity, and consistency of the data. First, data cleaning is performed. For the collected duplicate or missing records (such as the missing historical data of some developers or projects), they are completed through the data filling function of the OpenDigger toolset and time series interpolation technology, and the Z-score analysis method is used to detect and co-integrate abnormal indicators, such as abnormally high Activity values or distorted OpenRank values. Then, the Min-Max normalization method is adopted to unify the indicator values within the range of [0,1] for fair evaluation in the multi-attribute decision-making model. At the same time, time window alignment is performed on the time series data to ensure that the data has a consistent time span. Feature mapping is performed on the standardized data. First, the technical stack labels of developers are matched with project requirements to generate feature vectors to describe technical relevance, and then multi-level relationships between developers and projects, such as collaboration frequency and technical stack similarity, are extracted to generate a relationship matrix, providing basic data for multi-layer network construction. Finally, based on the principal component analysis PCA, dimensionality reduction is performed on the developer technical feature matrix, retaining the main information to reduce computational complexity, and Gaussian filtering is used to smooth the time series data to reduce the interference of data noise on the model.

[0074] Among them, PCA is a commonly used dimensionality reduction algorithm for extracting main features from high-dimensional data, reducing the data dimension at the same time, and reducing computational complexity. The main goal of PCA is to project the original data into a new coordinate system through linear transformation, so that in the new coordinate system, the main changes in the data are concentrated in a few dimensions. For the dataset X, it contains m samples, and each sample has n features. First, the data standardization step is performed, calculating the mean of each feature and removing it, and then the data is standardized so that the variance of each feature is 1. The specific method is expressed as: ; Then, the covariance matrix is calculated to represent the linear relationship between features. The dimension of the covariance matrix is expressed as , and the specific calculation method is: ; Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors. Among them, the eigenvalue represents the variance size explained by each principal component, and the eigenvector represents the component direction. Sort the eigenvalues from largest to smallest, select the first k principal components as the new unique features, and select the number of principal components whose cumulative explained variance reaches 95%. Finally, project the original data onto the principal components to obtain the dimensionality-reduced data: ; Among them, W is a matrix composed of k principal components.

[0075] (2)Multi-layer open-source collaboration network construction The construction of the multi-layer network is a key link in transforming the collected open-source collaboration data into a structured network representation, providing the basic input for the link prediction model. In view of the complex interaction characteristics of the open-source collaboration network, the present invention designs the following network construction method: 1. Definition of the multi-layer network The multi-layer network is a structure composed of multiple sub-networks, where each sub-network represents a specific type of relationship, and multiple sub-networks interact by sharing nodes or edges. The multi-layer network of the present invention is designed to include three levels, namely the developer collaboration layer, the developer-project layer, and the developer technical association layer. Among them, the developer collaboration layer is constructed based on the direct interaction relationships among developers in behaviors such as code submission, Pull Request, and Issue comment. The developer-project layer represents the participation relationship between developers and projects (such as contributing code, creating Issues, etc.). The developer technical association layer is generated based on the developer's technology stack tags and skill similarity, reflecting the potential technical relevance among developers. The interaction relationships between these layers are integrated by sharing developer nodes and inter-layer links (such as projects jointly participated in by technology).

[0076] 2. Network construction process The model first performs data parsing and layer initialization. Specifically, it extracts multi-dimensional data from data collection and initializes each network layer according to certain rules. For the developer collaboration layer, each developer represents a node, and an edge is established when there is code submission, Pull Request cooperation, or Issue comment between two developers. The edge weight is calculated based on the interaction frequency and is expressed as: ; For the developer-project layer, nodes are represented as two categories: developers and projects. The contribution behaviors of developers to projects (such as code submissions) form edges, and the edge weights are assigned according to the number of contributions or contribution type weights. For the developer technical association layer, it shares developer nodes with the developer collaboration layer and generates edges based on the similarity of developer technology stack tags, where the similarity calculation uses cosine similarity. In terms of constructing the inter-layer interaction relationship, the inter-layer interaction relationship is defined by sharing nodes and edges, providing additional information support for link prediction. In the relationship between the developer-project and the developer collaboration layer, it is defined that the two-layer sub-network shares developer nodes, representing the mapping from the personal behavior of developers to the project collaboration relationship. In the construction of the relationship between the developer technical association and the developer collaboration layer, it is defined that the two-layer self-network shares developer nodes, supplementing the implicit collaboration relationship based on technical similarity. Through the construction of this multi-layer network, the present invention provides accurate and comprehensive structured network input for link prediction and developer recommendation, can accurately depict the relationship between developer behavior and projects, and provides support for the recommendation algorithm. At the same time, it captures the implicit technical connections of developers in the collaboration network and discovers potential collaboration opportunities among developers.

[0077] (3)Feature Extraction and Analysis Feature extraction and analysis are important steps based on the multi-layer network link prediction algorithm. Its goal is to extract key association features from the constructed multi-layer network and reveal potential collaboration relationships among developers through comprehensive analysis. Aiming at the actual needs of open-source developer recommendation, the present invention designs an innovative feature extraction method based on the multi-layer network link prediction algorithm MCFMN-LP, combining node association features, inter-layer association features, and community association features to provide rich and high-quality input for the link prediction model. In the node association feature module, for open-source collaboration, the extraction of contribution type features is added, and the main contribution behaviors of developers in projects (such as code submissions, Issue creation, etc.) are represented by one-hot encoding.

[0078] One-hot encoding is a feature representation method used to convert categorical variables into a numerical vector form that can be processed by machine learning models. The core idea is to use a binary vector to represent each category, where the length of the vector is equal to the total number of categories, and only one position in the vector is 1, and the other positions are 0. Each category corresponds to the only "1" position in the vector. Suppose there is a categorical variable representing the four main programming languages ​​of developers, namely Python, Java, C++ and JavaScript. Then the result of one-hot encoding is [1, 0, 0, 0] for Python, [0, 1, 0,0] for Java, [0, 0, 1, 0] for C++, and [0, 0, 0, 1] for JavaScript. The use of one-hot encoding can adapt to the input requirements of machine learning models. Many models (such as neural networks, linear regression, etc.) cannot directly process categorical variables and need to be converted into numerical form. And one-hot encoding avoids the size order or numerical spacing between categorical variables. One-hot encoding does not introduce priority or weight assumptions between categories, ensuring the independence of features. And the binary vector form is suitable for machine learning models to calculate weights, especially for classification tasks.

[0079] In terms of inter-layer correlation features, we design a cross-layer technology similarity index based on the characteristics of open source collaborative networks, and extract features by combining the inter-layer correlation feature module in the MCFMN-LP algorithm. Specifically, cross-layer technology similarity is to calculate the technical similarity of developers through technology stack labels, capturing the potential collaborative relationship between developers with similar skills. The specific calculation method is expressed as: ; In terms of community association features, the FCM community detection method of the MCFMN-LP algorithm is used to divide the multi-layer open source collaboration network into communities and extract the community information of each developer. At the same time, combined with the calculation of community features, the community center row and the collaboration ratio inside and outside the community are calculated for each developer, and the community boundaries are dynamically adjusted based on the network structure.

[0080] (4) Open Source Collaborative Developer Recommendation Model Recommendation model and multi-attribute decision analysis are the core parts of developer recommendation based on multi-layer network features. This paper builds an efficient and scalable recommendation model by integrating node, inter-layer and community features and combining the multi-attribute decision analysis (MADM) framework. This model can mine potential developer collaboration relationships from multi-layer networks and provide accurate developer recommendations for open source projects.

[0081] In the open-source collaboration network, the goal of the recommendation model is to predict the potential new collaboration links between developers and projects, and generate a prioritized recommendation list for each target project. The core design of the model includes potential link prediction, recommendation priority scoring, and decision support. Among them, potential link prediction uses the multi-layer network link prediction algorithm (MCFMN-LP) to mine the potential collaboration links between developers and projects. Recommendation priority scoring assigns a collaboration priority score for each developer with the target project based on comprehensive feature analysis. Decision support optimizes and ranks the recommendation results through a multi-attribute decision analysis framework.

[0082] The core process of the recommendation model is divided into data input, potential link prediction, priority score generation, and recommendation list output. Among them, data input refers to using the node, inter-layer, and community features provided by the multi-layer network feature extraction module as input data. In potential link prediction, the model calculates the collaboration probability between developers and projects based on the multi-layer network link prediction algorithm. In priority score generation, the model constructs a multi-attribute decision matrix by combining multi-dimensional features to generate a collaboration priority score for each developer with the target project. Finally, in recommendation list output, the developers are ranked according to the priority score to generate a recommendation list.

[0083] The final recommendation results will not only present the matching degree scores obtained based on the multi-layer network link prediction algorithm MCFMN-LP, but also output the technology stack matching degree, community centrality, and collaboration frequency added by the model according to the characteristics of the open-source collaboration network in the feature design and extraction part. The result example is as follows:

[0084] Showing multiple matching degrees (such as technology stack matching degree, community centrality, collaboration frequency, etc.) and related metrics in the recommendation results provides users with comprehensive recommendation references. First, it improves the transparency of the recommendation. The recommendation results contain multi-dimensional feature scores, clearly showing the logic and reasons behind the recommendation, and enhancing the credibility of the results. For example, by showing the "technology stack matching degree", it illustrates the degree of fit between the developer's technical background and the project requirements, and by the "collaboration frequency", it reflects the intensity of the developer's past cooperation with the team. Users can compare and select based on multiple metrics, rather than relying solely on a single comprehensive score.

[0085] Secondly, the model supports personalized needs. Users can independently select the importance of certain metrics according to the specific requirements of the project. If the project requires developers with high technical fit, they can focus on the "technology stack matching degree" metric; if they need core members, they can focus on the "community centrality" metric. The display of multi-dimensional metrics enables the recommendation results to take into account different application scenarios, such as long-term collaboration, cross-team cooperation, or short-term task outsourcing.

[0086] Then, multi-dimensional recommendation metrics can improve the accuracy of recommendations. A single score may obscure the potential advantages of some developers, while multi-metric display can highlight the performance of developers in different dimensions. For example, a developer with a slightly lower "technology stack matching degree" but a higher "collaboration frequency" may still be an important candidate suitable for team collaboration in the recommendation list. The comprehensive score may be affected by extreme values of a single feature, while the display by dimension can reveal the specific distribution of these features and help users make more reasonable judgments.

[0087] Finally, the multi-dimensional result display conforms to the complexity of the open-source collaboration scenario. In the open-source collaboration network, features such as the technical background, community role, and historical contributions of developers may simultaneously affect the recommendation decision. The multi-metric display can comprehensively reflect these complex relationships. For different roles in open-source projects (such as project maintainers and community administrators), multi-dimensional metrics provide more extensive information support to meet the selection needs of different perspectives.

[0088] (5) Dynamic update The behaviors of developers and project requirements in the open-source collaboration network change at any time. Therefore, the recommendation model needs to have the ability to update dynamically to maintain the timeliness and accuracy of the recommendation results. The dynamic update mechanism of the present invention includes dynamic data collection, multi-layer network update, and real-time feedback mechanism. In dynamic data collection, the model will monitor the behaviors of developers in real time and update the behavior data of developers such as code commits, Issue participation, and PullRequest through the GitHub data interface of the OpenDigger platform. At the same time, synchronize the changes in project requirements, monitor the updates of the technology stack tags or task descriptions of the project, and dynamically adjust the recommendation parameters.

[0089] In the multi-layer network update process, the model will dynamically adjust the node and edge weights. According to the newly collected data, update the feature values of nodes (such as developer activity and technology stack similarity) and edge weights in the multi-layer network. For example, the more projects a developer participates in recently, the higher the edge weight of the developer in the developer-project network layer. The model will capture the changes in the relationships between layers. For example, the update of a developer's technology stack may trigger the recalculation of technology similarity.

[0090] Build a real-time feedback mechanism for changes in the open-source collaboration network, allowing users to evaluate the recommendation results (such as marking whether the recommendation is effective) and input the feedback data into the model optimization link. At the same time, dynamically adjust the feature weights according to the user's feedback on the use of the recommendation results. For example, when the user attaches more importance to the technology stack matching degree, increase the weight of this feature.

[0091] Example 6 Benchmark experiments are designed and conducted for the ILNC-LP algorithm and the MCFMN-LP algorithm respectively. First, for the benchmark experiment of the ILNC-LP algorithm, the link prediction methods based on node similarity (including CN, RA, PA, Jaccard) and the benchmark methods for multi-layer networks (MNE, HGAN, LPIS) are selected as controls to verify the effectiveness of ILNC-LP in multi-layer network link prediction. Figure 6 and Figure 7 respectively show the changes in the AUC values and Precision values of the ILNC-LP algorithm and each comparison method on all datasets. In the experimental settings, the influence parameters of the target layer and the supplementary layer in the LPIS method are set to 0.3. ILNC-LP uses the cosine similarity (CSL) as the method for extracting inter-layer association features and selects PA as the method for extracting node association features to fully evaluate the performance based on node similarity. All experimental results are averaged by running each method independently 20 times to ensure the reliability and robustness of the results.

[0092] From Figure 6 it can be clearly seen that the ILNC-LP algorithm is overall superior to other comparison methods in terms of the AUC value, especially on the Aarhus, Lazega, Vivkers, and KjwxData datasets. This is mainly due to the relatively high clustering coefficients of the multi-layer networks in these datasets, which makes the extraction of inter-layer association features more effective. In contrast, traditional link prediction methods based on node similarity such as CN, PA, RA, and Jaccard perform poorly when used as independent benchmarks because they fail to fully consider the hierarchical information of multi-layer networks, resulting in insufficient prediction accuracy. These results further verify the advantages and effectiveness of the ILNC-LP algorithm in multi-layer network link prediction.

[0093] Figure 7 Show the Precision value performance of each method on different datasets. In the experiment, the value in the calculation of Precision is 10. It can be seen from the figure that the ILNC-LP algorithm is overall superior to other methods in terms of precision, especially on the Lazega and Vivkers datasets with relatively high inter-layer similarity. Several multi-layer network methods achieve relatively high precision on these datasets. However, for the relatively large-scale TF and KjwxData datasets, regardless of the method used, their Precision scores are relatively low. Although the KjwxData dataset has a relatively high AUC value due to its high degree heterogeneity, its large network scale increases the difficulty of identifying effective potential links in the first ten predictions, thus reducing the prediction accuracy. Generally speaking, the ILNC-LP algorithm performs more excellently and stably than other benchmark methods on multi-layer network datasets.

[0094] Design a benchmark comparison experiment for the MCFMN-LP algorithm to systematically verify its performance. The experiment selects link prediction methods based on node similarity (such as CN, PA, RA, and Jaccard similarity coefficient), benchmark methods for multi-layer networks (MNE, HGAN, LPIS), and the proposed ILNC-LP algorithm that fuses node association features and inter-layer association features as controls, aiming to evaluate the effectiveness of introducing community association information in MCFMN-LP. Figure 8 and Figure 9 Respectively show the AUC values and Precision values of each method on different datasets.

[0095] In the experiment, the ILNC-LP algorithm uses CSL and PA to extract inter-layer association features and node association features, while the MCFMN-LP algorithm uses the EIGEN algorithm for community detection, uses the CSL method to extract inter-layer association features, and uses PA and RA respectively as the extraction methods for node association features. All experimental results are the averages after each method runs independently 20 times to ensure the robustness and reliability of the results.

[0096] From Figure 8 It can be seen that the MCFMN-LP algorithm significantly outperforms other comparison methods in terms of AUC on multiple datasets, especially when using RA as the node association feature extraction method. This is because community association features can highlight the connection patterns within the community, and the RA method allocates resources based on the number of common neighbors. Within the community, nodes usually have more common neighbors, thus enhancing the effectiveness of the RA method in link prediction within the community. In addition, the community structure usually means the existence of highly clustered node groups, and the RA method utilizes this clustering effect to effectively predict new links through the common neighbors of nodes within the community. Community association features provide a more explicit community boundary and internal structure for the RA method, thus significantly improving the prediction accuracy. Generally speaking, the introduction of community association features further strengthens the high-clustering characteristics of the network, improves the link prediction accuracy of the RA method within and between communities, enabling the algorithm to make full use of the clustering characteristics and inter-community interactions in the network, thereby improving the overall performance of link prediction.

[0097] Figure 9Show the performance of each method in terms of Precision. MCFMN-LP is also superior to other comparison methods. Notably, the performance of the ILNC-LP algorithm is relatively mediocre on large-scale datasets (such as TF, Friendfeed_ita, and KjwxData), while the performance of MCFMN-LP on the TF and KjwxData datasets is significantly optimized. Large-scale datasets usually contain rich community structures, including diverse communities and subgroups. By introducing community detection algorithms and extracting community association features, the MCFMN-LP algorithm can better capture these complex community structures, understand the relationships between nodes more accurately, help alleviate the sparsity and noise problems in large-scale network data, and thus improve the accuracy of link prediction. In large-scale networks, internal links within communities may dominate, and the introduction of community association features enables the algorithm to more effectively mine potential links.

[0098] In summary, the addition of community association features enables the MCFMN-LP algorithm to make full use of different types of features for link prediction, thus better adapting to the complexity and diversity of large-scale datasets. By comprehensively capturing community structures, improving the accuracy of internal link prediction within communities, and deeply mining the relationships between communities, the MCFMN-LP algorithm significantly enhances the overall accuracy of link prediction. Based on the experimental results, the MCFMN-LP algorithm shows extremely high applicability in the application scenarios of open-source collaboration networks. Especially after introducing community association features, it can be effectively used for the recommendation analysis of open-source developers. Open-source collaboration networks usually contain multi-level complex relationship structures, and the collaboration behaviors and contribution patterns among different developers form rich community features. Through community detection and association feature extraction, the MCFMN-LP algorithm can more accurately identify active communities and key developers, thus achieving precise developer recommendations. This is of great significance for improving the collaboration efficiency and community health of open-source projects, helping to identify potential core contributors, promoting more efficient collaboration processes, and driving the sustainable development of the open-source ecosystem.

Claims

1. A method for constructing a multi-layer network link prediction model taking into account multiple correlation features, characterized in that: The steps include: Step 1: Extract node association features, including common neighbors (CN), resource allocation (RA), Adamic / Adar (AA) and other similarity measurement methods; Step 2: extract inter-layer correlation features and use the common connectivity index (CCI) to calculate the cross-layer structural similarity; Step 3, extract community association features and use the improved Fluid Communities algorithm (FCM) to detect communities; Step 4: Construct a community association extraction algorithm, Community Association Likelihood Estimation (GALE), to extract community association scores of node pairs; Step 5: Construct a multi-attribute decision matrix and generate a link prediction score by integrating multiple features.

2. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 1, characterized in that: In step 1, the node association features include the common neighbors algorithm CN (Common Neighbors), the preferential search algorithm PA (Preferential Attachment), the resource allocation algorithm RA (Resource Allocation) and the Jaccard similarity coefficient Jaccard (Jaccard Coefficient).

3. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 1, characterized in that: In step 2, multiple inter-layer similarity measurement methods are used to extract inter-layer correlation features. ; This type of feature integrates the hierarchical topology information of the multi-layer network and considers the impact of the interaction between different sub-network layers in the multi-layer network on link prediction; in subsequent experiments, the impact of different methods on the link prediction results of the multi-layer network is compared and analyzed; the inter-layer correlation features are obtained by the following method: Step 2-1, first is the Tanimoto coefficient. This algorithm is based on the Jaccard similarity algorithm, adding the consideration of the unique elements of the two target sets, adding the difference between the union and the intersection in the algorithm, and emphasizing the influence of the unique elements of the two sets on the similarity; for the set , , the Tanimoto similarity coefficient is defined as: ; Step 2-2, followed by the Pearson correlation coefficient, is an indicator of the linear correlation between two variables and also a measure of set similarity. The algorithm converts the set into a variable representation, calculates the size of the variable itself, and adds a quantitative representation of the trend of change relative to the average value, taking into account the deviation of the variable value from its mean. The Pearson correlation coefficient of the node set Γ(A) and Γ(B) is defined as: ; Step 2-3, followed by cosine similarity CSL, is a common method for measuring the similarity between two vectors. It is similar to the Pearson correlation coefficient algorithm in that it also converts a set into a vector. The difference is that the CSL algorithm ignores the size of the vector and only measures the correlation by the directional relationship between the vectors. , , the vector representation of the set is , , the cosine similarity of two sets is calculated as follows: ; Steps 2-4, finally the Common Connection Index (CCI). In a multi-layer network, the number of common connections between different layers reflects the synergy and interdependence between layers, and reflects the closeness between layers. Based on the "like attracts like" principle in social network theory, two network layers with more common connections are more similar in structure or function. , ,node exist The number of connections is ,exist The number of connections is , the total number of connections is , the inter-layer similarity index of a two-layer network is defined as: .

4. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 1, characterized in that: In step 3, community association features are extracted through the following steps: Step 3-1, detect the community, use the community detection algorithm to divide the multi-layer network into communities and generate a community structure set: , where the community label of each node a is: ; Its algorithm uses a variety of community detection algorithms to divide the multi-layer network into communities; the DER algorithm uses the concept of entropy to simulate the community structure as a diffusion process; Step 3-2: After the community is divided, the community association feature extraction algorithm GALE (Community Association Likelihood Estimation) is constructed; the community association score of the node pair is extracted by the GALE algorithm, and the concept of the strong influence area of ​​the node is first proposed. , will be connected to the nodes in the network A node cluster with a distance less than or equal to 3 is defined as , indicating a node The strong influence area of ​​​​the node and the node The relationship is a strong connection; the specific definition is as follows: ; Step 3-3, then connect the network to the node The set of nodes whose distance is greater than 3 and less than 6 is defined as , i.e. node The weakly affected area of ​​​​the node in this area and the node There are weak connections between them; in this way, the distribution and scope of influence in the network can be divided more finely, enriching the complexity and depth of network analysis; the specific definition is as follows: ; For node pairs , define the node With Node The intersection of the strong influence areas is the common strong influence area , the intersection of weakly affected areas is the common weakly affected area ; The specific definition is as follows: ; ; Step 3-4, finally calculate the community association score based on the node distribution of common strong influence domains and weak influence domains: ; ; 。 5. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 4, characterized in that: In step 3-1, the core of the algorithm is that during this diffusion process, the total entropy of the network decreases, showing high stability; this entropy reduction is more significant in the densely connected areas that form communities; the core concept of the SURP algorithm is the measure of "surprise", which is used to evaluate the unexpectedness or unusualness of community divisions in the network; the algorithm is based on statistical significance and evaluates the difference between the observed community divisions and random divisions in the network; the core theory of the SIGNI algorithm is the statistical significance test, which attempts to identify node groups whose internal link density is significantly higher than random expectations by comparing the differences between the observed network structure and the expected structure of the random network model.

6. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 4, characterized in that: In step 3-1, the core theory of the EIGEN algorithm is eigenvector centrality, which measures the importance of a node in the network according to the number of its neighbors and the importance of its neighbors; the core theory of the Girvan-Newman algorithm is to remove the edges connecting different communities and detect hidden communities or modules; Its algorithm identifies the edges connecting different communities by calculating the betweenness of the edges in the network, and reveals the hidden communities in the network based on these edges; after gradually removing these edges, the network gradually decomposes into multiple internally tightly connected modules or communities.

7. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 4, characterized in that: In step 3-1, the Fluid Communities algorithm is based on the concept of dynamic fluid model and simulates the interaction of fluid in the environment. In this algorithm, each community is regarded as a "fluid" that flows and interacts with each other in the network. This concept of flow and interaction simulates the relationship between community members in the network. Finally, based on this algorithm, a Fluid Communities community detection algorithm (Fluid Communities for Multiplex, FCM) suitable for multi-layer networks is proposed.

8. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 7, characterized in that: The FCM algorithm first divides the multi-layer network into multi-layer sub-networks according to the hierarchical structure, and uses the Fluid Communities algorithm to perform community detection on each layer of the network to obtain community detection results at different network layers. Hyperparameters are introduced into the FCM algorithm. , at least Only when there are node clusters with the same community label in the layer sub-network can the final community be formed in the multi-layer network; in the actual calculation process of the FCM algorithm, The value of is set according to the hierarchical structure of the original multi-layer network. The larger the value, the sparser the distribution between communities and the closer the connections within the community.

9. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 4, characterized in that: In step 3-2, the community association feature extraction algorithm GALE is based on the community label distribution of the community detection algorithm result: the community structure divided by the community detection algorithm (CDM) provides basic information for community association feature extraction; Weight adjustment: The influence of a node in a community is inversely proportional to the size of the community. This normalized weight adjustment ensures that the contribution of a single node in a large community will not be overestimated due to its large size. When it is larger, its contribution to the target node's community association score is adjusted according to the following rules: ; This normalization method reduces the calculation bias caused by differences in community size; Multi-dimensional evaluation of community influence: By dividing the strong influence area into weak influence areas, the GALE method not only focuses on the direct community connection between nodes, but also comprehensively considers the distribution characteristics of nodes across communities to more comprehensively evaluate the possibility of node pairs forming potential links; Adaptability to dynamic networks: The GALE method handles dynamic network changes through incremental calculation. When new nodes or edges affect the network community structure, the strong and weak influence areas and community association scores of the nodes are quickly adjusted to maintain the real-time and accuracy of the link prediction results.

10. The method for constructing a multi-layer network link prediction model taking into account multiple correlation features according to claim 9, characterized in that: In the common strong influence area and the common weak influence area, the node distribution and its community label are important factors in evaluating the node-to-community association score; If the common strong influence area of ​​the target node pair contains multiple nodes with different community labels, and the community size of these nodes is large, it indicates that the community structure of the area is relatively complex and has high heterogeneity; this complexity reduces the possibility of node pairs forming potential links in this area, resulting in a lower community association score; On the contrary, when the community labels in the common strong influence area of ​​the target node pair are more uniformly distributed or mainly concentrated in the same community, the probability of link formation is higher and the community association score is correspondingly higher.

Citation Information

Cited By

  • Open source platform task matching method based on developer behavior map and double-layer filtering

    CN121658534A