Software ecological community discovery method and system based on semantic embedding
By selecting heterogeneous nodes, designing metapaths, building heterogeneous information networks in heterogeneous information networks, and using the Skip-Gram model and Word2Vec method, heterogeneous software ecological community was discovered, which solved the problem that existing technology was difficult to detect communities in heterogeneous information networks, and achieved effective community division of heterogeneous networks.
Patent Information
- Application Number
- CN202111216249.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The prior art is difficult to detect communities in heterogeneous information networks and cannot effectively utilize semantic information in heterogeneous networks.
By selecting heterogeneous nodes, designing metapaths, building heterogeneous information networks, and using the Skip-Gram model and Word2Vec method, the occurrence probability and embedded vector expression of nodes are obtained, and the heterogeneous software ecological community is discovered.
It realizes the effective division of software ecological communities in heterogeneous networks, considers the structural information and semantic information of heterogeneous networks, conforms to the heterogeneity of communities in real life, and has practical significance.
Smart Images

Figure CN113987084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining and analysis, and in particular to a method and system for discovering software ecological communities based on semantic embedding. Background Art
[0002] In real-world applications, data can be more intuitively represented as network structures, such as social relationship networks, cooperative relationship networks, etc. Clustering and community detection is a basic unsupervised data analysis task in data mining. A community usually refers to a group of densely connected nodes that are more sparsely connected to the rest of the network. Detecting communities in large networks can better understand their internal structures. Community detection helps analyze the structural correlations of different entities in large-scale networks and recommend members of social groups with the same interests. Studying the evolution of communities in software ecosystems and the development trends of their technologies can guide the process of software development and regulate the allocation of personnel, etc., which is of practical significance.
[0003] In recent years, a large number of community discovery algorithms have been proposed, most of which simply make full use of network topology information, including hierarchical clustering, modularity optimization, spectral clustering and other methods. In addition to using network topology, node semantics or node attributes can also be used to improve the quality of generated communities. However, most of the existing work is based on homogeneous information networks, while real-world information networks are generally heterogeneous. Compared with homogeneous networks, heterogeneous information networks store rich semantic information and have received widespread attention in search, clustering, data mining and other fields in recent years.
[0004] A patent document with publication number CN104866781A discloses a privacy protection method for publishing social network data for community detection applications. After initializing the data and performing community detection, the nodes in the network graph are arranged; the data is anonymized through K to form super nodes, and the edges are generalized into weighted super edges; the super nodes and super edges are split; and the anonymized social network data is published.
[0005] A method for detecting community in a smart community big data knowledge graph network is disclosed in a patent document with publication number CN112925920A, which includes the following steps: S1) constructing a smart community big data knowledge graph; S2) reconstructing the smart community big data knowledge graph to obtain a low-dimensional encoding of the smart community knowledge graph network; S3) using a deep learning method to construct a deep model for training, and finally obtaining a community detection result.
[0006] A patent document with publication number CN111696345A discloses an intelligent algorithm for fast prediction of width learning of coupled large-scale data streams based on network community detection and GCN. The algorithm includes the following steps: Step 1: community detection; Step 2: spatiotemporal feature extraction; Step 3: fast prediction of width learning; Step 4: large-scale real-time prediction of spatiotemporal coupled width learning neural network.
[0007] With respect to the above-mentioned related technologies, the inventor believes that the above-mentioned solutions do not detect communities in heterogeneous information networks. Therefore, it is necessary to propose a technical solution to improve the above-mentioned technical problems. Summary of the invention
[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for discovering software ecological communities based on semantic embedding.
[0009] According to a method for discovering software ecological communities based on semantic embedding provided by the present invention, the method comprises the following steps:
[0010] Step S1: Select heterogeneous nodes, design meta-paths, and build a software ecological heterogeneous information network;
[0011] Step S2: According to the meta-path and the constructed heterogeneous information network, the Skip-Gram model is used to obtain the occurrence probability of the node, and the node with a large occurrence probability is set as the seed node;
[0012] Step S3: Taking the seed node as the center, similar nodes are obtained as members of the seed node, and finally a heterogeneous software community is obtained.
[0013] Preferably, the step S1 selects warehouses, organizations and members as heterogeneous nodes on the GitHub open source software development platform; designs a meta-path as member-organization-warehouse; and constructs a heterogeneous information network according to the meta-path.
[0014] Preferably, the semantic information associations between heterogeneous nodes and meta-path-guided nodes are used as edges to construct a software ecological heterogeneous information network.
[0015] Preferably, the input in step S2 is a meta-path, and the output is the occurrence probability of heterogeneous nodes; the meta-path is regarded as hidden semantic information; the Skip-Gram model of the Word2Vec method is used to obtain the embedded vector expression of the heterogeneous nodes through training, and the occurrence probability of all heterogeneous nodes is obtained through the mapping of the softmax layer, and the nodes with a large occurrence probability are set as seed nodes.
[0016] Preferably, the step S3 takes the meta-path and the heterogeneous information network as input; uses the Word2Vec method to find similar nodes around the seed node; uses the similar nodes as community members of the seed node, and finally obtains a heterogeneous community.
[0017] The present invention also provides a software ecological community discovery system based on semantic embedding, the system comprising the following modules:
[0018] Module M1: Select heterogeneous nodes, design meta-paths, and build a software ecosystem heterogeneous information network;
[0019] Module M2: Based on the meta-path and the constructed heterogeneous information network, the Skip-Gram model is used to obtain the probability of node occurrence, and nodes with high probability of occurrence are set as seed nodes;
[0020] Module M3: Taking the seed node as the center, similar nodes are obtained as members of the seed node, and finally a heterogeneous software community is obtained.
[0021] Preferably, the module M1 selects warehouses, organizations and members as heterogeneous nodes on the GitHub open source software development platform; designs a meta-path as member-organization-warehouse; and constructs a heterogeneous information network according to the meta-path.
[0022] Preferably, the semantic information associations between heterogeneous nodes and meta-path-guided nodes are used as edges to construct a software ecological heterogeneous information network.
[0023] Preferably, the input in the module M2 is a meta-path, and the output is the occurrence probability of heterogeneous nodes; the meta-path is regarded as hidden semantic information; the Skip-Gram model of the Word2Vec system is used to obtain the embedded vector expression of the heterogeneous nodes through training, and the occurrence probability of all heterogeneous nodes is obtained through the mapping of the softmax layer, and the nodes with a large occurrence probability are set as seed nodes.
[0024] Preferably, the module M3 takes the meta-path and the heterogeneous information network as input; uses the Word2Vec system to find similar nodes around the seed node; uses the similar nodes as community members of the seed node, and finally obtains a heterogeneous community.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. The present invention selects heterogeneous nodes on the GitHub open source software development platform, and obtains the correlation between heterogeneous nodes by defining meta-paths as hidden semantic information, constructs a heterogeneous information network, and uses the Word2Vec method based on this network to select seed nodes and perform community discovery around the seed nodes, and finally obtains a heterogeneous software ecological community;
[0027] 2. The present invention utilizes meta-paths to discover heterogeneous communities, which not only considers the structural information of heterogeneous networks, but also the semantic information;
[0028] 3. The present invention can divide software ecological communities in heterogeneous networks, which is in line with the heterogeneity of communities in real life and has more practical significance;
[0029] 4. By dividing the heterogeneous communities in the software ecosystem, we can better understand the development trend of software technology and lay the foundation for future work in areas such as software development, technical personnel allocation and recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0031] Figure 1 The schematic diagram of the heterogeneous information network constructed by selecting heterogeneous nodes, defining meta-paths, and constructing the heterogeneous information network on the GitHub open source software development platform of the present invention;
[0032] Figure 2 The figure is a schematic diagram showing an example of the process of the software ecological community discovery method based on semantic embedding of the present invention. DETAILED DESCRIPTION
[0033] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0034] In this paper, we mainly detect communities in heterogeneous information networks. First, we obtain data on GitHub, select heterogeneous nodes and define specific meta-paths to build heterogeneous information networks. Secondly, we use the Skip-Gram model to find seed nodes in heterogeneous information networks. Finally, with the seed nodes as the center, we get the embedded expressions of different nodes according to the Word2Vec method, find similar nodes as community members, and finally detect heterogeneous communities.
[0035] like Figure 1 The heterogeneous information network of the present invention is implemented by the following technical solutions:
[0036] A heterogeneous information network G can be defined as a graph G = {V, E, φ, ψ} with a vertex type mapping function φ: V→T and an edge type mapping function ψ: E→R, where V = {v 1 , v 2 , …, vn} is the vertex set of the heterogeneous information network, the i-th vertex v i ∈V belongs to a specific vertex type T, satisfying φ(v i )∈T, where i∈[1,n]. Similarly, E={e 1 , e 2 ,…,e m} is the edge set of the network, the jth edge e j ∈E belongs to a certain edge type R, satisfying ψ(e j )∈R, where j∈[1,m]. If the node type T and edge type R satisfy |T|+|R|>2, the network is called a heterogeneous information network, otherwise it is a homogeneous information network.
[0037] The meta-path P is a directed acyclic graph defined based on a heterogeneous information network, where different types of vertices are connected by corresponding edge types. In the form of i ∈T and R i ∈ R. Meta-paths can associate heterogeneous nodes, and different meta-paths can represent different semantic information.
[0038] For example, Figure 1 (2) It vividly shows an example of a heterogeneous information network on the GitHub open source software development platform, which contains three types of vertices, including repositories, organizations, and members of organizations. Figure 1 In (3), a meta-path is listed, member-organization-warehouse, whose semantic information is that members constitute the organization and the organization creates the warehouse.
[0039] like Figure 2 , the heterogeneous software ecological community discovery of the present invention can be achieved through the following technical solutions:
[0040] The present invention considers a heterogeneous software ecological community C composed of different types of vertices, which are structurally or semantically related. The heterogeneous software ecological community can be represented as C = {v 1 , v 2 , …, v k},in And node v i and node v j Satisfy φ(v i )≠φ(v j ).
[0041] In practical applications, most data are heterogeneous, so the discovery of heterogeneous communities is more practical.
[0042] The heterogeneous software ecological community can be represented as follows: given a heterogeneous information network G = {v 1 , v2 , …, v n}, our goal is to divide the network into different communities C = {c 1 , c 2 , …, c n}, where each heterogeneous community c i The internal heterogeneous nodes are closely related. In the invention, the detected communities can be overlapping, which is more in line with the actual situation because some heterogeneous nodes are difficult to be divided into a specific community.
[0043] The specific implementation steps of the software ecological community discovery method based on semantic embedding provided by the present invention are as follows:
[0044] Step S1: Select heterogeneous nodes on the GitHub open source software development platform: warehouse, organization, organization member, design the meta-path: organization member-organization-warehouse, and use the semantic information correlation between heterogeneous nodes and nodes guided by the meta-path as edges to build a software ecosystem heterogeneous information network.
[0045] Step S2: According to the meta-path and the constructed heterogeneous information network, the Skip-Gram model of the Word2Vec method is used to obtain the occurrence probability of heterogeneous nodes, and the nodes with large occurrence probability (the top 20%) are set as seed nodes.
[0046] Heterogeneous nodes are treated as words and input in onehot encoding form, and each word forms a vector w of V*1 c , V is the number of nodes. In the subsequent hidden layer, the input sparse onehot vector is reduced to a d-dimensional space, i.e., the word embedded expression.
[0047] Use the Skip-Gram model for training, obtain the word matrix W, the dimension is d*V, and take out the word vector v of the central word c =Ww c , a vector of dimension d*1, u x Represents the word vector of the xth word in the window outside the target word, and their inner product u x T v c That is the similarity between two nodes, and the calculation results of each node are obtained in turn.
[0048] The softmax function is used to obtain the probability of other nodes appearing when all given nodes appear. The formula is as follows:
[0049]
[0050] Traverse all given nodes, count the results, and set the nodes with the top 20% probability as seed nodes as input for subsequent community discovery.
[0051] Step S3: With the seed node as the center, the embedded expression of the node is obtained through the Word2Vec method, and similar nodes are used as members of the seed node to finally obtain a heterogeneous software ecological community.
[0052] The present invention also provides a software ecological community discovery system based on semantic embedding, comprising the following modules:
[0053] Module M1: Select heterogeneous nodes, design meta-paths, and build a software ecosystem heterogeneous information network; select warehouses, organizations, and members as heterogeneous nodes on the GitHub open source software development platform; design the meta-path as member-organization-warehouse; build a heterogeneous information network based on the meta-path; use the semantic information correlation between heterogeneous nodes and nodes guided by the meta-path as edges to build a software ecosystem heterogeneous information network.
[0054] Module M2: According to the meta-path and the constructed heterogeneous information network, the Skip-Gram model is used to obtain the occurrence probability of the node, and the nodes with a high occurrence probability are set as seed nodes; the input is the meta-path, and the output is the occurrence probability of the heterogeneous node; the meta-path is regarded as hidden semantic information; the Skip-Gram model of the Word2Vec system is used to obtain the embedded vector expression of the heterogeneous nodes through training, and the occurrence probability of all heterogeneous nodes is obtained through the mapping of the softmax layer, and the nodes with a high occurrence probability are set as seed nodes.
[0055] Module M3: With the seed node as the center, similar nodes are obtained as members of the seed node, and finally a heterogeneous software community is obtained; meta-paths and heterogeneous information networks are taken as input; around the seed node as the center, the Word2Vec system is used to find similar nodes; similar nodes are used as community members of the seed node, and finally a heterogeneous community is obtained.
[0056] The present invention selects heterogeneous nodes on the GitHub open source software development platform, and obtains the correlation between heterogeneous nodes by defining meta-paths as hidden semantic information, constructs a heterogeneous information network, and uses the Word2Vec method based on this network to select seed nodes and discover communities around the seed nodes, and finally obtains a heterogeneous software ecological community; the present invention uses meta-paths to discover heterogeneous communities, which not only considers the structural information of the heterogeneous network, but also the semantic information.
[0057] The present invention can divide software ecological communities in heterogeneous networks, which conforms to the heterogeneity of communities in real life and has more practical significance. By dividing heterogeneous communities in the software ecology, we can better understand the development trend of software technology and lay the foundation for future work in the fields of software development, technical personnel allocation and recommendation.
[0058] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0059] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A method for discovering software ecological communities based on semantic embedding, It is characterized in that The method comprises the following steps: Step S1: Select heterogeneous nodes, design meta-paths, and build a software ecological heterogeneous information network; Step S2: According to the meta-path and the constructed heterogeneous information network, the Skip-Gram model is used to obtain the occurrence probability of the node, and the node with a large occurrence probability is set as the seed node; Step S3: Taking the seed node as the center, similar nodes are obtained as members of the seed node, and finally a heterogeneous software community is obtained; The step S1 selects warehouses, organizations and members as heterogeneous nodes on the GitHub open source software development platform; designs a meta-path of member-organization-warehouse; and constructs a heterogeneous information network according to the meta-path; The input in step S2 is a meta-path, and the output is the occurrence probability of a heterogeneous node; the meta-path is regarded as hidden semantic information; the Skip-Gram model of the Word2Vec method is used to obtain an embedded vector expression of the heterogeneous node through training, and the occurrence probability of all heterogeneous nodes is obtained through mapping of the softmax layer, and the node with a large occurrence probability is set as a seed node; The step S3 takes the meta-path and the heterogeneous information network as input; Centering on the seed node, the Word2Vec method is used to find similar nodes. Similar nodes are used as community members of the seed node, and finally a heterogeneous community is obtained.
2. According to the method for discovering software ecological communities based on semantic embedding according to claim 1, It is characterized in that The semantic information correlation between heterogeneous nodes and meta-path-guided nodes is used as edges to construct a software ecological heterogeneous information network.
3. A software ecological community discovery system based on semantic embedding, It is characterized in that The system includes the following modules: Module M1: Select heterogeneous nodes, design meta-paths, and build a software ecosystem heterogeneous information network; Module M2: Based on the meta-path and the constructed heterogeneous information network, the Skip-Gram model is used to obtain the probability of node occurrence, and nodes with high probability of occurrence are set as seed nodes; Module M3: Taking the seed node as the center, similar nodes are obtained as members of the seed node, and finally a heterogeneous software community is obtained; The module M1 selects warehouses, organizations and members as heterogeneous nodes on the GitHub open source software development platform; designs a meta-path of member-organization-warehouse; and constructs a heterogeneous information network based on the meta-path; The input of the module M2 is the meta-path, and the output is the occurrence probability of the heterogeneous nodes; the meta-path is regarded as hidden semantic information; the Skip-Gram model of the Word2Vec system is used to obtain the embedded vector expression of the heterogeneous nodes through training, and the occurrence probability of all heterogeneous nodes is obtained through the mapping of the softmax layer, and the nodes with the largest occurrence probability are set as seed nodes; The module M3 takes the meta-path and the heterogeneous information network as input; Centered around the seed node, the Word2Vec system is used to find similar nodes. Similar nodes are used as community members of the seed node, and finally a heterogeneous community is obtained.
4. According to the software ecological community discovery system based on semantic embedding according to claim 3, It is characterized in that The semantic information correlation between heterogeneous nodes and meta-path-guided nodes is used as edges to construct a software ecological heterogeneous information network.
Citation Information
Patent Citations
Privacy protection method for community detection application-oriented social network data publication
CN104866781A
Coupling large-scale data flow width learning rapid prediction intelligent algorithm based on network community detection and GCN
CN111696345A
Smart community big data knowledge graph network community detection method
CN112925920A
Multi-source heterogeneous data entity alignment method oriented to field of public security
CN111753024A
Domain concept expression method and system based on Word2vec and LPA
CN112487267A