A third-party library recommendation method and device based on high-order collaborative filtering

By constructing application hypergraphs and third-party library hypergraphs, and using hypergraph neural networks to extract high-order neighborhood information, the problem of noise introduction in existing technologies is solved, and more accurate and diverse third-party library recommendations are achieved.

CN116821521BActive Publication Date: 2026-01-16GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310858893.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2026-01-16
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

Existing collaborative filtering algorithms based on graph convolutional neural networks rely on graph structures when recommending third-party libraries, which introduces noise and results in insufficient accuracy and diversity of recommendation results.

Method used

We construct application hypergraphs and third-party library hypergraphs, extract high-order neighborhood information through hypergraph neural networks, generate node vectors containing high-order neighborhood information, and recommend third-party libraries by calculating probability scores through inner product.

Benefits of technology

This improves the accuracy and diversity of recommendation results and generates more precise node embedding vector representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821521B_ABST
    Figure CN116821521B_ABST
Patent Text Reader

Abstract

The application discloses a kind of third-party library recommendation method and device based on high-order collaborative filtering, method includes: obtaining the interaction information between target application and third-party library, similarity is calculated, and then determine application hypergraph and third-party library hypergraph;Application node and library node are embedded in latent semantic space, and the latent semantic vector of application and third-party library is obtained;High-order neighborhood information of hypergraph is extracted by hypergraph neural network, and node vector is generated;The new node latent semantic vector is obtained by weighting aggregation to node vector and initial latent semantic vector;The inner product of the latent semantic vector of application and third-party library is calculated, and the possibility score is obtained, and then third-party library is recommended for application;The high-order neighborhood information is extracted by hypergraph in the application, so that node can obtain the high-order neighborhood information containing less noise, to generate more accurate node embedding vector representation, produce more accurate and diverse recommendation results, can be widely applied in computer technology application field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer technology application, and particularly to a third-party library recommendation method and device based on high-order collaborative filtering. BACKGROUND

[0002] Due to the popularity of mobile smart phones, the demand for mobile applications (Apps) has increased dramatically. The number of Apps has also grown rapidly over time, with an increasing iteration speed, which puts App developers in a challenging environment, where they must quickly develop or improve their products to remain competitive. It is unnecessary to reinvent the wheel, and third-party libraries (TPLs) can provide many functions required for App development, such as location services, web requests, and push advertising. In addition, many bugs have been found and fixed in these libraries, which helps to optimize App development.

[0003] TPLs have become an important part of the Android ecosystem, but the proliferation of TPLs has also made it a major challenge to find TPLs that meet the needs of developers. Finding suitable TPLs among a large number of TPLs and evaluating their effectiveness from multiple perspectives, such as functionality, performance, and compatibility, requires a considerable amount of time. TPL recommendation can recommend TPLs that meet the development needs of software developers, reducing the time they spend searching for TPLs.

[0004] The best existing third-party library recommendation algorithm is a collaborative filtering algorithm based on graph convolutional neural networks. This algorithm models the interaction records of Apps and TPLs as an App-TPL interaction graph, then represents each App and TPL node using an embedding vector of dimension d, and then uses a graph convolutional neural network for information propagation to capture high-order neighborhood information of Apps and TPLs, thereby generating node embedding vectors that contain high-order neighborhood information. This algorithm produces recommendation results that are more accurate and diverse than traditional third-party library recommendation algorithms.

[0005] The main deficiency of the prior art is that the collaborative filtering algorithm based on graph convolutional neural networks relies on the graph structure, and the information propagation based on the graph structure introduces noise when the graph nodes obtain high-order neighborhood information, resulting in a significant oversmoothing phenomenon in the graph node vectors formed, which makes the accuracy and diversity of the recommendation results still significantly insufficient. SUMMARY

[0006] Therefore, the embodiments of the present application provide a third-party library recommendation method based on high-order collaborative filtering with high accuracy.

[0007] In one aspect, the embodiments of the present application provide a third-party library recommendation method based on high-order collaborative filtering:

[0008] obtaining interaction information between the target application and the third-party library, calculating application similarity and third-party library similarity from the interaction information;

[0009] determining an application hypergraph by comparing the application similarity with a preset first threshold value, and determining a third-party library hypergraph by comparing the third-party library similarity with a preset second threshold value;

[0010] embedding application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain application latent semantic vectors and third-party library latent semantic vectors;

[0011] iteratively extracting high-order neighborhood information of nodes of the application hypergraph and the third-party library hypergraph through a hypergraph neural network to generate node vectors containing high-order neighborhood information;

[0012] performing weighted aggregation on the node vectors containing high-order neighborhood information and initial latent semantic vectors of the nodes to obtain new node latent semantic vectors;

[0013] calculating an inner product of the application latent semantic vectors and the third-party library latent semantic vectors to obtain a possibility score, and recommending a third-party library to the application according to the possibility score.

[0014] Optionally, in the step of obtaining interaction information between the target application and the third-party library, calculating application similarity and third-party library similarity from the interaction information, the calculation formula of the application similarity and the third-party library similarity is as follows:

[0015]

[0016]

[0017] wherein, L(A i ), L(A j ) respectively represent a set of third-party libraries used by the target application A i and the target application A j , |L(A i )∩L(A j )| represents a number of third-party libraries commonly used by the target application A i and the target application A j , |L(A i )| represents a number of third-party libraries used by the target application A i , and |L(A j )| represents a number of third-party libraries used by the target application A j ; A(L i ) and A(L jrespectively represent the target application set using the third-party library L i and the third-party library L j , |A(L i )∩A(L j )| represents the number of target applications using the third-party library L i and the third-party library L j , |A(L i )| represents the number of target applications using the third-party library L i , |A(L j )| represents the number of target applications using the third-party library L j ; SimApp(A i , A j ) is the application similarity, and SimTPL(L i , L j ) is the third-party library similarity.

[0018] The application further comprises a third threshold value, and the application supergraph is determined by comparing the application similarity with the first threshold value, and the third-party library supergraph is determined by comparing the third-party library similarity with the second threshold value.

[0019] The application similarity between the first application and the second application is compared with the first threshold value, and when the application similarity is greater than or equal to the first threshold value, the first application and the second application are taken as two adjacent nodes of the application supergraph.

[0020] The third-party library similarity between the first third-party library and the second third-party library is compared with the second threshold value, and when the third-party library similarity is greater than or equal to the second threshold value, the first third-party library and the second third-party library are taken as two adjacent nodes of the third-party library supergraph.

[0021] Each application node is taken as a centroid in turn, a superedge is used to connect the centroid node and all adjacent nodes of the centroid node, and the similarity between the adjacent nodes and the centroid node is taken as the weight of the adjacent nodes on the superedge, that is, the similarity between the nodes connected by the superedge and the centroid node of the superedge is taken as the weight of the nodes on the superedge, so as to construct the application supergraph.

[0022] Each third-party library node is taken as a centroid in turn, a superedge is used to connect the centroid node and all adjacent nodes of the centroid node, and the similarity between the adjacent nodes and the centroid node is taken as the weight of the adjacent nodes on the superedge, that is, the similarity between the nodes connected by the superedge and the centroid node of the superedge is taken as the weight of the nodes on the superedge, so as to construct the third-party library supergraph.

[0023] Optionally, the embedding of the application node of the application hypergraph and the library node of the third-party library hypergraph into a latent semantic space to obtain an application latent semantic vector and a third-party library latent semantic vector comprises:

[0024] inputting the application node of the application hypergraph into an initializer to obtain the application latent semantic vector through initialization;

[0025] inputting the library node of the third-party library hypergraph into the initializer to obtain the third-party library latent semantic vector through initialization.

[0026] Optionally, the iterative extraction of high-order neighborhood information of the nodes of the application hypergraph and the third-party library hypergraph through the hypergraph neural network to generate a node vector containing high-order neighborhood information comprises:

[0027] each hyperedge aggregates information of the connected nodes according to the weight of the connected nodes on the hyperedge to generate a hyperedge vector;

[0028] After the generation of the hyperedge vector, the nodes connected to the hyperedge extract hyperedge information according to the weight of the nodes on the hyperedge to generate a node vector containing high-order neighborhood information.

[0029] Optionally, the step of generating the hyperedge vector by each hyperedge aggregating information of the connected nodes according to the weight of the connected nodes on the hyperedge, the calculation formula of the hyperedge vector is:

[0030]

[0031] wherein, is the hyperedge vector, is the number of target applications, A i is the i-th application node in the application hypergraph, e k is the k-th hyperedge in the application hypergraph, is the application latent semantic vector, H1(A i , e k ) is an element in the application hypergraph association matrix H1, if A i and A k are adjacent nodes, the value is SimApp(A i , A k ), if A i and A k are not adjacent nodes, the value is 0, representing the weight of the application node A i on the hyperedge e k .

[0032] Optionally, the possibility score is obtained by calculating the inner product of the application latent semantic vector and the third-party library latent semantic vector, and the third-party library is recommended for the application according to the possibility score.

[0033] obtaining a possibility score by calculating an inner product of the application latent semantic vector and the target node vector;

[0034] sorting the possibility scores in descending order to obtain a third-party library recommendation list;

[0035] eliminating third-party libraries used by the target application from the third-party library recommendation list, and recommending a number of third-party libraries with the highest possibility scores from the remaining third-party libraries in the third-party library recommendation list.

[0036] In another aspect, an embodiment of the present application provides a third-party library recommendation device based on high-order collaborative filtering, comprising:

[0037] A first module is configured to obtain interaction information between a target application and a third-party library, and calculate application similarity and third-party library similarity based on the interaction information;

[0038] A second module is configured to determine an application hypergraph by comparing the application similarity with a preset first threshold value, and determine a third-party library hypergraph by comparing the third-party library similarity with a preset second threshold value;

[0039] A third module is configured to embed application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain an application latent semantic vector and a third-party library latent semantic vector;

[0040] A fourth module is configured to iteratively extract high-order neighborhood information of nodes of the application hypergraph and the third-party library hypergraph by using a hypergraph neural network to generate node vectors containing the high-order neighborhood information;

[0041] A fifth module is configured to aggregate the node vectors containing the high-order neighborhood information and initial latent semantic vectors of the nodes by weighting to obtain new node latent semantic vectors;

[0042] A sixth module is configured to obtain a possibility score by calculating an inner product of the application latent semantic vector and the third-party library latent semantic vector, and recommend third-party libraries for the application based on the possibility score.

[0043] In another aspect, an embodiment of the present application further provides an electronic device comprising a processor and a memory; the memory is configured to store a program; and the processor is configured to execute the program to implement a third-party library recommendation method based on high-order collaborative filtering.

[0044] In another aspect, an embodiment of the present application further provides a computer readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement a third-party library recommendation method based on high-order collaborative filtering.

[0045] The embodiment of the application further discloses a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device can read the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing method.

[0046] The embodiment of the application has at least the following beneficial results: according to the interaction information between the application and the third-party library, the embodiment of the application constructs a hypergraph, extracts high-order neighborhood information through a hypergraph neural network, can make the node obtain high-order neighborhood information containing less noise when obtaining the high-order neighborhood information, and thus generate more accurate node embedding vector representation and more accurate and diverse recommendation results. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0048] Figure 1 is a flowchart of a third-party library recommendation method based on high-order collaborative filtering provided by the embodiment of the application;

[0049] Figure 2 is a flowchart of obtaining APP-TPL interaction information, calculating similarity and then finding neighbor nodes provided by the embodiment of the application;

[0050] Figure 3 is a schematic diagram of an application hypergraph provided by the embodiment of the application;

[0051] Figure 4 is a schematic diagram of a third-party library hypergraph provided by the embodiment of the application;

[0052] Figure 5 is a schematic diagram of a third-party library recommendation device based on high-order collaborative filtering provided by the embodiment of the application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0054] In one aspect, the embodiment of the application provides a third-party library recommendation method based on high-order collaborative filtering, which is described with reference toFigure 1 The method includes but is not limited to steps S100-S600:

[0055] S100: Obtain the interaction information between the target application and the third-party library, and calculate the application similarity and the third-party library similarity through the interaction information.

[0056] Optionally, the application (APP) is a mobile application, and the third-party library (TPL) is used to provide functions required for APP development, such as location service, web request, and push advertisement. The interaction information between the application and the third-party library is obtained, and the application similarity and the third-party library similarity are calculated through the interaction information. The calculation formula of the application similarity and the third-party library similarity is:

[0057]

[0058]

[0059] Wherein, L(A i ), L(A j ) respectively represent the set of third-party libraries used by the target application A i and the target application A j , |L(A i )∩L(A j )| represents the number of third-party libraries commonly used by the target application A i and the target application A j , |L(A i )| represents the number of third-party libraries used by the target application A i , |L(A j )| represents the number of third-party libraries used by the target application A j ; A(L i ) and A(L j ) respectively represent the set of target applications using the third-party library L i and the third-party library L j , |A(L i )∩A(L j )| represents the number of target applications using the third-party library L i and the third-party library L j , |A(L i )| represents the number of target applications using the third-party library L i , |A(L j )| represents the number of target applications using the third-party library L j ; SimApp(A i , A j ) is the application similarity, SimTPL(L i , L j() represents the similarity of third-party libraries.

[0060] Optionally, one embodiment of the present invention uses 6 APPs and 7 TPLs for interaction information as an example, the interaction information being as follows: Figure 1 As shown in (a), A k For the target application (APP), L k For third-party libraries (TPLs), a number 1 indicates that the app has a user-license relationship with the TPL, while 0 indicates that the app has no direct connection with the TPL. The similarity between the app and the third-party library is calculated, and the resulting table is shown below. Figure 2 As shown in (b).

[0061] S200: By comparing the application similarity with a preset first threshold, an application hypergraph is determined; by comparing the third-party library similarity with a preset second threshold, a third-party library hypergraph is determined.

[0062] Optionally, step S200 includes, but is not limited to, steps S210-240:

[0063] S210: Compare the application similarity between the first application and the second application with a preset first threshold. If the application similarity is greater than or equal to the first threshold, then the first application and the second application are regarded as two neighboring nodes of the application hypergraph.

[0064] Optionally, a first threshold α1 is set to measure whether two apps are similar. If the app similarity between the first app and the second app is greater than or equal to the first threshold, the two apps are considered similar; if the app similarity between the first app and the second app is less than the first threshold, the two apps are considered dissimilar. If the nodes of the two apps are similar, then the two app nodes are considered as two neighbor nodes in the application hypergraph, App A i The neighbor set is denoted as N(A) i ).

[0065] Optionally, in one embodiment of the present invention, the first threshold is set to 60%, such as... Figure 2 (c) with A k The relevant table is the neighbor set of the APP. The nodes marked with "√" in the figure are application nodes with application similarity greater than or equal to the first threshold. Two nodes are neighbor nodes.

[0066] S220: Compare the similarity between the first third-party library and the second third-party library with a preset second threshold. If the similarity is greater than or equal to the second threshold, then the first third-party library and the second third-party library are regarded as two neighboring nodes of the third-party library hypergraph. For example, N(A2) is A1 and A3, that is, the neighboring nodes of A2 are A1 and A3.

[0067] Optionally, a second threshold α2 is set to measure whether two TPLs are similar. If the similarity between the first and second third-party libraries is greater than or equal to the second threshold, the two TPLs are considered similar; if the similarity is less than the second threshold, the two TPLs are considered dissimilar. If two TPL nodes are similar, they are considered as two neighbor nodes in the third-party library hypergraph. i The neighbor set is denoted as N(L). j ).

[0068] Optionally, in one embodiment of the present invention, the second threshold is set to 60%, such as... Figure 2 (c) with L k The relevant table is the set of neighbor nodes of TPL. In the figure, the nodes marked with "√" are library nodes whose similarity to the third-party library is greater than or equal to the second threshold. The two nodes are neighbors of each other.

[0069] S230: Using each application node as its centroid, connect the centroid node with all its neighboring nodes using a hyperedge, and use the similarity between the neighboring node and the centroid node as the weight of the neighboring node on this hyperedge. In other words, use the similarity between the node connected by the hyperedge and the centroid node of the hyperedge as its weight on the hyperedge, thereby constructing the application hypergraph.

[0070] Optionally, through step S210, each application node A in the application hypergraph can be obtained. j The neighbor node set N(A) j ), sequentially using each application node A j Using a hyperedge e as the centroid j Connect the centroid node to all its neighboring nodes N(A) j We will use A j The application hyperedge constructed around the centroid is named e. j and neighbor node A i With centroid node A j Similarity SimApp(A) i A j As a neighboring node A i On this superedge e j The weights on the graph are used to construct the application hypergraph. (Refer to...) Figure 3 This is a schematic diagram of an application hypergraph according to an embodiment of the present invention. The hypergraph is composed of... Figure 2 The table (c) is constructed from.

[0071] S240: Using each third-party library node as its centroid, connect the centroid node with all its neighboring nodes using a hyperedge, and use the similarity between the neighboring node and the centroid node as the weight of the neighboring node on this hyperedge. In other words, use the similarity between the node connected by the hyperedge and the centroid node of the hyperedge as its weight on the hyperedge, thereby constructing the third-party library hypergraph.

[0072] Optionally, each third-party library node L in the third-party library hypergraph can be obtained through step S220. j The neighbor node set N(L) j ), sequentially using each third-party library node L j Using a hyperedge e as the centroid j Connect the centroid node to all its neighboring nodes N(L) j We will use L j The superedge of the third-party library built for the centroid is named e. j And the similarity between neighboring nodes and centroid nodes is SimTPL(L) i L j ) as a neighbor node L i On this superedge e j The weights on the graph are used to construct a third-party library hypergraph, referring to... Figure 4 This is a schematic diagram of a third-party library hypergraph according to an embodiment of the present invention. The hypergraph is composed of... Figure 2 The table (c) is constructed from.

[0073] Optionally, in one embodiment of the present invention, a hypergraph is defined as... in It is the set of nodes in a hypergraph. It is the hyperedge set of the hypergraph; a constructed hypergraph can also use shapes of The correlation matrix H represents a hypergraph correlation matrix that differs from the traditional hypergraph correlation matrix. It assigns a weight to each node of the hyperedge. In this invention, the similarity between a node and the centroid node of the hyperedge is used as its weight on the hyperedge.

[0074] Alternatively, one embodiment of the present invention represents the application of a hypergraph as follows: It can also be represented by applying the hypergraph incidence matrix H1, with the expression:

[0075]

[0076] Among them, A i ∈V1 is The i-th App node, yes The j-th superedge, which connects A i and and included in N(A) j A in ) iSimApp(A) i A j ) represents A i With A j The application similarity between them is also the similarity between node A and node A. i In the super-edge e j The weights on the graph, where V1 is the set of application nodes in the application hypergraph. For applying the hyperedge set of the hypergraph.

[0077] Optionally, in one embodiment of the present invention, the third-party library hypergraph is represented as follows: It can also be represented as the hypergraph association matrix H2 of a third-party library:

[0078]

[0079] in, yes The i-th TPL node in yes The j-th superedge, which connects L j and and included in N(L) j L in ) i SimTPL(L) i L j ) represents L i and L j The similarity between third-party libraries is also the similarity between nodes L. i In the super-edge e j The weights on, For the third-party library hypergraph, the set of third-party library nodes. This is the set of hyperedges for a hypergraph from a third-party library.

[0080] S300: Embed the application nodes of the application hypergraph and the library nodes of the third-party library hypergraph into the latent semantic space to obtain the application latent semantic vector and the third-party library latent semantic vector.

[0081] Optionally, step S300 includes, but is not limited to, steps S310-S320:

[0082] S310: Input the application nodes of the application hypergraph into the initializer, and obtain the application latent semantic vector through initialization.

[0083] Optionally, in one embodiment of the present invention, the initializer is a Xavier initializer. The application nodes of the application hypergraph are input into the Xavier initializer, and the nodes are initialized into latent semantic vectors of dimension d. For example, application node Ai is initialized into an application latent semantic vector.

[0084] S320: Input the library nodes of the third-party library hypergraph into the initializer, and obtain the implicit semantic vector of the third-party library through initialization.

[0085] Optionally, the library nodes of the third-party library hypergraph are input into the Xavier initializer. The nodes are initialized by the initializer into latent semantic vectors of dimension d. For example, library node L i It is initialized as a third-party library implicit semantic vector

[0086] Optionally, latent semantic vectors can be continuously learned through model training, aiming to represent the features of an App or TPL, such as functionality, performance, security features, data storage methods, versatility, and compatibility. App A i Application of latent semantic vectors With TPLL j Third-party library implicit semantic vectors inner product It can be used to represent App A i For TPL L j Potential for use The larger the value, the more it indicates that A i For L i The higher the potential for use, the better.

[0087] S400: Iteratively extract the high-order neighborhood information of the nodes of the application hypergraph and the third-party library hypergraph through a hypergraph neural network, and generate a node vector containing the high-order neighborhood information.

[0088] Optionally, step S400 includes, but is not limited to, S410-S430:

[0089] S410: Extract higher-order neighborhood information of nodes and generate node vectors with higher-order neighborhood information.

[0090] Optionally, the higher-order neighborhood information of nodes in the application hypergraph and the hypergraph of the third-party library is extracted by iteratively executing the following two processes: (1) Each hyperedge aggregates the information of the connected nodes according to the weight of the connected nodes on the hyperedge to generate a hyperedge vector; (2) After generating the hyperedge vector, the nodes connected to the hyperedge extract the hyperedge information according to their weight on the hyperedge to generate a node vector with higher-order neighborhood information.

[0091] S420: Each hyperedge generates a hyperedge vector by aggregating the information of the connected nodes according to the weights of the nodes on the hyperedge.

[0092] Optionally, each hyperedge of the hypergraph generates a hyperedge vector by aggregating information of the connected nodes according to the weight of the connected nodes on the hyperedge. The hyperedge vector of the hyperedge e2 is generated in one embodiment of the application For example, the hyperedge e2 connects A2 and N(A2), i.e. A1 and A3, so the hyperedge vector of e2 is It can be obtained by the following expression:

[0093]

[0094] wherein When i = 4, 5, 6, H1(A i , e2) are all 0, so It can be expressed as:

[0095]

[0096] The above expression is extended to the entire application hypergraph, and any hyperedge vector is:

[0097]

[0098] wherein, wherein, is the hyperedge vector, is the number of nodes in the application hypergraph, A i is the i-th node in the application hypergraph, is the application latent semantic vector, H1(A i , e k ) is an element of the application hypergraph adjacency matrix, if A i and A k are adjacent nodes, its value is SimApp(A i , A k ), if A i and A k are not adjacent nodes, its value is 0, representing the weight of the application node A i on the hyperedge e k .

[0099] Optionally, the application node set vector represents The process of obtaining the hyperedge set vector representation of the application hypergraph can be expressed as:

[0100] E = H T X

[0101] wherein, E is the application hypergraph hyperedge set vector, H is the application hypergraph adjacency matrix, and X is the application node set vector.

[0102] Optionally, this process is extended to the multi-layer network layer of the hypergraph neural network, and the process of applying the hyperedge set vector generated in the 1th network layer can be represented as:

[0103] E l T X (l)

[0104] where X (l) is the application node set vector of the 1th network layer, and E l is the hyperedge set vector of the 1th network layer.

[0105] S430: After generating the hyperedge vector, the nodes connected by the hyperedge extract the hyperedge information according to their weights on the hyperedge to generate the node vector with high-order neighborhood information.

[0106] Optionally, the nodes connected by the hyperedge extract the hyperedge information according to their weights on the hyperedge, and the process can be represented as:

[0107] X( l+1 )=HE l

[0108] where X (l+1) is the application node set vector of the (l+1)th network layer, i.e. the node vector with high-order neighborhood information, which is also the target application node set vector formed after the 1th network layer extracts the hyperedge set vector.

[0109] Optionally, the entire process of extracting high-order neighborhood information can be summarized as:

[0110] X (l+1) =HH T X (l)

[0111] Optionally, the extracted hyperedge information and node information are normalized to obtain the standard expression for extracting high-order neighborhood information

[0112]

[0113] where D v and D e are the node degree matrix and hyperedge degree matrix of the application hypergraph, respectively, and play the role of normalization, and the degree of the node of the application hypergraph or the third-party library hypergraph and the degree of the hyperedge are defined as

[0114]

[0115] ​

[0116] where H(·) and v i may be H1(·) and A i or H2(·) and L i , d(v i ) is the degree of the application node or the degree of the third-party library node, and d(e i ) is the degree of the hyperedge in the application hypergraph or the degree of the hyperedge in the third-party library hypergraph.

[0117] S500: Weighted aggregation is performed on the node vector containing high-order neighborhood information and the initial latent semantic vector of the node to obtain a new node latent semantic vector.

[0118] Optionally, the calculation formula of the step of weighted aggregation of the node vector containing high-order neighborhood information and the initial latent semantic vector of the node to obtain a new node latent semantic vector is as follows:

[0119]

[0120] where X is a new node set latent semantic vector, m represents the number of network layers in the hypergraph neural network, X (l) represents the node set vector representation of the lth network layer, X (0) is the initial latent semantic vector of the node set, a l is the weight of the node set vector representation of the lth network layer, and a l ≥ 0, when we set the node set vector representation of each network layer to have equal weights, a l may take a value of

[0121] S600: A possibility score is obtained by calculating the inner product of the application latent semantic vector and the third-party library latent semantic vector, and a suitable third-party library is recommended for the application according to the possibility score.

[0122] Optionally, step S600 includes but is not limited to S610-S630:

[0123] S610: A possibility score is obtained by calculating the inner product of the application latent semantic vector and the third-party library latent semantic vector.

[0124] Optionally, the possibility score is obtained by calculating the inner product of the application latent semantic vector and the third-party library latent semantic vector

[0125] S620: The possibility score is sorted in descending order to obtain a third-party library recommendation list.

[0126] Optionally, the possibility score is sorted in descending order to obtain a third-party library recommendation list.

[0127] S630: Discard the third-party library used by the application in the third-party library recommendation list, and recommend a plurality of third-party libraries with the highest possibility score of the remaining third-party libraries in the third-party library recommendation list.

[0128] Optionally, the third-party library used by the application in the third-party library recommendation list is discarded, and a plurality of third-party libraries with the highest possibility score of the remaining third-party libraries in the third-party library recommendation list are recommended, and the recommendation is completed.

[0129] The execution and application of a third-party library recommendation method based on high-order collaborative filtering will be described below through an embodiment of the application:

[0130] 1. Obtain the interaction information between the application and the third-party library, calculate the application similarity and the third-party library similarity from the interaction information, compare the application similarity with a preset first threshold value to determine the application hypergraph, and compare the third-party library similarity with a preset second threshold value to determine the third-party library hypergraph;

[0131] 2. Embed the application nodes of the application hypergraph and the library nodes of the third-party library hypergraph into the latent semantic space to obtain the application latent semantic vector and the third-party library latent semantic vector; and iteratively extract the high-order neighborhood information of the nodes of the application hypergraph and the third-party library hypergraph through the hypergraph neural network to generate node vectors containing high-order neighborhood information.

[0132] 3. Weight and aggregate the node vectors containing high-order neighborhood information and the initial latent semantic vectors of the nodes to obtain new node latent semantic vectors; calculate the inner product of the application latent semantic vector and the third-party library latent semantic vector to obtain a possibility score, and recommend appropriate third-party libraries for the application according to the possibility score.

[0133] The third-party library recommendation method based on high-order collaborative filtering has the following beneficial effects: according to the interaction information between the application and the third-party library, the hypergraph is constructed, and the high-order neighborhood information is extracted through the hypergraph neural network, so that the node can obtain high-order neighborhood information containing less noise when obtaining the high-order neighborhood information, thereby generating more accurate node embedding vector representation and producing more accurate and diverse recommendation results.

[0134] On the other hand, with reference to Figure 5 , the embodiment of the application provides a third-party library recommendation device based on high-order collaborative filtering, which comprises:

[0135] The first module 501 is configured to obtain the interaction information between the application and the third-party library, and calculate the application similarity and the third-party library similarity from the interaction information.

[0136] The second module 502 is configured to determine an application hypergraph by comparing the application similarity with a preset first threshold value, and determine a third-party library hypergraph by comparing the third-party library similarity with a preset second threshold value.

[0137] The third module 503 is configured to embed application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain application latent semantic vectors and third-party library latent semantic vectors.

[0138] The fourth module 504 is configured to iteratively extract high-order neighborhood information of nodes of the application hypergraph and the third-party library hypergraph by a hypergraph neural network, and generate node vectors containing the high-order neighborhood information.

[0139] The fifth module 505 is configured to perform weighted aggregation on the node vectors containing the high-order neighborhood information and initial latent semantic vectors of the nodes to obtain new node latent semantic vectors.

[0140] The sixth module 506 is configured to obtain a possibility score by calculating an inner product of the application latent semantic vectors and the third-party library latent semantic vectors, and recommend a suitable third-party library for the application according to the possibility score.

[0141] In another aspect, an electronic device is provided, which includes a processor and a memory. The memory is configured to store a program. The processor is configured to execute the program to implement the foregoing method for recommending a third-party library based on high-order collaborative filtering.

[0142] In another aspect, a computer readable storage medium is provided, which stores a program. The program is executed by a processor to implement the foregoing method for recommending a third-party library based on high-order collaborative filtering.

[0143] The embodiments of the present application also disclose a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device can read the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the method shown in the embodiments of the present application. Figure 1 The method shown in the embodiments of the present application.

[0144] In alternative embodiments, the functions / operations in the flow diagrams can occur in different orders and / or concurrently with each other. For example, two operations shown in succession can in fact be executed substantially concurrently or the operations can sometimes be executed in the reverse order, depending upon the functionality / operations involved. Furthermore, embodiments of the present application are presented in terms of exemplary implementations, which are provided as examples. Other implementations are possible and contemplated by the inventors; therefore, the following methods should not be interpreted as being limited to the particular implementation illustrated herein. Alternative implementations might be constructed by those skilled in the art, which will depend on the particular application or applications addressed and designed to operate depending upon the input received. With reference to the appended drawings is where an exemplary implementation is depicted. It should be understood that the illustrated implementation is shown as an example only and not limiting of the present application. Many implementation variations are possible, which will depend on the particular application or applications addressed. It should also be understood that the exemplary implementation is not necessarily indicative of all of the possible desirable implementations.

[0145] Furthermore, although the present application is described in the context of functional modules, it is understood that one or more of the functions and / or features described can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also understood that detailed discussion of the actual implementation of each module is unnecessary to an understanding of the present application. Rather, the actual implementation is within the routine skill of engineers familiar with the attributes, functions, and internal relationships of the various functional modules disclosed herein. Accordingly, details concerning the actual implementation are not set forth herein other than understanding that such details are within the scope of one of ordinary skill in the art and that the claimed application is not limited to the details of the implementation. It is also understood that the disclosed specific concepts are merely illustrative and are not intended to limit the scope of the present application, which is defined by the appended claims and their equivalents.

[0146] If the functions are implemented in software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a USB flash disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0147] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be embodied in non-transitory computer-readable media, executed by one or more computing devices, and / or in any other way. The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a form that can be later executed by a computer. In this context, a "computer-readable medium" can be any means that can store the program for use by or in connection with the instruction execution system, apparatus, or device.

[0148] The foregoing description of various embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed, and various modifications and variations are possible in light of the above teachings or can be acquired from practice of the application. For example, while a particular feature of the application can have been described with respect to only one or more embodiments thereof, the feature is not necessarily limited to that one or more embodiments. Rather, applicants have provided various embodiments of the application and combinations thereof and candidates can combine them in various combinations to produce yet other embodiments of the application. It is intended that the specification and examples be considered as exemplary only, with a true scope of the application being indicated by the following claims.

[0149] It is understood that various portions of the application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following technologies, known in the art, or a combination thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions upon an application of data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0150] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0151] While the embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary and that changes in form and detail can be made without departing from the principles and spirit of the application. The scope of the application is defined by the claims and their equivalents.

[0152] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the described embodiment, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

Claims

1. A third-party library recommendation method based on high-order collaborative filtering, characterized in that, The application comprises the following steps: obtaining interaction information between a target application and a third-party library, calculating application similarity and third-party library similarity from the interaction information; comparing the application similarity with a preset first threshold to determine an application hypergraph, and comparing the third-party library similarity with a preset second threshold to determine a third-party library hypergraph; embedding application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain application latent semantic vectors and third-party library latent semantic vectors; iteratively extracting high-order neighborhood information of nodes of the application hypergraph and the third-party library hypergraph through a hypergraph neural network to generate node vectors containing high-order neighborhood information; performing weighted aggregation on the node vectors containing high-order neighborhood information and initial latent semantic vectors of the nodes to obtain new node latent semantic vectors; calculating the inner product of the application latent semantic vectors and the third-party library latent semantic vectors to obtain a possibility score, and recommending a third-party library to the application according to the possibility score; In the step of obtaining interaction information between a target application and a third-party library, and calculating application similarity and third-party library similarity from the interaction information, the calculation formula of the application similarity and the third-party library similarity is as follows: wherein, , respectively represent the target application and the target application using the set of third-party libraries, represent the target application and the target application using the number of third-party libraries in common, represent the target application using the number of third-party libraries, represent the target application using the number of third-party libraries; and respectively represent the set of target applications using the third-party library and the third-party library , represent the number of target applications using the third-party library and the third-party library , represent the number of target applications using the third-party library , represent the number of target applications using the third-party library ; is the application similarity, is the third-party library similarity; The step of comparing the application similarity with a preset first threshold to determine an application hypergraph, and comparing the third-party library similarity with a preset second threshold to determine a third-party library hypergraph comprises the following steps: comparing the application similarity between a first application and a second application with a preset first threshold, and regarding the first application and the second application as two neighbor nodes of an application hypergraph when the application similarity is greater than or equal to the first threshold; comparing the third-party library similarity between a first third-party library and a second third-party library with a preset second threshold, and regarding the first third-party library and the second third-party library as two neighbor nodes of a third-party library hypergraph when the third-party library similarity is greater than or equal to the second threshold; sequentially taking each application node as a centroid, connecting the centroid node and all neighbor nodes of the centroid node with a hyperedge, and regarding the similarity between the neighbor nodes and the centroid node as the weight of the neighbor nodes on the hyperedge, i.e. regarding the similarity between the nodes connected by the hyperedge and the centroid node of the hyperedge as the weight of the nodes on the hyperedge, thereby constructing an application hypergraph; sequentially taking each third-party library node as a centroid, connecting the centroid node and all neighbor nodes of the centroid node with a hyperedge, and regarding the similarity between the neighbor nodes and the centroid node as the weight of the neighbor nodes on the hyperedge, i.e. regarding the similarity between the nodes connected by the hyperedge and the centroid node of the hyperedge as the weight of the nodes on the hyperedge, thereby constructing a third-party library hypergraph.

2. The third-party library recommendation method based on high-order collaborative filtering according to claim 1, characterized in that, The step of embedding application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain application latent semantic vectors and third-party library latent semantic vectors comprises the following steps: inputting the application nodes of the application hypergraph into an initializer to obtain application latent semantic vectors through initialization; inputting the library nodes of the third-party library hypergraph into an initializer to obtain third-party library latent semantic vectors through initialization. 3.The third-party library recommendation method based on high-order collaborative filtering according to claim 1, wherein, The high-order neighborhood information of the nodes of the application hypergraph and the third-party library hypergraph is iteratively extracted by the hypergraph neural network to generate node vectors containing high-order neighborhood information, including: Each hyperedge aggregates information of the connected nodes according to the weight of the connected nodes on the hyperedge to generate a hyperedge vector; After generating the hyperedge vector, the connected nodes of the hyperedge extract hyperedge information according to the weight of the connected nodes on the hyperedge to generate node vectors with high-order neighborhood information.

4. The method of claim 3, wherein the method further comprises: The calculation formula for generating the hyperedge vector is: in, For hyperedge vectors, For the number of target applications, For the i-th application node in the application hypergraph, For the application of hypergraphs A super edge, To apply latent semantic vectors, To apply the hypergraph incidence matrix One of the elements, if and A node that is a neighbor to another node has a value of ,like and If nodes are not neighbors, its value is 0, representing the application node. exist The weight on this superedge.

5. The method of claim 1, wherein the method is based on high-order collaborative filtering. The inner product of the application latent semantic vector and the third-party library latent semantic vector is calculated to obtain a possibility score, and the third-party library is recommended to the application according to the possibility score, including: The inner product of the application latent semantic vector and the third-party library latent semantic vector is calculated to obtain a possibility score; The possibility scores are sorted in descending order to obtain a third-party library recommendation list; The third-party libraries that have been used by the target application are removed from the third-party library recommendation list, and the third-party libraries with the highest possibility scores in the remaining third-party library recommendation list are recommended.

6. An apparatus for implementing the method of recommending third-party libraries based on high-order collaborative filtering according to any one of claims 1-5, characterized in that, Including: The first module is configured to obtain interaction information between a target application and a third-party library, and calculate an application similarity and a third-party library similarity based on the interaction information; The second module is configured to determine an application hypergraph by comparing the application similarity with a preset first threshold value, and determine a third-party library hypergraph by comparing the third-party library similarity with a preset second threshold value; The third module is configured to embed application nodes of the application hypergraph and library nodes of the third-party library hypergraph into a latent semantic space to obtain an application latent semantic vector and a third-party library latent semantic vector; The fourth module is configured to iteratively extract high-order neighborhood information of nodes of the application hypergraph and the third-party library hypergraph by a hypergraph neural network to generate node vectors containing high-order neighborhood information; The fifth module is configured to weight and aggregate the node vectors containing high-order neighborhood information and initial latent semantic vectors of the nodes to obtain new node latent semantic vectors; The sixth module is configured to calculate the inner product of the application latent semantic vector and the third-party library latent semantic vector to obtain a possibility score, and recommend a third-party library to the application according to the possibility score.

7. An electronic device, comprising: A processor and a memory are included; The memory is configured to store a program; The processor executes the program to implement the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by a processor to implement the method of any one of claims 1 to 5. The storage medium stores a program, and the program is executed by a processor to implement the method of any one of claims 1 to 5.