A system and method for discovering enterprise collaboration groups applicable to multiple industrial chains

The invisible inherent features and invisible topological features of enterprises in multiple industrial chains are extracted through NMF technology, and combined with the multi-view fuzzy clustering algorithm DVHVC-MVFC, the shortcomings of the multi-view clustering algorithm in enterprise collaboration group discovery are solved, and more accurate enterprise collaboration group discovery and cooperation decision support are achieved.

CN116776180BActive Publication Date: 2025-10-03SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310633896.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2025-10-03
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

Existing multi-view clustering algorithms fail to effectively combine the inherent characteristics and topological features of enterprises when discovering collaborative groups of enterprises in multiple industrial chains, resulting in inaccurate clustering results.

Method used

The non-negative matrix factorization (NMF) technology is used to extract the common hidden inherent features and hidden topological features of enterprises, and collaborative learning is carried out through the multi-view fuzzy clustering algorithm. A new multi-view fuzzy clustering algorithm DVHVC-MVFC is designed to combine the inherent features and topological features of enterprises for collaborative group discovery.

Benefits of technology

It achieves more accurate discovery of enterprise collaboration groups in multiple industrial chains, improves the reliability of clustering results and the quality of the basis for enterprise cooperation decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776180B_ABST
    Figure CN116776180B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for discovering enterprise collaboration groups applicable to multiple industrial chains. The system includes an enterprise feature extraction module and a multi-view clustering module. The enterprise feature extraction module is implemented using non-negative matrix factorization technology. The multi-view clustering module includes a weighted integration of two types of visible and invisible views: inherent features and topological features, and multi-view learning constraints based on the spatial topological relationships of individuals and the subordinate pattern trends of individuals obtained through fuzzy partitioning. The present invention can simultaneously mine two types of visible and invisible view information of an enterprise, including inherent feature perspective-specific (visible) and shared (invisible) information, as well as topological feature perspective-specific (visible) and shared (invisible) information brought about by enterprise connection relationships. It also completes collaborative learning of the two types of perspectives and obtains better enterprise collaboration group discovery results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent computing, and in particular relates to an enterprise collaboration group discovery method applicable to multiple industrial chains. Background Art

[0002] Discovering collaborative enterprise groups within multi-chains involves identifying clusters of enterprises, or collaborative groups, with similar behaviors or business objectives within a large network of multi-chains, based on the connections between enterprises. A collaborative enterprise group within a multi-chain network is a tightly packed structure consisting of enterprises and the edges connecting them. Collaborative enterprise group discovery involves segmenting all enterprises based on the principle that enterprises within a group are tightly connected and enterprises between groups are sparsely connected. Uncovering potential collaborative enterprise groups within a multi-chain structure is crucial for understanding complex systems and for corporate decision-making.

[0003] The current manufacturing industry chain exhibits multiple nested features within the chain and multidimensional entanglements beyond it, involving numerous enterprises. Analyzing these features using big data analysis techniques—clustering, in particular multi-view clustering—is a current mainstream trend. In multi-view clustering, each network layer corresponds to a view, and each enterprise corresponds to a cluster sample. Multi-view collaborative clustering then fuses the diverse information from multiple networks to discover distinct cluster structures, namely, enterprise collaboration groups. Enterprises within multiple industry chains not only possess inherent characteristics such as business characteristics, geographic location, business scope, and cooperative tendencies, but also, due to development needs, form collaborative relationships with other enterprises across different industry chains, thus forming the topological characteristics of the complex network system. Both of these characteristics influence the discovery of potential collaboration groups within complex systems. Therefore, constructing a multi-view clustering algorithm that simultaneously considers these two characteristics to accurately discover enterprise collaboration groups within multiple industry chains is a practical and pressing issue.

[0004] Multi-view clustering algorithms are currently a hot topic in the field of intelligent computing, and a significant amount of related work has been proposed. However, most existing multi-view clustering algorithms focus on only one type of information: either based solely on the inherent characteristics of the individual entities or solely on the topological characteristics formed by the connections between them. This means that the modeling of clustering object characteristics is incomplete, leading to biased pattern recognition in multi-network structures. Regarding these two types of information, most existing work uses visible views (given clustering object characteristics or connectivity relationships) to perform analysis, or exploits hidden views shared by multiple visible views to achieve data fusion of multiple networks, ultimately performing clustering on only one hidden view. This approach ignores the coupling relationship between multiple views and fails to consider the influence of the visible view, resulting in poor clustering results. Some work is beginning to consider the collaborative learning of visible and hidden views for clustering, but lacks comprehensive consideration of both the inherent and topological characteristics of the clustering objects. Efficient fusion of these two types of feature views can further improve multi-view clustering performance. Enterprises in multi-industry chains possess distinct attributes of both types, making existing multi-view clustering algorithms unsuitable for directly discovering collaborative groups of enterprises within these chains. In order to more accurately explore potential enterprise collaboration groups in multiple industrial chains and improve the reliability of subsequent decision-making basis, the present invention discloses a method suitable for discovering enterprise collaboration groups in multiple industrial chains - a multi-view fuzzy clustering algorithm based on the collaboration of two types of visible and hidden views of enterprises (Multiview Fuzzy Clustering with Double Visible-Hidden View Cooperation, DVHVC-MVFC). Summary of the Invention

[0005] Technical Issues: Discovering collaborative groups of enterprises across multiple industrial chains based on multi-view clustering faces three major challenges: 1) While the multi-layer visible inherent features of an enterprise are known information, existing multi-view clustering techniques have shown that shared hidden views exist within these layers. Multi-view clustering that combines visible and hidden views offers higher performance. 2) The topological relationships of enterprises within each layer of the industrial chain are predetermined, but they must be decomposed into multi-layer visible topological features and shared hidden topological features to fully exploit the topological features derived from these relationships and thus comprehensively model enterprise characteristics. 3) Existing multi-view clustering techniques lack the fusion of inherent feature views and topological feature views, making them unsuitable for structural analysis across multiple industrial chains. In summary, discovering collaborative groups of enterprises across multiple industrial chains still requires improvement in both feature modeling and multi-view clustering.

[0006] Technical solution:

[0007] To address the challenges of current technologies, the present invention's technical solution consists of two components: an enterprise feature extraction module and a multi-view clustering module. The enterprise feature extraction module aims to obtain shared implicit intrinsic feature views between visible intrinsic feature views of an enterprise, as well as shared implicit topological feature views derived from inter-enterprise connections and visible topological feature views unique to each view. Collaborative learning of these two types of views is then used to discover collaborative enterprise groups across multiple industry chains.

[0008] There are m existing industrial chains, each of which has N enterprises. The multi-view enterprise set is represented by X = {X [1] ,X [2] ,…,X [m] ], the kth view is represented by the matrix X [k] express, d [k] is the number of features in the kth visible inherent feature view. The connection relationship of enterprises in the kth layer of the industrial chain is represented by the weight matrix W [k] Indicates that Representative companies and enterprises The degree of closeness of the relationship in the k-th layer of the industrial chain. The larger the value, the closer the connection between the two companies, and vice versa. A value of 0 means that the two companies have no connection. Based on the above known information, the specific solution of the present invention is as follows:

[0009] The first is the enterprise feature extraction module, which is mainly based on the non-negative matrix factorization (NMF) technology to extract the invisible inherent features, invisible topological features and visible topological features of the enterprise in each layer of the industrial chain, and complete the enterprise feature modeling of two types of views. NMF technology can be used to decompose the matrix. Depending on the given constraint environment, the meaning of the decomposition matrix obtained is also different. Because the inherent features of the enterprise may have negative values, in order to use NMF technology, it is necessary to convert X [k] Normalization makes each element non-negative. Then, based on the NMF technology, the following objective function can be established to extract the hidden inherent feature view information shared by each view:

[0010]

[0011] in Is the mapping matrix, used to map the data of the inherent invisible space, that is, B∈R r×N , to the feature space of the k-th view. r represents the number of hidden space features shared by m inherent feature spaces, satisfying 1≤r≤min{d [1] ,…,d [m] By minimizing the objective function The corresponding optimal solution B obtained can complete the extraction of the invisible inherent features shared by the enterprise.

[0012] As for the extraction of enterprise topological features, the NMF technology can be used to degrade the enterprise topological relationship W in multiple industrial chains into common invisible topological features and visible topological features under each layer of the industrial chain. This is achieved by minimizing the objective function:

[0013]

[0014] Among them H [k] ∈R N×l is the visible topological feature space unique to each layer of view, S∈R l×N is the hidden topological feature space shared by all views. l is the dimension of the topological feature space. Minimize The optimal solution H [k] and S are the visible and invisible topological feature extraction results of the enterprise.

[0015] Given a multi-view enterprise dataset X and a topological connection relationship W, the hidden space B shared by the inherent feature space and the visible topological feature space H and its hidden space S can be obtained through NMF technology. Then, a multi-view clustering algorithm that combines these two types of views for collaborative learning is used to obtain enterprise collaboration groups in multiple industrial chains. Fuzzy clustering algorithms have been widely used due to their simple process and significant effects, and have been extended to multi-view clustering. However, the existing multi-view fuzzy clustering algorithms lack the fusion of the inherent features and topological features of the clustering objects, resulting in inaccurate clustering results. Therefore, the present invention redefines the objective function of the multi-view fuzzy clustering algorithm as follows:

[0016]

[0017]

[0018] Where C is the number of enterprise collaboration groups, N is the total number of enterprises included in the entire complex system of multiple industrial chains, and u ij Is Enterprise x j The degree of belonging to group i. Group i is also divided into inherent eigenvectors q i and the topological eigenvector p i η and 1-η measure the importance of the intrinsic feature view and the topological feature view, respectively, and are set in advance. and φ=[φ1,φ2,…,φ m+1 ] are weight vectors of different intrinsic feature views and different topological feature views, respectively, each containing m visible views and 1 invisible view, and their values ​​can be adaptively adjusted during the learning process. Invisible view data X [m+1] =B and H [m+1]=S has been obtained through the previous NMF technology. The design ideas of this objective function can be explained in the following aspects:

[0019] 1) Collaborative learning of intrinsic feature view and topological feature view: the first term of the objective function This optimization of an existing multi-view fuzzy clustering algorithm incorporates consideration of topological feature views and controls the influence of these two types of views on clustering through the parameter η. The partitioning matrix U constrains the uniformity of multiple visible views and one invisible view. Furthermore, the distance between individuals and cluster centers in each layer of view is used to leverage enterprise-specific information. This approach adaptively achieves collaborative learning of intrinsic and topological feature views.

[0020] 2) Collaborative learning based on the least absolute shrinkage and selection operator (LASSO) mechanism: The second term in the objective function The network LASSO mechanism is used to further enhance the learning between different views. j and u z It is the unified partitioning information shared by all 2m visible views and 2 invisible views of enterprises i and z. It is the fusion of the topological relationships of all views. By minimizing the network LASSO term, we can absorb as much as possible the membership of individuals in different views relative to the connected adjacent nodes, thereby explicitly considering the connection relationship between each enterprise.

[0021] 3) Adaptive Learning of View Weights: The Third Item represents the non-negative Shannon entropy, which is used to balance the importance of each view and reconcile their differences. From an optimization perspective, if this term is removed, the importance of clearly discriminative views will be infinitely close to 1, while the importance of other views will be negligible. Alternatively, directly minimizing this term will make the importance of each view equal. By setting penalty factors ζ1 and ζ2 to balance these effects, the weights of different views can be adaptively adjusted.

[0022] By iteratively optimizing the above clustering objective function, the optimal fuzzy partitioning matrix U can be obtained, where each element u ij It indicates the degree of conformity between the characteristics of enterprise j and enterprise collaboration group i. The larger the value, the higher the degree of connection between the enterprise and the members in the group, and the more similar the enterprise characteristics are. Therefore, the enterprise collaboration group belonging of each enterprise can be expressed as label(x j )=arc max(u ij ) i=1,…,C, that is, the collaboration group corresponding to the maximum value of fuzzy partitioning. Each enterprise collaboration group in the entire multi-industry chain network can be represented as a tuple Π i =(q i ,p i ,γ i ),i=1,…,C. Among them, q i and p i are the inherent characteristics and topological characteristics of enterprise collaboration group i obtained by optimizing the objective function, γ i ={j:label(x j )=i} represents the set of enterprises belonging to group i. Thus, the design of the multi-view clustering algorithm DVHVC-MVFC, which is applicable to the discovery of enterprise collaboration groups in multiple industrial chains, is completed.

[0023] Through the above-mentioned enterprise feature extraction and multi-view clustering scheme design, the discovery of enterprise collaboration groups in multiple industrial chains can be completed. Subsequently, the discovery results can be used to assist enterprises in selecting cooperation partners. For example, enterprises in a collaboration group can be given priority. This is because the clustering results are based on the fact that the objects in a group are as closely connected as possible and have similar characteristics, which means that they have a good foundation for cooperation.

[0024] Beneficial effects:

[0025] (1) The present invention does not focus unilaterally on the utilization of a certain type of view information of an enterprise, but extracts the visible and hidden features of the inherent feature view and the topological feature view through NMF technology, so that each enterprise can obtain a more comprehensive modeling representation.

[0026] (2) The present invention designs a new multi-view fuzzy clustering algorithm to solve the problem that the current multi-view clustering technology cannot simultaneously utilize the inherent characteristics and topological characteristics of the samples, thereby realizing the collaborative learning of the two types of enterprise views in the clustering process and taking into account the network LASSO restriction based on topological relationships, unifying the collaborative learning and clustering division of visible-invisible perspectives in one framework.

[0027] (3) Introducing Shannon entropy in multi-view clustering can not only effectively reduce the impact of noisy views, but also coordinate the differences among views. It fully considers the data quality issues and differences among multiple industrial chains in actual application scenarios, thereby enhancing the collaboration of different views and more accurately locating their potential enterprise collaboration groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is the overall framework diagram of the method of the present invention.

[0029] Figure 2It is a multi-perspective commonality representation diagram of the inherent feature perspective and the topological feature perspective, where the left picture is the inherent feature perspective and the right picture is the topological feature perspective. DETAILED DESCRIPTION

[0030] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0031] As shown in the figure, the enterprise collaboration group discovery method disclosed in the present invention is applicable to multiple industrial chains. The specific implementation process is as follows:

[0032] Step 1: Determine the dimension d of each visible inherent feature space [k] , invisible inherent feature space dimension r, visible and invisible topological feature space dimension l, the connection relationship W of enterprises in each industrial chain [k] , the number of clusters C of multi-view enterprise data to be clustered, the number of views m, and the total number of enterprises N.

[0033] Step 2: Standardize the elements in each enterprise's visible inherent characteristic view to positive values ​​using the min-max method. The specific calculation method is as follows.

[0034]

[0035] Step 3: Randomly initialize the mapping matrix F [k] , invisible intrinsic feature matrix B, visible topological feature matrix H [k] , invisible topological feature S.

[0036] Step 4: Solve F using alternating iterative optimization [k] , B, the specific process is as follows:

[0037] 4.1 Fix B and update F [k] .

[0038]

[0039] 4.2 Fixed F [k] , update B.

[0040]

[0041] 4.3 Iterate steps 4.1 and 4.2 until the objective function Converge and obtain the optimal F [k] .

[0042] Step 5: Use alternating iterative optimization to solve H [k] , S, the specific process is as follows:

[0043] 5.1 Fix S, Update H [k].

[0044]

[0045] 5.2 Fixed H [k] , update S.

[0046]

[0047] 5.3 Iterate steps 5.1 and 5.2 until the objective function Converge and obtain the optimal H [k] and S.

[0048] Step 6: Construct the objective function formula J of the initialized multi-view clustering DVHVC-MVFC and calculate its value as follows:

[0049]

[0050]

[0051] η, ζ1 and ζ2 are hyperparameters, which can usually be determined using a grid search strategy.

[0052] Step 7: Randomly initialize the fuzzy partition matrix U, dual variables Determines the maximum number of iterations.

[0053] Step 8: The present invention adopts alternating iterative optimization to solve each optimization item. The specific process is as follows:

[0054] 8.1 Fixed U and Update the inherent feature space cluster center Q [k] .

[0055]

[0056] 8.2 Fixing U and φ [k] , update the topological feature space cluster center P [k] .

[0057]

[0058] 8.3 Fixed Q [k] and U, update the intrinsic feature view weights

[0059]

[0060]

[0061] 8.4 Fixed P [k] and U, update the topological feature view weight φ k .

[0062]

[0063]

[0064] 8.5 Calculation of intermediate parameter a ij , where the set The enterprises in the k-th view that have topological connection relationships with enterprise j are recorded.

[0065]

[0066] 8.6 Calculation of u ij The approximate value v ijc .

[0067]

[0068]

[0069] 8.7 Calculation of two variables

[0070]

[0071] 8.8 Update the fuzzy partition matrix U.

[0072]

[0073]

[0074]

[0075] Step 9: Repeat steps 8.1 to 8.8 until the maximum number of iterations is reached. And the optimal fuzzy partitioning matrix U is obtained. j Affiliated corporate collaboration group label(x j )=arc max(u ij ) i=1,…,C and corporate collaboration groups i =(q i ,p i ,γ i ),i=1,…,C, where γ i ={j:label(x j )=i} represents the set of enterprises belonging to group i.

[0076] Step 10: Based on the results of the enterprise collaboration group classification, enterprises can clarify the implicit enterprise collaboration groups across the entire multi-faceted industrial chain and understand the behavioral tendencies of current partners or competitors, which will facilitate future behavioral decisions. For example, when selecting subsequent partners within each multi-faceted industrial chain, priority should be given to enterprises belonging to the same collaboration group.

Claims

1. A collaborative enterprise discovery system applicable to multiple industry chains, characterized by: The system includes an enterprise feature extraction module and a multi-view clustering module. The feature extraction module is used to obtain shared hidden intrinsic feature views between visible intrinsic feature views of enterprises, as well as shared hidden topological feature views brought about by the connection relationships between enterprises and visible topological feature views unique to each view. The system uses the collaborative learning of these two types of views to achieve enterprise collaborative group discovery in multiple industrial chains. The enterprise feature extraction module is based on the non-negative matrix factorization (NMF) technology to extract the common invisible inherent features, invisible topological features and visible topological features of each layer of the industrial chain, and complete the enterprise feature modeling of two types of views; because the inherent features of the enterprise X [k] There may be negative values. In order to use NMF technology, X [k] Normalization makes each element non-negative; then based on the NMF technology, the following objective function can be established to extract the hidden inherent feature view information shared by each view: in Is the mapping matrix, used to map the data of the inherent invisible space, that is, B∈R r×N , to the feature space of the kth view; r represents the number of hidden space features shared by m inherent feature spaces, satisfying 1≤r≤min{d [1] ,…,d [m] By minimizing the objective function The corresponding optimal solution B obtained can complete the extraction of the invisible inherent features shared by the enterprise; The extraction of topological features uses NMF technology to degrade the enterprise topological relationship W in multiple industrial chains into common invisible topological features and visible topological features under each layer of the industrial chain. This is achieved by minimizing the objective function: Among them H [k] ∈R N×l is the visible topological feature space unique to each layer of view, S∈R l×N It is the invisible topological feature space shared by all views; l is the dimension of the topological feature space; minimize The optimal solution H [k] and S are the visible and invisible topological feature extraction results of the enterprise; The multi-view clustering module obtains the hidden space B and visible topological feature space H and its hidden space S shared by the inherent feature space and the known visible inherent feature space X through the enterprise feature extraction module. The multi-view clustering algorithm of collaborative learning of these two types of views obtains enterprise collaboration groups in multiple industrial chains. The existing multi-view fuzzy clustering algorithm lacks the fusion of the inherent features and topological features of the clustering objects, resulting in inaccurate clustering results. Therefore, the objective function of the multi-view fuzzy clustering algorithm is redefined as follows: Where C is the number of enterprise collaboration groups, N is the total number of enterprises included in the entire complex system of multiple industrial chains, and u ij Is Enterprise x j The degree of belonging to group i; group i is also divided into inherent eigenvectors q i and the topological eigenvector p i ; η and 1-η measure the importance of the intrinsic feature view and the topological feature view respectively, and are set in advance; and φ=[φ1,φ2,…,φ m+1 ] are weight vectors of different intrinsic feature views and different topological feature views, respectively, each containing m visible views and 1 invisible view, and their values ​​can be adaptively adjusted during the learning process; the invisible view data X [m+1] =B and H [m+1] =S has been obtained in advance; By iteratively optimizing the above clustering objective function, the optimal fuzzy partitioning matrix U can be obtained, where each element u ij It indicates the degree of conformity between the characteristics of enterprise j and enterprise collaboration group i. The larger the value, the higher the degree of connection between the enterprise and the members in the group, and the more similar the enterprise characteristics are. Therefore, the enterprise collaboration group belonging of each enterprise can be expressed as label(x j )=arcmax(u ij ) i=1, … ,C , that is, the collaboration group corresponding to the maximum value of fuzzy partitioning; and each enterprise collaboration group in the entire multi-industry chain network can be represented as a tuple Π i =(q i ,p i ,Υ i ),i=1,…,C; where q i and p i are the inherent characteristics and topological characteristics of the enterprise collaboration group i obtained by optimizing the objective function, i ={j:label(x j )=i} represents the set of enterprises belonging to group i; thus, the design of the multi-view clustering algorithm DVHVC-MVFC, which is suitable for discovering collaborative groups of enterprises in multiple industrial chains, is completed; The design method of the multi-view clustering algorithm is as follows: 1) Collaborative learning of intrinsic feature view and topological feature view: the first term of the objective function This algorithm optimizes the existing multi-view fuzzy clustering algorithm by adding the consideration of topological feature views and controlling the influence of these two types of views on clustering through the parameter η. It also restricts the unity of multiple visible views and one invisible view by partitioning the matrix U, while utilizing the unique information of each enterprise by calculating the distance between individuals and cluster centers in each layer of views. This method can adaptively realize the collaborative learning of intrinsic feature view and topological feature view; 2) Collaborative learning based on the network minimum absolute shrinkage rate and the selection operator LASSO mechanism: the second term in the objective function The network LASSO mechanism is used to further enhance the learning between different views; the vector u j and u z It is the unified partitioning information shared by all 2m visible views and 2 invisible views of enterprises i and z; It is the fusion of the topological relationships of all views. By minimizing the network LASSO term, it can absorb as much as possible the membership of individuals in different views relative to the connected adjacent nodes, thereby explicitly considering the connection relationship between each enterprise. 3) Adaptive Learning of View Weights: The Third Item represents the non-negative Shannon entropy, which is used to balance the importance of each view and coordinate the differences between them. From an optimization perspective, if this term is removed, the importance of the obviously distinctive view will be infinitely close to 1, while the importance of other views will be negligible. In addition, directly minimizing this term will make the importance of each view equal. By setting penalty factors ζ1 and ζ2 to balance these effects, the weights of different views can be adaptively adjusted.