Method and apparatus for multi-view cancer gene data clustering ensemble based on width learning
Patent Information
- Application Number
- CN202410311105.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-03-19
AI Technical Summary
但是,癌症基因多视图数据聚类需要综合考虑不同视图之间的信息,权衡不同视图之间的重要性和解决高维数据的处理,导致现有方法的计算量大、聚类质量不足,限制了多视图聚类方法在现实中的应用
[0044]本发明提出的基于宽度学习的多视图的癌症基因数据聚类集成方法利用宽度学习网络的性能优势,进行多视图的癌症基因数据的聚类处理,得到兼具效率与性能的自编码器模型和聚类集成模型,此方法无需消耗大规模计算资源,即使是在一台普通的计算机上,都可以轻松运行。其中在自编码器模型中引入了子空间自表达结构,全面考虑样本高维信息,并且在集成步骤进行模糊处理和置信度计算,丰富信息,有效提升网络模型的鲁棒性和准确性,因而在实际场景中更具适用性。
Smart Images

Figure CN118155731B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning, and specifically to a method and apparatus for clustering and integrating cancer gene data based on width learning and multiple views. Background Technology
[0002] Clustering ensemble methods can classify cancer genes into different subtypes, each with unique gene expression patterns and biological characteristics, corresponding to different clinical phenotypes and potentially varying treatment responses. Therefore, identifying the differences and similarities between different types of cancer is crucial. Multi-view clustering integrates heterogeneous data from multiple views of cancer genes. Different views can provide different information, leading to a more comprehensive and objective analysis and understanding of cancer genes. However, multi-view clustering of cancer genes requires comprehensive consideration of information from different views, weighing their importance, and handling high-dimensional data. This results in high computational costs and insufficient clustering quality in existing methods, limiting the practical application of multi-view clustering methods. Summary of the Invention
[0003] The purpose of this application is to propose a method and apparatus for clustering and integrating cancer gene data based on width learning and multiple views, addressing the aforementioned technical problems.
[0004] In a first aspect, the present invention provides a multi-view cancer gene data clustering and ensemble method based on width learning, comprising the following steps:
[0005] Obtain multi-view cancer gene data;
[0006] An autoencoder model is constructed, which includes a first-width learning network and a subspace self-expression structure connected in sequence. The subspace self-expression structure is trained to determine the coefficient matrix of the trained subspace self-expression structure. The autoencoder model is trained based on the coefficient matrix of the trained subspace self-expression structure to determine the weights of the trained autoencoder model. Multi-view cancer gene data is input into the trained autoencoder model to obtain a sample matrix after feature processing.
[0007] Cluster the sample matrix after feature processing to obtain multiple basic clustering results. Use the basic clustering results as ensemble members in the ensemble pool. Construct a fuzzy partitioning matrix based on the fuzzy membership degree of the ensemble members. Randomly set a basic clustering result as a pseudo-label. Construct a confidence matrix based on the confidence degree of the sample matrix after feature processing using the pseudo-label.
[0008] A clustering ensemble model based on a second-width learning network is constructed. The clustering ensemble model is trained according to the confidence matrix to obtain the trained clustering ensemble model. The fuzzy partitioning matrix is input into the trained clustering ensemble model to obtain the soft ensemble result. The soft ensemble result is then clustered to obtain the clustering result of the multi-view cancer gene data.
[0009] As a preferred embodiment, the objective function of the subspace self-expression structure during training is shown in the following equation:
[0010]
[0011] in, Let represent a sample dataset consisting of cancer gene data from multiple views, n represent the number of samples in the sample dataset, d represent the feature dimension of the samples in the sample dataset, γ represent the trade-off parameter of the subspace self-expression structure, T represent the transpose of the matrix, and Q represent the coefficient matrix.
[0012] As a preferred embodiment, the objective function of the autoencoder model during training is shown in the following equation:
[0013]
[0014] Among them, A p This represents the output features of the hidden layer of the first-width learning network. The sample dataset passes sequentially through the input layer and hidden layer of the first-width learning network to obtain the output features of the hidden layer. This represents a sample dataset consisting of cancer gene data from multiple views, where n represents the number of samples in the sample dataset, d represents the feature dimension of the samples in the sample dataset, and W... p The weights represent the weights of the trained autoencoder model, α represents the weighting parameter for the regularization term, and β represents the weighting parameter for the self-expression error term. Let Frobenius norm be calculated, and Q be the coefficient matrix of the trained subspace self-expression structure;
[0015] The formula for calculating the sample matrix after feature processing is as follows:
[0016]
[0017] Among them, X * This represents the sample matrix after feature processing.
[0018] As a preferred method, a fuzzy partitioning matrix is constructed based on the fuzzy membership degrees of the integrated members, specifically including:
[0019] The similarity between two clusters in each ensemble member is calculated using a fuzzy membership function, as shown in the following equation:
[0020]
[0021] in, Let a represent class a of the i-th basic clustering result; This represents the cluster center of class a in the i-th basic clustering result. Let b represent the i-th basic clustering result; This represents the cluster center of class b in the i-th basic clustering result. Let j represent the cluster center of class j in the i-th basic clustering result, where j = 1, 2, ..., k. i k i This indicates that the i-th basic clustering result includes k non-overlapping clusters; Let P represent the Euclidean distance between the centers of two clusters. Based on the fuzzy membership function, a fuzzy partitioning matrix P is constructed as shown in the following equation:
[0022]
[0023] Where, x a This represents a sample belonging to class a in the i-th basic clustering result.
[0024] Preferably, a confidence matrix is constructed based on the confidence of the sample matrix after feature processing using pseudo-labels, specifically including:
[0025] The confidence level of the pseudo-label on the feature-processed sample matrix is calculated using the following formula:
[0026]
[0027]
[0028] Where, x i X represents the sample matrix after feature processing. * In the i-th sample data, conf() represents the confidence level of the sample data, and q represents the similarity measure between the sample data and the pseudo-label; SC(x i ) indicates that in the pseudo-tag, x i The set of samples assigned to the same cluster; This indicates that in the pseudo-tag, x i Sample sets assigned to different clusters; CA ij ∈[0,1] indicates that x is in the ensemble pool i and x j The proportion of basic clusters to which individuals are assigned to the same category, x j X represents the sample matrix after feature processing. * The j-th sample data in the dataset; θ represents the confidence trade-off parameter.
[0029] As a preferred option, the objective function of the clustering ensemble model during training is shown in the following equation:
[0030]
[0031] Where A represents the output feature of the hidden layer of the second-width learning network, the fuzzy partitioning matrix P is passed sequentially through the input layer and hidden layer of the second-width learning network to obtain the output feature of the hidden layer of the second-width learning network, and W represents the weights of the trained clustering ensemble model. X represents the confidence matrix, which is the sample matrix after feature processing. * The confidence scores are constructed using diagonal elements, where diag() represents the diagonal matrix; G represents the guidance information matrix, which is a binary partitioning matrix of pseudo-labels; L = DS is the graph Laplacian matrix, where S is a sparse pairwise similarity matrix, and the fuzzy partitioning matrix P is calculated using the K-nearest neighbor algorithm; D represents the degree matrix, which is a diagonal matrix about S, and its diagonal elements are generated by... The calculation shows that n represents the number of samples in the sample dataset; ∈ represents the trade-off parameter of the manifold constraint;
[0032] The calculation process for the soft integration results is as follows:
[0033] O = AW;
[0034] Where O represents the soft integration result.
[0035] Secondly, the present invention provides a multi-view cancer gene data clustering and integration device based on width learning, comprising:
[0036] The data acquisition module is configured to acquire cancer gene data from multiple views;
[0037] The autoencoder model building module is configured to build an autoencoder model, which includes a first-width learning network and a subspace self-expression structure connected in sequence. The subspace self-expression structure is trained to determine the coefficient matrix of the trained subspace self-expression structure. The autoencoder model is trained based on the coefficient matrix of the trained subspace self-expression structure to determine the weights of the trained autoencoder model. Multi-view cancer gene data is input into the trained autoencoder model to obtain a feature-processed sample matrix.
[0038] The integration module is configured to cluster the sample matrix after feature processing to obtain multiple basic clustering results. The basic clustering results are used as integration members in the integration pool. A fuzzy partitioning matrix is constructed based on the fuzzy membership degree of the integration members. A basic clustering result is randomly set as a pseudo-label. A confidence matrix is constructed based on the confidence degree of the sample matrix after feature processing based on the pseudo-label.
[0039] The clustering module is configured to construct a clustering ensemble model based on a second-width learning network. The clustering ensemble model is trained according to the confidence matrix to obtain a trained clustering ensemble model. The fuzzy partitioning matrix is input into the trained clustering ensemble model to obtain a soft ensemble result. The soft ensemble result is then clustered to obtain the clustering result of the multi-view cancer gene data.
[0040] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0041] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.
[0042] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the implementations in the first aspect.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] This invention proposes a multi-view cancer gene data clustering and ensemble method based on width learning. Leveraging the performance advantages of width learning networks, it performs clustering processing on multi-view cancer gene data, resulting in an efficient and high-performance autoencoder model and clustering ensemble model. This method does not require large-scale computing resources and can be easily run even on a standard computer. The autoencoder model incorporates a subspace self-expression structure to comprehensively consider high-dimensional information from the samples. Furthermore, fuzzing and confidence calculations are performed in the ensemble step to enrich the information and effectively improve the robustness and accuracy of the network model, making it more applicable in real-world scenarios. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating a multi-view cancer gene data clustering and integration method based on width learning, as an embodiment of this application.
[0047] Figure 2This is a block diagram illustrating a multi-view cancer gene data clustering and ensemble method based on width learning, as an embodiment of this application.
[0048] Figure 3 This is a schematic diagram of a multi-view cancer gene data clustering and integration device based on width learning, as an embodiment of this application.
[0049] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0051] Figure 1 This application illustrates an embodiment of a multi-view cancer gene data clustering and ensemble method based on width learning, comprising the following steps:
[0052] S1, acquire multi-view cancer gene data.
[0053] Specifically, different subtypes of cancer genes have different views. Taking lung cancer as an example, different subtypes of lung cancer genes, such as lung adenocarcinoma and lung squamous cell carcinoma, not only have different images, but also differ in their location in the trachea. Therefore, they can be divided into cancer gene data based on image views and cancer gene data based on the location of the disease. Taking this as an example, multi-view cancer gene data is collected. The multi-view cancer gene dataset is obtained as samples in the sample dataset, where each sample is unlabeled.
[0054] S2, Construct an autoencoder model, which includes a first-width learning network and a subspace self-expression structure connected in sequence. Train the subspace self-expression structure to determine the coefficient matrix of the trained subspace self-expression structure. Train the autoencoder model based on the coefficient matrix of the trained subspace self-expression structure to determine the weights of the trained autoencoder model. Input the multi-view cancer gene data into the trained autoencoder model to obtain the feature-processed sample matrix.
[0055] In a specific embodiment, the objective function of the subspace self-expression structure during the training process is shown in the following equation:
[0056]
[0057] in, Let represent a sample dataset consisting of cancer gene data from multiple views, n represent the number of samples in the sample dataset, d represent the feature dimension of the samples in the sample dataset, γ represent the trade-off parameter of the subspace self-expression structure, T represent the transpose of the matrix, and Q represent the coefficient matrix.
[0058] In a specific embodiment, the objective function of the autoencoder model during training is shown in the following equation:
[0059]
[0060] Among them, A p This represents the output features of the hidden layer of the first-width learning network. The sample dataset passes sequentially through the input layer and hidden layer of the first-width learning network to obtain the output features of the hidden layer. This represents a sample dataset consisting of cancer gene data from multiple views, where n represents the number of samples in the sample dataset, d represents the feature dimension of the samples in the sample dataset, and W... p The weights represent the weights of the trained autoencoder model, α represents the weighting parameter for the regularization term, and β represents the weighting parameter for the self-expression error term. Let Frobenius norm be calculated, and Q be the coefficient matrix of the trained subspace self-expression structure;
[0061] The formula for calculating the sample matrix after feature processing is as follows:
[0062]
[0063] Among them, X * This represents the sample matrix after feature processing.
[0064] For details, please refer to Figure 2 An autoencoder model based on a first-width learning network and a subspace self-representation structure is constructed to extract features and reduce dimensionality from a sample dataset. The first-width learning network consists of an input layer, hidden layers, and an output layer. It calculates the weights of each feature node and augmentation node by taking a pseudo-inverse, eliminating the need for backpropagation to change the kernel of the feature extractor, thus making the training process more efficient and flexible. The first-width learning network significantly outperforms existing deep neural networks in training speed. A subspace self-representation structure is added to the first-width learning network, and an objective function is used to train the subspace self-representation structure. The objective function used during training aims to calculate the optimal sample matrix Q of the subspace self-representation structure, i.e., the trained sample matrix Q. This trained sample matrix Q is then used as a parameter in the autoencoder model to enrich the subspace information. The objective function used to train the autoencoder model aims to calculate the optimal weight W. p,pass Obtain the sample matrix X after feature processing * .
[0065] S3. Cluster the sample matrix after feature processing to obtain multiple basic clustering results. Use the basic clustering results as ensemble members in the ensemble pool. Construct a fuzzy partitioning matrix based on the fuzzy membership degree of the ensemble members. Randomly set a basic clustering result as a pseudo-label. Construct a confidence matrix based on the confidence degree of the sample matrix after feature processing using the pseudo-label.
[0066] In a specific embodiment, a fuzzy partitioning matrix is constructed based on the fuzzy membership degree of the integrated members, specifically including:
[0067] The similarity between two clusters in each ensemble member is calculated using a fuzzy membership function, as shown in the following equation:
[0068]
[0069] in, Let a represent class a of the i-th basic clustering result; This represents the cluster center of class a in the i-th basic clustering result. Let b represent the i-th basic clustering result; This represents the cluster center of class b in the i-th basic clustering result. Let j represent the cluster center of class j in the i-th basic clustering result, where j = 1, 2, ..., k. i k i This indicates that the i-th basic clustering result includes k non-overlapping clusters; Let P represent the Euclidean distance between the centers of two clusters. Based on the fuzzy membership function, a fuzzy partitioning matrix P is constructed as shown in the following equation:
[0070]
[0071] Where, x a This represents a sample belonging to class a in the i-th basic clustering result.
[0072] In a specific embodiment, a confidence matrix is constructed based on the confidence of the sample matrix after feature processing using pseudo-labels, specifically including:
[0073] The confidence level of the pseudo-label on the feature-processed sample matrix is calculated using the following formula:
[0074]
[0075]
[0076] Where, xi X represents the sample matrix after feature processing. * In the i-th sample data, conf() represents the confidence level of the sample data, and q represents the similarity measure between the sample data and the pseudo-label; SC(x i ) indicates that in the pseudo-tag, x i The set of samples assigned to the same cluster; This indicates that in the pseudo-tag, x i Sample sets assigned to different clusters; CA ij ∈[0,1] indicates that x is in the ensemble pool i and x j The proportion of basic clusters to which individuals are assigned to the same category, x j X represents the sample matrix after feature processing. * The j-th sample data in the dataset; θ represents the confidence trade-off parameter.
[0077] Specifically, the sample matrix X after feature processing * The reduced-dimensional data needs further fuzzing to enrich the ensemble pool information as ensemble members, and to calculate the pseudo-label pair for the feature-processed sample matrix X. * The confidence level. The fuzzy partitioning matrix P is generated from the fuzzy membership degrees of the ensemble members. The K-Means clustering algorithm is used to process the feature-processed sample matrix X. * Clustering is performed to obtain multiple basic clustering results, which are called ensemble members. Their set is called the ensemble pool. For each ensemble member, a fuzzy membership function can be used to measure the similarity between two clusters. Confidence is used to determine the consistency between the sample partitioning in the pseudo-label and the partitioning in the ensemble pool. Therefore, a basic clustering result is randomly selected as the pseudo-label. The pseudo-label is applied to the feature-processed sample matrix X. * The confidence levels are different. Therefore, it is necessary to calculate the pseudo-label pair for the feature-processed sample matrix X. * The confidence level. Further, based on the pseudo-labels, the feature-processed sample matrix X... * Construct a confidence matrix based on the confidence scores.
[0078] S4. Construct a clustering ensemble model based on a second-width learning network. Train the clustering ensemble model according to the confidence matrix to obtain the trained clustering ensemble model. Input the fuzzy partitioning matrix into the trained clustering ensemble model to obtain the soft ensemble result. Cluster the soft ensemble result to obtain the clustering result of the multi-view cancer gene data.
[0079] In a specific embodiment, the objective function of the clustering ensemble model during the training process is shown in the following equation:
[0080]
[0081] Where A represents the output feature of the hidden layer of the second-width learning network, the fuzzy partitioning matrix P is passed sequentially through the input layer and hidden layer of the second-width learning network to obtain the output feature of the hidden layer of the second-width learning network, and W represents the weights of the trained clustering ensemble model.
[0082] X represents the confidence matrix, which is the sample matrix after feature processing. * The confidence scores are constructed using diagonal elements, where diag() represents the diagonal matrix; G represents the guidance information matrix, which is a binary partitioning matrix of pseudo-labels. If sample x a If it belongs to class B in the pseudo-tags, then let its corresponding element G be... aB If the value is 1, then let the corresponding element G be 1; otherwise, let the value be 1. aB The value is 0; L = DS is the graph Laplacian matrix, where S is a sparse pairwise similarity matrix, calculated from the fuzzy partition matrix P using the K-nearest neighbor algorithm; D represents the degree matrix, a diagonal matrix about S, whose diagonal elements are derived from... The calculation shows that n represents the number of samples in the sample dataset; ∈ represents the trade-off parameter of the manifold constraint;
[0083] The calculation process for the soft integration results is as follows:
[0084] O = AW;
[0085] Where O represents the soft integration result.
[0086] Specifically, this clustering ensemble model employs a second-width learning network, which also includes an input layer, an output layer, and hidden layers. The sample matrix X after feature processing using pseudo-labels... * A confidence matrix is constructed using the confidence scores as diagonal elements. This confidence matrix is used in the calculation of the objective function of the clustering ensemble model. Combined with the guidance information matrix and the graph Laplacian matrix, it integrates local sample adjacency information and global ensemble member guidance, allowing the fuzzy partitioning matrix P to be input into the trained clustering ensemble model. After data processing, the soft ensemble result is output. The K-Means clustering algorithm is used to cluster this soft ensemble result, ultimately obtaining the clustering results for the multi-view cancer gene data.
[0087] The steps S1-S4 above do not necessarily represent the order of the steps, but are represented by step symbols. The order of the steps can be adjusted.
[0088] Further reference Figure 3 As an implementation of the methods shown in the above figures, this application provides an embodiment of a multi-view cancer gene data clustering and integration device based on width learning. This device embodiment is similar to... Figure 1Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0089] This application provides a multi-view cancer gene data clustering and integration device based on width learning, including:
[0090] Data acquisition module 1 is configured to acquire cancer gene data from multiple views;
[0091] Autoencoder model building module 2 is configured to build an autoencoder model, which includes a first-width learning network and a subspace self-expression structure connected in sequence. The subspace self-expression structure is trained to determine the coefficient matrix of the trained subspace self-expression structure. The autoencoder model is trained based on the coefficient matrix of the trained subspace self-expression structure to determine the weights of the trained autoencoder model. Multi-view cancer gene data is input into the trained autoencoder model to obtain a sample matrix after feature processing.
[0092] Integration module 3 is configured to cluster the sample matrix after feature processing to obtain multiple basic clustering results, use the basic clustering results as integration members in the integration pool, construct a fuzzy partitioning matrix based on the fuzzy membership degree of the integration members, randomly set a basic clustering result as a pseudo label, and construct a confidence matrix based on the confidence degree of the sample matrix after feature processing using the pseudo label.
[0093] Clustering module 4 is configured to construct a clustering ensemble model based on a second-width learning network. The clustering ensemble model is trained according to the confidence matrix to obtain a trained clustering ensemble model. The fuzzy partitioning matrix is input into the trained clustering ensemble model to obtain a soft ensemble result. The soft ensemble result is then clustered to obtain the clustering result of the multi-view cancer gene data.
[0094] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 4 As shown, the electronic device in this embodiment includes a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; and the processor 401 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0095] Alternatively, the memory 402 can be either standalone or integrated with the processor 401.
[0096] When the memory 402 is set up independently, the electronic device also includes a bus 403 for connecting the memory 402 and the processor 401.
[0097] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by a processor, implement the method described above.
[0098] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0099] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0100] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0101] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0102] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0103] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0104] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0105] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0106] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0107] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.
[0108] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-view clustering and ensemble method for cancer gene data based on width learning, characterized in that, Includes the following steps: Obtain multi-view cancer gene data; An autoencoder model is constructed, comprising a first-width learning network and a subspace self-expression structure connected in sequence. The subspace self-expression structure is trained to determine the coefficient matrix of the trained subspace self-expression structure. Based on the coefficient matrix of the trained subspace self-expression structure, the autoencoder model is trained to determine the weights of the trained autoencoder model. The multi-view cancer gene data is input into the trained autoencoder model to obtain a feature-processed sample matrix. Cluster the sample matrix after feature processing to obtain multiple basic clustering results. Use the basic clustering results as ensemble members in the ensemble pool. Construct a fuzzy partitioning matrix based on the fuzzy membership degree of the ensemble members. Randomly set a basic clustering result as a pseudo-label. Construct a confidence matrix based on the confidence degree of the sample matrix after feature processing using the pseudo-label. A clustering ensemble model based on a second-width learning network is constructed. The clustering ensemble model is trained according to the confidence matrix to obtain a trained clustering ensemble model. The fuzzy partitioning matrix is input into the trained clustering ensemble model to obtain a soft ensemble result. The soft ensemble result is then clustered to obtain the clustering result of the multi-view cancer gene data.
2. The cancer gene data clustering and ensemble method based on width learning with multiple views according to claim 1, characterized in that, The objective function of the subspace self-expression structure during training is shown in the following equation: in, Let n represent the sample dataset consisting of cancer gene data from multiple views, d represent the feature dimension of the samples in the sample dataset, γ represent the trade-off parameter of the self-expression structure of the subspace, T represent the transpose of the matrix, and Q represent the coefficient matrix.
3. The cancer gene data clustering and ensemble method based on width learning with multiple views according to claim 1, characterized in that, The objective function of the autoencoder model during training is shown in the following equation: Among them, A p This represents the output features of the hidden layer of the first-width learning network. The sample dataset passes sequentially through the input layer and hidden layer of the first-width learning network to obtain the output features of the hidden layer. This represents a sample dataset consisting of cancer gene data from multiple views, where n represents the number of samples in the sample dataset, d represents the feature dimension of the samples in the sample dataset, and W... p The weights represent the weights of the trained autoencoder model, α represents the weighting parameter for the regularization term, and β represents the weighting parameter for the self-expression error term. Let Frobenius norm be calculated, and Q be the coefficient matrix of the trained subspace self-expression structure; The formula for calculating the sample matrix after feature processing is as follows: Among them, X * This represents the sample matrix after feature processing.
4. The cancer gene data clustering and ensemble method based on width learning with multiple views according to claim 1, characterized in that, The construction of the fuzzy partitioning matrix based on the fuzzy membership degrees of the integrated members specifically includes: The similarity between two clusters in each ensemble member is calculated using a fuzzy membership function, as shown in the following equation: in, Let a represent class a of the i-th basic clustering result; This represents the cluster center of class a in the i-th basic clustering result. Let b represent the i-th basic clustering result; This represents the cluster center of class b in the i-th basic clustering result. Let j represent the cluster center of class j in the i-th basic clustering result, where j = 1, 2, ..., k. i k i This indicates that the i-th basic clustering result includes k non-overlapping clusters; Let P represent the Euclidean distance between the centers of two clusters. Based on the fuzzy membership function, a fuzzy partitioning matrix P is constructed as shown in the following equation: Where, x a This represents a sample belonging to class a in the i-th basic clustering result.
5. The method for clustering and integrating cancer gene data based on width learning with multiple views according to claim 1, characterized in that, The construction of a confidence matrix based on the confidence of the sample matrix after feature processing using the pseudo-labels specifically includes: The confidence level of the pseudo-label on the feature-processed sample matrix is calculated using the following formula: Where, x i X represents the sample matrix after feature processing. * In the i-th sample data, conf() represents the confidence level of the sample data, and q represents the similarity measure between the sample data and the pseudo-label; SC(x i ) indicates that in the pseudo-tag, x i The set of samples assigned to the same cluster; This indicates that in the pseudo-tag, x i Sample sets assigned to different clusters; CA ij ∈[0,1] indicates that x in the integrated pool i and x j The proportion of basic clusters to which individuals are assigned to the same category, x j X represents the sample matrix after feature processing. * The j-th sample data in the dataset; θ represents the confidence trade-off parameter.
6. The method for clustering and integrating cancer gene data based on width learning with multiple views according to claim 5, characterized in that, The objective function of the clustering ensemble model during training is shown in the following equation: Where A represents the output feature of the hidden layer of the second-width learning network, the fuzzy partitioning matrix P is passed sequentially through the input layer and hidden layer of the second-width learning network to obtain the output feature of the hidden layer of the second-width learning network, and W represents the weights of the trained clustering ensemble model. X represents the confidence matrix, which is the sample matrix after feature processing. * The confidence scores are constructed using diagonal elements, where diag() represents the diagonal matrix; G represents the guidance information matrix, which is a binary partitioning matrix of pseudo-labels; L = DS is the graph Laplacian matrix, where S is a sparse pairwise similarity matrix, calculated from the fuzzy partitioning matrix P using the K-nearest neighbor algorithm; D represents the degree matrix, a diagonal matrix about S, whose diagonal elements are generated by... The calculation shows that n represents the number of samples in the sample dataset; ∈ represents the trade-off parameter of the manifold constraint; The calculation process for the soft integration result is as follows: O = AW; Where O represents the soft integration result.
7. A multi-view cancer gene data clustering and integration device based on width learning, characterized in that, include: The data acquisition module is configured to acquire cancer gene data from multiple views; The autoencoder model building module is configured to build an autoencoder model, which includes a first-width learning network and a subspace self-expression structure connected in sequence. The subspace self-expression structure is trained to determine the coefficient matrix of the trained subspace self-expression structure. Based on the coefficient matrix of the trained subspace self-expression structure, the autoencoder model is trained to determine the weights of the trained autoencoder model. The multi-view cancer gene data is input into the trained autoencoder model to obtain a feature-processed sample matrix. An integration module is configured to cluster the feature-processed sample matrix to obtain multiple basic clustering results, use the basic clustering results as integration members in an integration pool, construct a fuzzy partitioning matrix based on the fuzzy membership degree of the integration members, randomly set a basic clustering result as a pseudo-label, and construct a confidence matrix based on the confidence degree of the feature-processed sample matrix using the pseudo-label. The clustering module is configured to construct a clustering ensemble model based on a second-width learning network, train the clustering ensemble model according to the confidence matrix to obtain a trained clustering ensemble model, input the fuzzy partitioning matrix into the trained clustering ensemble model to obtain a soft ensemble result, and cluster the soft ensemble result to obtain the clustering result of the multi-view cancer gene data.
8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Pneumovirus gene data multi-view clustering integration method and device and electronic equipment
CN119513631A