A Chronic Disease Comorbidity Pattern Recognition Method, Device, Electronic Device and Storage Medium Incorporating Graph Convolution Constraint Enhancement for Low-Rank Representation
By building a chronic disease comorbidity network and using a low-rank representation method with fusion graph convolution constraints, the comorbidity pattern of chronic disease is identified, which solves the problem of insufficient recognition accuracy of comorbidity pattern in the existing technology, and achieves higher recognition accuracy and network structure feature retention.
Patent Information
- Application Number
- CN202411366928.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The prior art is difficult to effectively explore the deep relationships and complex relationships between chronic diseases, resulting in insufficient accuracy in identifying comorbid disease patterns.
A low-rank representation method with fusion graph convolution constraints is adopted to build a chronic disease comorbidity network, and comorbidity patterns are identified through low-rank representation learning and community division mechanisms, and the network structure characteristics are maintained in combination with graph convolutional regularity and graph Laplace regularity methods.
The accuracy and community detection capabilities of comorbid disease pattern recognition are improved, the global nonlinear and local feature representation capabilities of the network are maintained, and the numerical non-negative characteristics of the model are ensured.
Smart Images

Figure CN119312111B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of pattern recognition, and in particular to a chronic disease comorbidity pattern recognition method, device, electronic device and storage medium that fuse graph convolution constraints to enhance low-rank representation. Background Art
[0002] Chronic non-communicable diseases (hereinafter referred to as "chronic diseases") have become the main threats to global public health and human health, bringing a heavy economic burden to patients and society. Due to the characteristics of complex etiology, insidious onset, long course of disease and protracted illness, some patients often suffer from multiple chronic diseases at the same time. The World Health Organization refers to this phenomenon as "comorbidity". Under the trend of global population aging, the problem of chronic disease comorbidity has gradually emerged and become a major challenge in the global health field. The complex interactions between diseases pose great challenges to disease diagnosis and treatment. Understanding the specific factors and processes of comorbidity, the interactions and possible synergistic effects between diseases is very important for promoting diagnosis, improving the quality of life of patients, improving prevention and reducing the costs of the healthcare system. Therefore, using new technologies such as artificial intelligence and big data to explore the complex comorbidity patterns between chronic diseases provides important decision-making basis for predicting disease risks, disease prevention, treatment and management, and has practical significance for the sustainable development of human health.
[0003] Diseases do not cluster randomly. Comorbidity patterns reveal the statistically significant associations and pathophysiological relationships between diseases. Mining comorbidity patterns and exploring their distribution laws help disease classification, reduce the complexity of chronic disease prevention and control, and make prevention measures more accurate and efficient. A systematic review of multi-disease epidemiological studies shows that different studies have different findings on the nature and patterns of comorbidity. Currently, there is no unified definition for the method of identifying comorbidity patterns. According to different data analysis principles and technical means, the common methods are summarized into the following four categories:
[0004] (1) Odds ratio: used to measure the ratio of the probability of something happening to not happening, and is often used to analyze the relative risk of two diseases in comorbidity pattern analysis. A study in the Netherlands showed that the relative risks of the correlations between depression and anxiety, coronary heart disease and heart failure, and chronic obstructive pulmonary disease and heart failure are the largest. The odds ratio algorithm is simple and easy to use and has become one of the common methods in comorbidity pattern research.
[0005] (2) Factor analysis: aims to identify the complex relationships between multiple diseases, combines multiple related factors into a few comprehensive indicators, simplifies the data structure and reveals the variable associations. A study on Japanese adults identified 5 comorbidity patterns of 17 chronic diseases through factor analysis. Factor analysis does not rely on traditional grouping assumptions, extracts important information by dimension reduction and data simplification, and can more effectively reveal variable relationships.
[0006] (3) Cluster analysis: A data mining technique used to identify data clusters. Common methods include hierarchical clustering and K-means clustering. The comorbidity patterns of patients with chronic obstructive pulmonary disease may include comorbidities with cardiovascular diseases, allergic diseases, and other representative diseases. Cluster analysis can explore the commonalities of potential factors in these comorbidity patterns to help understand the phenomenon of chronic disease comorbidity.
[0007] (4) Association rules: Used to analyze the direct associations between multiple chronic diseases and reveal strong associations in comorbidity patterns. A study in the UK evaluated the comorbidity patterns of middle-aged and elderly people through cluster analysis and association rule mining, and found 3 clusters and 30 disease patterns, revealing the complex relationships and central roles among different diseases, such as diabetes, hypertension, and asthma. There are complex interactions and associations among these chronic diseases, and comprehensive evaluation and treatment are needed to achieve better health outcomes.
[0008] Due to the complex interactions and influence relationships among diseases, traditional methods are difficult to mine the deep relationships and complex connections among diseases and difficult to obtain accurate comorbidity patterns. With the development of complex science, complex network analysis describes complex relationships from an overall and systematic perspective, which is conducive to comprehensively revealing the complex influence laws among diseases from an overall perspective and has become the mainstream tool for exploring chronic disease comorbidity patterns. Community detection algorithms can discover closely connected clusters in the network and thus be used to identify disease co-occurrence patterns.
[0009] The present invention designs a low-rank representation modeling method fused with graph convolution constraints to extract non-linear features in the target comorbidity network, and at the same time incorporates graph convolution regularization and graph Laplacian regularization methods to maintain global and local structural features, thereby enhancing the model's representation learning ability to obtain higher community detection accuracy and comorbidity pattern recognition accuracy. Summary of the Invention
[0010] Based on this, it is necessary to provide a chronic disease comorbidity pattern recognition method, device, electronic device, and storage medium that fuse graph convolution constraints to enhance low-rank representation for existing problems.
[0011] In a first aspect, an embodiment of the present application provides a chronic disease comorbidity pattern recognition method that fuses graph convolution constraints to enhance low-rank representation, characterized by including:
[0012] Obtain medical data resources, where the medical data resources include data records of chronic disease prevalence sorted out from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources;
[0013] Construct a chronic disease comorbidity network using medical data resources. Among them, the chronic disease comorbidity network is used to describe the correlation and influence between diseases, and can be constructed by calculating the correlation between diseases or patients based on multi-source data. The multi-source data includes patient basic information, as well as test, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data. Among them, the chronic disease comorbidity network is represented by a graph G=(V, E), where V={v i |i∈{1,...,n}} represents a set containing n disease nodes, and E={e ij |i,j∈{1,...,m}} represents a set containing m edges; the adjacency matrix A=[a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between disease nodes v i and v j , and its value is equal to I(s i , s j );
[0014] Perform low-rank representation learning on the chronic disease comorbidity network to mine and output the community structure therein. Among them, the low-rank representation learning uses a low-rank representation method that combines graph convolution constraints to represent the comorbidity network;
[0015] Use the community partitioning mechanism to identify the community structure in the comorbidity network and discover the comorbidity pattern.
[0016] Preferably, the construction method of the chronic disease comorbidity network includes:
[0017] Sort out the chronic disease records and sort out the data records from the test, diagnosis, treatment, prescription, electronic medical records, and clinical data;
[0018] Count the frequency of disease occurrence, and count the frequency Q(s i ) of each disease occurrence and the frequency Q(s i , s j ) of the co-occurrence of disease combinations from the data records. Among them, if a certain disease s i appears in a disease record, it is counted once, and if a certain two diseases s i and s j co-appear in a disease record, it is counted once, where i,j = 1,2,...,n represents the number of diseases of concern;
[0019] Calculate the mutual information value between two diseases. The calculation method is as follows:
[0020]
[0021] Among them, p(s i ) = Q(si ) / F, p(s j ) = Q(s j ) / F are used to calculate the observed frequencies of diseases s i and s j respectively, where F represents the total number of observations, and p(s i , s j ) = Q(s i , s j ) / F is used to calculate the frequency of observing diseases s i and s j simultaneously;
[0022] Construct a comorbidity network for describing the complex influence relationships of chronic diseases.
[0023] Preferably, the method for representing the comorbidity network using a low-rank representation method with fused graph convolution constraints includes:
[0024] Design an objective function, and the objective function is as follows:
[0025]
[0026]
[0027] U, H, W (1) , W (2) ≥0
[0028] where A is the adjacency matrix of the comorbidity network, U is a feature matrix of size n×K, and n and K are the number of nodes and communities in the comorbidity network respectively; is used to calculate the F-norm of the matrix; Tr(·) is used to calculate the trace of the matrix; U T represents the transpose of matrix U; L = D - A is the Laplacian matrix of A, where D = ∑ l A il is the degree matrix of A, where A il is the element value at the corresponding subscript in A; represents the feature matrix of the hidden layer of the GCN module, where h i corresponds to the eigenvector of node v i , and d represents the feature dimension; and represent the weight matrices of the GCN module; is the normalized adjacency matrix, where is the adjacency matrix with self-connections added, is 's degree matrix, is 's element value at the corresponding subscript; λ1 is the graph regularization coefficient; σ(·) is the non-linear activation function of the GCN module;
[0029] Convert the graph convolution constraint into the following regularization form:
[0030]
[0031] s.t. U, H, W (1) , W (2) ≥ 0
[0032] Solve the following optimization problem to obtain the parameter matrices U, H, and W (1) and W (2) :
[0033] (U, H, W (1) , W (2) ) ← argmin J GCGS , s.t. U, H, W (1) , W (2) ≥ 0
[0034] Combine the Lagrangian function method and the KKT conditions of the inequality constraint to derive the learning rule of the optimization problem:
[0035]
[0036]
[0037]
[0038]
[0039] where λ1 is the graph regularization coefficient, λ2 is the graph convolution regularization coefficient, g1 and g1′ (the derivative of g1) respectively represent and g2 and g2′ (the derivative of g2) respectively represent and σ′(·) is the derivative of σ(·); β is the linear adjustment coefficient that controls the learning scale factor of the feature factor matrix U and takes values in the interval (0, 1); ⊙ represents the Hadamard product of matrices; Iteratively update the feature and weight matrices U, H, and W respectively according to the learning rule (1) and W (2) , until the convergence condition is reached and stopped, and finally obtain the feature matrices U, H, and the weight W (1) , W (2) as the solution of the optimization problem, and output the feature matrix U as the indication matrix for community division.
[0040] Preferably, the use of the community division mechanism to identify the community structure in the comorbidity network to discover the comorbidity pattern includes:
[0041] After obtaining the feature parameters U and H, using U as the indication matrix for dividing the node-disease comorbidity module attribution relationship, judge each node in the comorbidity network one by one and assign it to the corresponding comorbidity module. The assignment rules are as follows:
[0042]
[0043] Among them, V represents the set of nodes in the comorbidity network, and v i represents any one of them; C s is the comorbidity combination s, s ∈ {1, 2,..., K}, and K is the total number of comorbidity combinations; u ik is an element in the low-rank matrix U, indicating the probability that the node v i is assigned to the comorbidity combination k; if u is is the maximum value among them, then the node v i has the greatest possibility of being assigned to the comorbidity combination s. After dividing the comorbidity combinations of all nodes in the target comorbidity network, all possible identified comorbidity combinations C = {C1, C2,..., C K} can be obtained.
[0044] In a second aspect, an embodiment of the present application provides a device for identifying chronic disease comorbidity patterns based on relaxed constraint symmetric low-rank representation, which is characterized by including:
[0045] An acquisition unit for acquiring medical data resources, where the medical data resources include data records of chronic disease prevalence conditions sorted out from inspections, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources;
[0046] A construction unit for constructing a chronic disease comorbidity network using medical data resources, where the chronic disease comorbidity network is used to describe the correlation and influence between diseases and can be constructed by calculating the correlation between diseases or patients based on multi-source data. The multi-source data includes patient basic information and inspection, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data. Among them, the chronic disease comorbidity network is represented by a graph G = (V, E), where V = {v i |i ∈ {1,..., n}} represents a set containing n disease nodes, and E = {e ij |i, j ∈ {1,..., m}} represents a set containing m edges; the adjacency matrix A = [a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between the disease nodes v i and v j , and its value is equal to I(s i , s j );
[0047] A characterization unit for performing low-rank characterization learning on the chronic disease comorbidity network to mine and output the community structure therein, where the low-rank characterization learning uses a low-rank characterization method with fused graph convolution constraints to characterize the comorbidity network;
[0048] A partitioning unit for using a community partitioning mechanism to identify the community structure in the comorbidity network and discover comorbidity patterns.
[0049] In a third aspect, an embodiment of the present application provides an electronic device, which is characterized by including:
[0050] A processor;
[0051] A memory for storing executable instructions executable by the processor;
[0052] The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the above method steps.
[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium stores a computer program for executing the above method.
[0054] Compared with the prior art, the present invention has the following advantages: (1) Compared with the existing non-negative matrix factorization algorithm, the present invention not only has the good interpretable clustering characteristics of the low-rank factorization model, but also integrates the non-linear and non-Euclidean characterization learning ability of GCN, improving its representation learning ability; (2) The GCN module can learn the global structural features of the network, and by introducing the graph regularization technology, the inherent geometric structure characteristics of the network are effectively maintained. Therefore, the technical solution of the present invention can theoretically obtain better global non-linear and local feature representation capabilities; (3) The designed model's parameter solution under non-negative constraints is realized by adopting an optimization scheme based on NMU, ensuring the interpretability of the model for the non-negative numerical characteristics of the comorbidity network. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] By referring to the following drawings, the exemplary embodiments of the present invention can be more completely understood. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0056] Figure 1 It is a flowchart of a method for identifying chronic disease comorbidity patterns with enhanced low-rank characterization by fusing graph convolution constraints according to an exemplary embodiment of the present application;
[0057] Figure 2Schematic diagram of a chronic disease comorbidity pattern recognition device that fuses graph convolutional constraints to enhance low-rank representation according to an exemplary embodiment of the present application. Detailed implementation manners
[0058] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0059] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0060] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0061] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0062] The embodiments of the present application provide a chronic disease comorbidity pattern recognition method based on relaxed constraint symmetric low-rank representation, which will be described below with reference to the accompanying drawings.
[0063] Refer to Figure 1 , which shows a chronic disease comorbidity pattern recognition method based on relaxed constraint symmetric low-rank representation provided by some embodiments of the present application. As Figure 1 shown, the method may include the following content:
[0064] S101: Obtain medical data resources, where the medical data resources include data records of chronic disease prevalence sorted out from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources;
[0065] S102: Construct a chronic disease comorbidity network using the medical data resources, where the chronic disease comorbidity network is used to describe the correlation and influence between diseases and can be constructed by calculating the correlation between diseases or patients based on multi-source data, and the multi-source data includes patient basic information and test, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data;
[0066] Specifically, the chronic disease comorbidity network is represented by a graph G=(V, E), where V={v i |i∈{1,...,n}} represents a set containing n disease nodes, and E={e ij |i,j∈{1,...,m}} represents a set containing m edges; the adjacency matrix A=[a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between disease nodes v i and v j , and its value is equal to I(s i , s j );
[0067] S103: Perform low-rank representation learning on the chronic disease comorbidity network to mine and output the community structure therein, where the low-rank representation learning uses a low-rank representation method that combines graph convolutional constraints to represent the comorbidity network;
[0068] S104: Use the community partitioning mechanism to identify the community structure in the comorbidity network and discover the comorbidity patterns.
[0069] Specifically, the construction method of the chronic disease comorbidity network includes the following steps:
[0070] (1) Sort out the chronic disease records and sort out the data records from the test, diagnosis, treatment, prescription, electronic medical records, and clinical data;
[0071] (2) Count the frequency of disease occurrence, and count the frequency Q(s i ) of each disease occurrence and the frequency Q(s i , s j ) of the co-occurrence of disease combinations from the data records, where if a certain disease s i appears in a disease record, it is counted once, and if two certain diseases s i and s j co-appear in a disease record, it is counted once, where i,j = 1,2,...,n represent the number of diseases of concern;
[0072] Calculate the mutual information value between two diseases pairwise, and the calculation method is as follows:
[0073]
[0074] Among them, p(s i ) = Q(s i ) / F, p(s j ) = Q(s j ) / F are used to calculate the observed frequencies of diseases s i and s j respectively, where F represents the total number of observations. p(s i , s j ) = Q(s i , s j ) / F is used to calculate the frequency of simultaneously observing diseases s i and s j ;
[0075] (3) Construct a comorbidity network for describing the complex influence relationships of chronic diseases.
[0076] Specifically, the method of representing the comorbidity network by using a low-rank representation method with fused graph convolutional constraints includes the following steps:
[0077] (1) Design an objective function, and the objective function is as follows:
[0078]
[0079]
[0080] U, H, W (1) , W (2) ≥0
[0081] Among them, A is the adjacency matrix of the comorbidity network, U is a feature matrix of size n×K, and n and K are the numbers of nodes and communities in the comorbidity network respectively; is used to calculate the Frobenius norm of the matrix; Tr(·) is used to calculate the trace of the matrix; U T represents the transpose of matrix U; L = D - A is the Laplacian matrix of A, where D = ∑ l A il is the degree matrix of A, where A il is the element value at the corresponding subscript in A; represents the feature matrix of the hidden layer of the GCN module, where h i corresponds to the eigenvector of node v i , and d represents the feature dimension; and represent the weight matrices of the GCN module; is the normalized adjacency matrix, where is the adjacency matrix with self-connections added, is 's degree matrix, is The element value at the corresponding subscript; λ1 is the graph regularization coefficient; σ(·) is the non-linear activation function of the GCN module;
[0082] (2) Transform the graph convolution constraint into the following regularization form:
[0083]
[0084] s.t. U, H, W (1) , W (2) ≥0
[0085] Solve the following optimization problem to obtain the parameter matrices U, H, W (1) and W (2) :
[0086] (U, H, W (1) , W (2) ) ← argmin J GCGS , s.t. U, H, W (1) , W (2) ≥0
[0087] (3) Combine the Lagrangian function method and the KKT conditions of the inequality constraint to derive the learning rule of the optimization problem:
[0088]
[0089]
[0090]
[0091]
[0092] where λ1 is the graph regularization coefficient, λ2 is the graph convolution regularization coefficient, g1 and g1′ (the derivative of g1) respectively represent and g2 and g2′ (the derivative of g2) respectively represent and σ′(·) is the derivative of σ(·); β is the linear adjustment coefficient that controls the learning scale factor of the feature factor matrix U and takes values in the interval (0, 1); ⊙ represents the Hadamard product of matrices; Update the feature and weight matrices U, H, W respectively according to the learning rule (1) and W (2) , until the convergence condition is reached and stop, and finally obtain the feature matrices U, H and the weights W (1) , W (2) as the solution of the optimization problem, and output the feature matrix U as the indication matrix for community division.
[0093] Specifically, the application of the community division mechanism to identify the community structure in the comorbidity network to discover comorbidity patterns includes:
[0094] After obtaining the feature parameters U and H through training, using U as the indication matrix for the division of the node-disease comorbidity module attribution relationship, each node in the comorbidity network is judged one by one and assigned to the corresponding comorbidity module. The assignment rules are as follows:
[0095]
[0096] Among them, V represents the set of nodes in the comorbidity network, and v i represents any one of them; C s is the comorbidity combination s, s ∈ {1, 2,..., K}, and K is the total number of comorbidity combinations; u ik is an element in the low-rank matrix U, representing the probability that the node v i is assigned to the comorbidity combination k; if u is is the maximum value among them, then the node v i has the greatest possibility of being assigned to the comorbidity combination s. After dividing the comorbidity combinations of all nodes in the target comorbidity network, all possible identified comorbidity combinations C = {C1, C2,..., C K} can be obtained.
[0097] Compared with the prior art, the present invention has the following advantages: (1) Compared with the existing non-negative matrix factorization algorithm, the present invention not only has the good interpretable clustering characteristics of the low-rank decomposition model, but also incorporates the non-linear and non-Euclidean representation learning ability of GCN, improving its representation learning ability. (2) The GCN module can learn the global structural features of the network, and by introducing the graph regularization technology, the inherent geometric structure characteristics of the network are effectively maintained. Therefore, the technical solution of the present invention can theoretically obtain better global non-linear and local feature representation capabilities. (3) The optimization scheme based on NMU is adopted to realize the parameter solution of the designed model under non-negative constraints, ensuring the interpretability of the model for the non-negative numerical characteristics of the comorbidity network.
[0098] In the above embodiment, a method is provided. Correspondingly, the present application also provides a device. The device provided in the embodiments of the present application can implement the above method, and the device can be implemented in a software, hardware, or a combination of software and hardware manner. For example, the device can include integrated or separate functional modules or units to execute the corresponding steps in the above methods.
[0099] In some implementation manners of the embodiments of the present application, the device provided in the embodiments of the present application and the method provided in the foregoing embodiments of the present application are based on the same inventive concept and have the same beneficial effects.
[0100] As Figure 2 shown, the device 20 may include:
[0101] An acquisition unit 201, configured to acquire medical data resources, where the medical data resources include data records of chronic disease prevalence sorted out from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources;
[0102] A construction unit 202, configured to construct a chronic disease comorbidity network by using the medical data resources, where the chronic disease comorbidity network is used to describe the correlation and influence between diseases and can be constructed by calculating the correlation between diseases or patients based on multi-source data, and the multi-source data includes patient basic information and examination, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data;
[0103] Among them, the chronic disease comorbidity network is represented by a graph G=(V, E), where V={v i |i∈{1,…,n}} represents a set containing n disease nodes, and E={e ij |i,j∈{1,…,m}} represents a set containing m edges; the adjacency matrix A=[a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between disease nodes v i and v j , and its value is equal to I(s i ,s j );
[0104] A characterization unit, configured to perform low-rank characterization learning on the chronic disease comorbidity network to mine and output the community structure therein, where the low-rank characterization learning uses a low-rank characterization method that fuses graph convolution constraints to characterize the comorbidity network;
[0105] A partitioning unit, configured to identify the community structure in the comorbidity network by using a community partitioning mechanism and discover comorbidity patterns.
[0106] An embodiment of the present application also provides an electronic device corresponding to the method provided in the foregoing embodiment. The device can be an electronic device for a server, such as a server, including an independent server and a distributed server cluster, etc., to execute the above method; the electronic device can also be an electronic device for a client to execute the above method.
[0107] The electronic device includes: a processor, a memory, a bus, and a communication interface. The processor, the communication interface, and the memory are connected through the bus; a computer program that can run on the processor is stored in the memory, and when the processor runs the computer program, it executes the method described above in the present application.
[0108] Among them, the memory may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0109] The bus can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory is used to store programs, and after receiving the execution instruction, the processor executes the program. Any implementation manner of the method disclosed in any embodiment of the present application can be applied to the processor or implemented by the processor.
[0110] The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly reflected as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. The electronic device provided by the embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0111] The embodiment of the present application also provides a computer-readable medium corresponding to the method provided by the foregoing embodiment, on which a computer program (i.e., a program product) is stored, and when the computer program is run by the processor, it will execute the foregoing method.
[0112] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated herein one by one.
[0113] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0115] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0116] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.
[0118] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0119] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of each embodiment of the present application, and they should all be covered by the scope of the claims and the description of the present application.
Claims
1. A method for identifying chronic disease comorbidity patterns by fusing graph convolutional constraints to enhance low-rank representation, characterized in that, including: Obtain medical data resources, where the medical data resources include data records of the prevalence of chronic diseases sorted out from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources; Construct a chronic disease comorbidity network using medical data resources. Among them, the chronic disease comorbidity network is used to describe the correlation and influence between diseases and can be constructed by calculating the correlation between diseases or patients based on multi-source data. The multi-source data includes patient basic information and test, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data. Among them, the chronic disease comorbidity network is represented by a graph G=(V, E), where V={v i |i∈{1,…,n}} represents a set containing n disease nodes, and E={e ij |i,j∈{1,…,m}} represents a set containing m edges; the adjacency matrix A=[a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between disease nodes v i and v j , and its value is equal to I(s i , s j ), where s i and s j represent two different diseases respectively; Perform low-rank representation learning on the chronic disease comorbidity network to mine and output the community structure therein, where the low-rank representation learning uses a low-rank representation method with fused graph convolution constraints to represent the comorbidity network; the use of the low-rank representation method with fused graph convolution constraints to represent the comorbidity network includes: Design an objective function, and the objective function is as follows: Among them, A is the adjacency matrix of the comorbidity network, U is the feature matrix of size n×K, where n and K are the numbers of nodes and communities in the comorbidity network, respectively; is to calculate the F-norm of the matrix; Tr(·) is to calculate the trace of the matrix; U T represents the transpose of the matrix U; L = D - A is the Laplacian matrix of A, where D = ∑ l A il is the degree matrix of A, where A il is the element value at the corresponding subscript in A; represents the feature matrix of the hidden layer of the GCN module, where h i corresponds to the eigenvector of node v i , and d represents the feature dimension; and represent the weight matrix of the GCN module; is the normalized adjacency matrix, is the adjacency matrix with self-connections added, is 's degree matrix, is 's element value at the corresponding subscript; λ1 is the graph regularization coefficient; σ(·) is the non-linear activation function of the GCN module; Use a community division mechanism to identify the community structure in the comorbidity network and discover comorbidity patterns.
2. The method according to claim 1, characterized in that, Convert the graph convolution constraint into the following regularization form: Solve the following optimization problem to obtain the feature matrices U, H, and the weight matrix W (1) , W (2) : (U, H, W (1) , W (2) ) ← arg min J GCGS , s.t. U, H, W (1) , W (2) ≥ 0 Combine the Lagrangian function method and the KKT conditions of inequality constraints to derive the learning rules for the optimization problem: Among them, λ1 is the graph regularization coefficient, λ2 is the graph convolution regularization coefficient, g1 and g1′ respectively represent and g2 and g2′ respectively represent and σ′(·) is the derivative of σ(·); β is a linear adjustment coefficient for controlling the learning scale factor of the feature matrix U, taking values in the interval (0, 1); ⊙ represents the Hadamard product of matrices; the feature matrices U, H, and the weight matrix W are iteratively updated respectively according to the learning rules (1) and W (2) , and stop until the convergence condition is reached, and the finally obtained feature matrices U, H, and the weight matrix W (1) , W (2) are used as the solutions to the optimization problem, and the output feature matrix U is used as the indication matrix for community division.
3. The method according to claim 1, characterized in that The method for constructing the chronic disease comorbidity network includes: Sort out the chronic disease medical records and sort out the data records from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data; Count the frequency of occurrence of diseases, and count the frequency Q(s i ) of each disease occurrence and the frequency Q(s i , s j ) of co-occurrence of disease combinations from the data records, where a disease s i appearing in a disease record is counted once, and two diseases s i and s j co-appearing in a disease record are counted once; Calculate the mutual information value between every two diseases, and the calculation method is as follows: where p(s i ) = Q(s i ) / F, and p(s j ) = Q(s j ) / F are used to calculate the observed frequencies of diseases s i and s j respectively, where F represents the total number of observations. p(s i , s j ) = Q(s i , s j ) / F is used to calculate the frequency of observing both diseases s i and s j ; Construct a comorbidity network for describing the complex influence relationship of chronic diseases.
4. The method according to claim 1, wherein The use of the community division mechanism to identify the community structure in the comorbidity network to discover comorbidity patterns includes: After training to obtain the feature matrices U and H, use U as the indication matrix for the division of the node-disease comorbidity module attribution relationship, and judge each node in the comorbidity network one by one and assign it to the affiliated comorbidity module. The assignment rules are as follows: Among them, V represents the set of nodes in the comorbidity network, and v i represents any one of the nodes; C s is a comorbidity combination, s ∈ {1, 2, …, P}, where P is the total number of comorbidity combinations; u ik is an element in the feature matrix U, representing the probability that the node v i is assigned to the comorbidity combination C k ; if u is is the maximum value among them, then the node v i has the highest possibility of being assigned to the comorbidity combination C s . After assigning the comorbidity combinations of all nodes in the target comorbidity network, all identified comorbidity combinations {C1, C2, …, C P} can be obtained.
5. A chronic disease comorbidity pattern recognition device that fuses graph convolutional constraints to enhance low-rank representation, characterized in that, including: An acquisition unit for obtaining medical data resources, where the medical data resources include data records of the prevalence of chronic diseases sorted out from examinations, diagnoses, treatments, prescriptions, electronic medical records, and clinical data resources; Building unit, constructing a chronic disease comorbidity network by utilizing medical data resources, wherein the chronic disease comorbidity network is used to describe the correlation and influence between diseases and can be constructed by calculating the correlation between diseases or patients based on multi-source data, and the multi-source data includes patient basic information and test, diagnosis, treatment, prescription data, electronic medical records, clinical data, and omics big data. Among them, the chronic disease comorbidity network is represented by a graph G=(V, E), where V = {v i | i ∈ {1, …, n}} represents a set containing n disease nodes, and E = {e ij | i, j ∈ {1, …, m}} represents a set containing m edges; the adjacency matrix A = [a ij is used to represent and store the target comorbidity network G, where a ij represents the interaction relationship between disease nodes v i and v j , and its value is equal to I(s i , s j ), where s i and s j represent two different diseases respectively; A representation unit for performing low-rank representation learning on the chronic disease comorbidity network to mine and output the community structure therein, where the low-rank representation learning uses a low-rank representation method with fused graph convolution constraints to represent the comorbidity network; the use of the low-rank representation method with fused graph convolution constraints to represent the comorbidity network includes: Design an objective function, and the objective function is as follows: Among them, A is the adjacency matrix of the comorbidity network, and U is the feature matrix of size n×K, where n and K are the numbers of nodes and communities in the comorbidity network, respectively; is to calculate the F-norm of the matrix; Tr(·) is to calculate the trace of the matrix; U T represents the transpose of the matrix U; L = D - A is the Laplacian matrix of A, where D = ∑ l A il is the degree matrix of A, where A il is the element value at the corresponding subscript in A; represents the feature matrix of the hidden layer of the GCN module, where h i corresponds to the eigenvector of node v i and d represents the feature dimension; and represent the weight matrix of the GCN module; is the normalized adjacency matrix, where is the adjacency matrix with self-connections added, is 's degree matrix, is 's element value at the corresponding subscript; λ1 is the graph regularization coefficient; σ(·) is the non-linear activation function of the GCN module; A division unit for using a community division mechanism to identify the community structure in the comorbidity network and discover comorbidity patterns.
6. The device according to claim 5, characterized in that Convert the graph convolution constraint into the following regularization form: Solve the following optimization problem to obtain the feature matrices U, H, and the weight matrix W (1) , W (2) : (U, H, W (1) , W (2) ) ← arg min J GCGS , s.t. U, H, W (1) , W (2) ≥ 0 Combine the Lagrangian function method and the KKT conditions of inequality constraints to derive the learning rules for the optimization problem: Among them, λ1 is the graph regularization coefficient, λ2 is the graph convolution regularization coefficient, g1 and g1′ respectively represent and g2 and g2′ respectively represent and σ′(·) is the derivative of σ(·); β is a linear adjustment coefficient for controlling the learning scale factor of the feature matrix U, taking values in the interval (0,1); ⊙ represents the Hadamard product of matrices; the feature matrices U, H, and the weight matrix W are iteratively updated according to the learning rules respectively (1) and W (2) , until the convergence condition is reached and then stop, and use the finally obtained feature matrices U, H, and the weight matrix W (1) , W (2) as the solution of the optimization problem, and output the feature matrix U as the indication matrix for community division.
7. An electronic device, characterized in that, including: A processor; A memory for storing the executable instructions that can be executed by the processor; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method according to any one of claims 1 to 4 above.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 4 above.
Citation Information
Patent Citations
Functional module identification method and device, terminal equipment and storage medium
CN118072833A