Digestive system disease-based risk prediction method and device

By combining cross-modal attention mechanisms and bidirectional interaction with graph convolutional networks with SHAP value analysis, the problem of feature fusion in risk prediction of gastroenterological diseases is solved, improving prediction accuracy and interpretability, and assisting clinical decision-making.

CN122050796APending Publication Date: 2026-05-15ZHENGZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511971567.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and integrate multiple features in predicting the risk of gastrointestinal diseases, particularly the interaction between clinical data and gut microbiome data, leading to insufficient predictive accuracy.

Method used

A cross-modal attention mechanism was used to perform bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix. Combined with graph convolutional networks and SHAP value analysis, fused features were obtained and risk prediction was performed.

Benefits of technology

It improves the accuracy and interpretability of risk prediction for gastroenterological diseases, assists physicians in identifying the impact of key clinical indicators and gut microbiota characteristics on risk, and supports clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a risk prediction method and device based on digestive system department diseases, and the method comprises the steps: obtaining clinical data and intestinal microbiome data of a target object, and carrying out the preprocessing and standardization of a clinical data set and the intestinal microbiome data; inputting the preprocessed clinical data into a first encoder to obtain a clinical feature matrix; inputting the preprocessed intestinal microbiome into a second encoder to obtain an intestinal microbiome characteristic matrix; based on a cross-modal attention mechanism, performing bidirectional attention calculation on the clinical feature matrix and the intestinal microbiome feature matrix to obtain fusion features; the method comprises the following steps: performing two-way interaction and fusion on clinical data and microbiome data of a target object, calculating an SHAP value of a fusion feature, taking the fusion feature as an input of a preset risk prediction module, and obtaining a risk prediction value of a disease of the digestive system department, and obtaining the risk prediction value of the disease of the digestive system department through two-way interaction and fusion of the clinical data and the microbiome data of the target object. The method can assist doctors in interpreting which key clinical indexes or intestinal microorganism characteristics have relatively large multi-risk prediction influence, assist clinical decision, and further improve the interpretability of digestive system department disease risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical data analysis technology, specifically a risk prediction method and device based on digestive diseases. Background Technology

[0002] The pathogenesis of gastroenterological diseases is often not the result of a single factor, but rather the result of the interaction and combined influence of multiple internal and external factors. Due to the complexity of the etiology, the characteristic dimensions of gastroenterological diseases are correspondingly high. These characteristics may include multiple aspects such as clinical data and microbial community analysis. Each aspect may contain multiple specific indicators or parameters, making the analysis of disease characteristics exceptionally complex. Among the numerous characteristics, accurately identifying the effective characteristics that play a key role in predicting the risk of gastroenterological diseases is an urgent problem to be solved. Some characteristics often have complex interaction relationships, which may directly affect the accuracy of disease prediction. Therefore, considering the interactions between characteristics in risk prediction models is crucial. In summary, current risk prediction for gastroenterological diseases faces multiple challenges, making it a complex and arduous task.

[0003] Therefore, overcoming the aforementioned technical problems and defects has become a key issue that needs to be addressed. Summary of the Invention

[0004] To overcome the aforementioned problems in the prior art, this application provides a risk prediction method and device based on gastrointestinal diseases, employing the following technical solution: Firstly, this application provides a risk prediction method based on gastrointestinal diseases, including: Acquire clinical data and gut microbiome data of the target subjects, and preprocess and standardize the clinical dataset and gut microbiome data respectively; The preprocessed clinical data is input into the first encoder to obtain the clinical feature matrix; The preprocessed gut microbiome is input into the second encoder to obtain the gut microbiome feature matrix; Based on the cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features; Calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterology diseases.

[0005] Further, the preprocessed gut microbiome is input into the second encoder to obtain the gut microbiome feature matrix, including: constructing an ecological network of the microbial community based on correlation analysis between microbial species; for each target object's gut microbiome sample, based on the constructed microbial ecological network, mapping microbial abundance information onto network nodes to form a node feature matrix, preserving the network's adjacency matrix, and representing the sample as a graph data structure containing the node feature matrix and adjacency matrix; reconstructing the input graph data structure based on the microbial ecological network topology to form graph structure data with multiple channels; performing parallel computation on the graph structure data of multiple channels using multi-scale convolution kernels in the convolutional layer to obtain multi-level and multi-dimensional features of the graph structure data of multiple channels, and obtaining multi-scale convolution features; further fusing the relationship information between multi-level and multi-dimensional feature nodes through GCN graph convolution to obtain graph convolution features; and performing a weighted average fusion of the multi-scale convolution features and graph convolution features to obtain the gut microbiome feature matrix.

[0006] Furthermore, based on the cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fusion features, including: obtaining microbiome fusion features guided by the clinical feature matrix, i.e., the first fusion feature; obtaining clinical fusion features guided by the microbiome features, i.e., the second fusion feature; and fusing the first fusion feature and the second fusion feature obtained by bidirectional attention using a weighted summation method to obtain the final fusion feature.

[0007] Secondly, this application also provides a risk prediction device based on gastrointestinal diseases, comprising: The data acquisition module is used to acquire clinical data and gut microbiome data of the target subjects, and to preprocess and standardize the clinical dataset and gut microbiome data respectively. The clinical feature acquisition module is used to input the preprocessed clinical data into the first encoder to obtain the clinical feature matrix. The gut microbiome feature acquisition module is used to input the preprocessed gut microbiome into the second encoder to obtain the gut microbiome feature matrix. The fusion feature acquisition module is used to perform bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix based on the cross-modal attention mechanism to obtain fusion features; The risk prediction value acquisition module is used to calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterological diseases.

[0008] Thirdly, this application provides an electronic device, comprising: One or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to perform the method as described in the first aspect.

[0009] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0010] Fifthly, this application provides a computer program that, when executed by a computer, performs the method described in the first aspect.

[0011] In one possible design, the program in the fifth aspect can be stored wholly or partially on a storage medium packaged with the processor, or it can be stored wholly or partially on a memory not packaged with the processor.

[0012] This application has the following beneficial effects: 1. This application obtains clinical data and gut microbiome data of the target subjects, and preprocesses and standardizes the clinical dataset and gut microbiome data respectively; the preprocessed clinical data is input into a first encoder to obtain a clinical feature matrix; the preprocessed gut microbiome data is input into a second encoder to obtain a gut microbiome feature matrix; based on a cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features. This application achieves bidirectional deep interaction between clinical and microbiome features through a bidirectional attention mechanism, and has dynamic weight allocation capabilities based on cross-modal query, key, and value information transmission. The bidirectional interaction and fusion of clinical and microbiome features is achieved through the cross-modal attention mechanism, effectively extracting complementary information from the two modalities.

[0013] 2. This application calculates the SHAP value of fused features and uses these fused features as input to a preset risk prediction module to obtain risk prediction values ​​for gastroenterological diseases. By combining SHAP value analysis, this application can assist physicians in interpreting which key clinical indicators or gut microbiota characteristics have a greater impact on risk prediction, thus aiding clinical decision-making and further improving the interpretability of risk prediction for gastroenterological diseases. Attached Figure Description

[0014] Figure 1 This is an exemplary system architecture diagram to which embodiments of this application can be applied; Figure 2 This is a flowchart illustrating a risk prediction method for gastroenterological diseases, as described in an embodiment of this application. Figure 3 This is a flowchart illustrating the process of obtaining gut microbiome features according to an embodiment of this application. Figure 4 This is a flowchart illustrating the process of obtaining fusion features according to an embodiment of this application. Figure 5 This is a system flowchart of an embodiment of this application; Figure 6 This is a schematic diagram of a computer device according to an embodiment of this application. Detailed Implementation

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0016] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0018] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0019] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0020] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.

[0021] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.

[0022] It should be noted that the risk prediction method based on gastroenterological diseases provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the risk prediction device based on gastroenterological diseases is generally set in the server / terminal device.

[0023] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0024] Reference Figure 2 , Figure 2 This disclosure provides an optional flowchart of a risk prediction method based on gastrointestinal diseases. This risk prediction method can be executed by a terminal, by a server, or by a combination of both. The method includes, but is not limited to, steps S1 to S5. The target audience in this application refers to gastrointestinal patients.

[0025] Step S1: Obtain clinical data and gut microbiome data of the target subjects, and preprocess and standardize the clinical dataset and gut microbiome data respectively.

[0026] In this embodiment, the clinical data includes basic information of the target subject, clinical diagnosis and treatment data, laboratory test data, and imaging data. Preprocessing and noise removal are performed on the clinical data and gut microbiome data of the target subject, which will not be elaborated here. The preprocessed data is then obtained.

[0027] It should be noted that the gut microbiome data is mainly composed of bacteria, but also includes archaea, fungi, viruses, and other microorganisms from multiple kingdoms.

[0028] It should be noted that the clinical data of patients are obtained based on the hospital's physical examination report, while the gut microbiome data is obtained based on the hospital's bioinformatics analysis of patient samples.

[0029] Step S2: Input the preprocessed clinical data into the first encoder to obtain the clinical feature matrix. The first encoder corrects the weight distribution of the disease association strength score based on the spatial constraints of the shape description operator and the offset direction of the abnormal offset vector.

[0030] In this embodiment, before inputting the preprocessed clinical data into the first encoder, the process includes: acquiring anatomical features of intestinal / abdominal images through imaging data and converting them into shape description operators; calculating the deviation values ​​of biochemical indicators in laboratory test data from their corresponding normal reference ranges to obtain abnormal offset vectors; and extracting keywords from historical diagnoses in clinical treatment data and converting them into disease association strength scores through a preset medical knowledge graph.

[0031] In this embodiment, the shape description operator is a geometric deformation parameter calculated from image data using keywords in clinical diagnostic data.

[0032] In this embodiment, anatomical features of intestinal / abdominal images are obtained through imaging data and converted into shape descriptive operators. This includes: semantic detection of text fields in clinical diagnostic data; when a keyword is detected, retrieving image data corresponding to the keyword from the images; performing geometric analysis on the image data to obtain geometric deformation parameters reflecting tissue morphology; and using these geometric deformation parameters as shape descriptive operators. Keywords include keywords related to gastrointestinal diseases, pathological descriptions, or structural abnormality trigger words.

[0033] In this embodiment of the application, geometric analysis is performed on the image data to obtain geometric deformation parameters reflecting tissue morphology, including: The retrieved image data is used to locate the target area. Based on the structural labels in the image data that are associated with keyword triggers, the target area to be analyzed is determined in the image.

[0034] Geometric features are extracted from the target region to obtain its spatial structure. It should be noted that the spatial feature extraction process includes extracting the boundary contours of the target organization, calculating the spatial distribution features within the region, and constructing a geometric representation of the target organization; the boundary contours are used to characterize the external structure, and the spatial distribution features are used to reflect the overall layout of the organizational form.

[0035] The geometric differences between the quantified spatial structure and the preset reference structure are used to obtain geometric deformation parameters that reflect changes in tissue morphology. These parameters include spatial offset parameters describing contour displacement, deformation amplitude parameters describing the degree of morphological expansion, contraction, or distortion, and blank direction parameters describing the directionality of morphological changes. It should be noted that the preset reference structure can be constructed based on historical images of the same target object, using standard structural templates for similar populations, or symmetrical or adjacent normal structural regions within the same image.

[0036] This application incorporates spatial constraints into the first encoder, using spatial constraint information from image data to participate in the encoding of clinical features, thereby achieving the fusion of image morphology information and clinical data, providing spatial coordinates of lesions for biochemical indicator analysis, and enabling the first encoder to focus on the characteristics of physically damaged areas.

[0037] In this embodiment of the application, the deviation value between the biochemical indicators in laboratory test data and the corresponding normal reference range is calculated to obtain the abnormal offset vector, including: The deviation direction and magnitude of various biochemical indicators in laboratory test data from their corresponding normal reference values ​​are calculated. Based on the comprehensive expression of the deviation direction and magnitude, an abnormal offset vector reflecting the degree of abnormality of the biochemical indicators is obtained.

[0038] The process for determining the direction of deviation is as follows: when the detected value is higher than the upper limit of the normal reference range, the deviation direction is determined to be positive; when the detected value is lower than the upper limit of the normal reference range, the deviation direction is determined to be negative; when the detected value is within the normal reference range, the deviation direction is determined to be no deviation. The deviation direction is used to represent the deviation trend of biochemical indicators relative to the normal physiological state.

[0039] The process of determining the offset amplitude is as follows: by quantifying the difference between the detected value and the boundary of the normal reference value interval, and normalizing the difference in combination with the scale of the reference value interval, the offset amplitude of different indicators is obtained.

[0040] The comprehensive expression of deviation direction and deviation magnitude includes: taking the deviation direction as a direction factor and the deviation magnitude as a offset intensity factor, and obtaining a deviation representation with directional attributes by combining the deviation direction and deviation magnitude.

[0041] In this embodiment of the application, keywords from historical diagnoses in clinical treatment data are extracted and converted into disease association strength scores using a preset medical knowledge graph, including: The diagnostic text in clinical diagnosis and treatment data is parsed to extract keyword information related to the disease state. The keywords are input into a preset medical knowledge graph. By mapping the relationship between keywords and disease nodes, a corresponding disease association strength score is generated to quantify the degree of association between different diseases and the current clinical state.

[0042] In this embodiment, preprocessed clinical data is input into a first encoder to obtain a clinical feature matrix. The first encoder, based on the spatial constraints of the shape description operator, corrects the weight distribution of the disease association strength score according to the offset direction of the abnormal offset vector. Specifically, this includes: The preprocessed shape descriptor, anomaly offset vector, and disease association strength score are input into the first encoder, which maps the shape descriptor to a spatial attention mask. Specifically, geometric deformation parameters are mapped to a two-dimensional coordinate matrix. In this matrix, regions with severe anatomical distortion are assigned high weight values, while regions with normal anatomical structures are assigned lower weight values. A Gaussian kernel function is used to smooth the deformation points, generating a continuous spatial weight field. This application uses a spatial attention mask to define key regions for feature extraction, concentrating the resources of the first encoder on abnormal lesion regions, thus improving the effectiveness of feature extraction.

[0043] The disease association strength score is adjusted based on the offset direction of the abnormal offset vector. Logical judgment is performed on the offset direction of the abnormal offset vector in relation to a certain biochemical indicator. Based on the result of this logical judgment, the weight coefficients of the corresponding scoring path are adjusted. For example, when the offset direction of the abnormal offset vector is positive in relation to a certain biochemical indicator, the logical judgment is performed: if the current disease association strength score corresponds to a validation disease, the weight coefficients of the corresponding association strength scoring path are adjusted upwards; when the offset direction of the abnormal offset vector is negative in relation to a certain biochemical indicator, the weight coefficients of the corresponding association strength scoring path are adjusted downwards.

[0044] The clinical data is filtered by element-wise multiplication of the spatial attention mask with the disease association strength score and offset magnitude adjusted by weight coefficients. This element-wise multiplication process strengthens abnormal anatomical regions, offset indicators, and logically related information while weakening information that does not meet these criteria. Linear projection is then used to perform uniform scale mapping on the filtered data, completing feature alignment and fusion to generate a clinical feature matrix. This clinical feature matrix encodes abnormal information in anatomical structures, offsets in laboratory test indicators, and the strength of associations between diseases.

[0045] Step S3: Input the preprocessed gut microbiome into the second encoder to obtain the gut microbiome feature matrix.

[0046] In this embodiment, the preprocessed gut microbiome is input into the second encoder to obtain the gut microbiome feature matrix. Please refer to [reference needed]. Figure 3 The specific content includes: Step 31: Based on correlation analysis among microbial species, construct an ecological network of the microbial community. Microbial species are used as network nodes, and the interactions between species are used as edges. The weight of each edge reflects the degree of connection between species. Simultaneously, combining known microbial functional information and metabolic pathway information, functional attribute labels are assigned to network nodes, constructing a composite network that includes both topological structure and functional attributes. In this embodiment, Spearman correlation coefficient and mutual information can be used for correlation analysis among microbial species.

[0047] Step 32: For each target subject's gut microbiome sample, based on the constructed microbial community ecological network, microbial abundance information is mapped onto network nodes to form a node feature matrix. Simultaneously, the network's adjacency matrix is ​​preserved to reflect the interactions between microorganisms. Finally, the sample is represented as a graph data structure containing both the node feature matrix and the adjacency matrix, avoiding the loss of inter-microbial relationship information caused by simple vector representation.

[0048] Step 33: Reconstruct the input graph data structure based on the microbial community ecological network topology to form graph structure data with multiple channels.

[0049] Step 34: In the convolutional layer, multi-scale convolutional kernels are used to perform parallel computation on the graph structure data of multiple channels in each channel to obtain multi-level and multi-dimensional features of the graph structure data of multiple channels, thus obtaining multi-scale convolutional features. The convolutional kernel size in this application can be set to 1×1, 3×3, 5×5, etc. A 1×1 convolutional kernel can be used to extract local detail features of a single microbial node, while larger graph convolutional kernels such as 3×3 and 5×5 are used to extract local structural features of the microbial community.

[0050] Step 35: Further fuse multi-level and multi-dimensional relational information between feature nodes through GCN graph convolution to obtain graph convolutional features. This application designs graph convolution operations to enable each microbial node to aggregate feature information from its neighboring nodes, while considering the weights of interactions between nodes, thus achieving adaptive learning of the complex network structure of microbial communities.

[0051] Step 36: Perform a weighted average fusion of multi-scale convolutional features and graph convolutional features to obtain the gut microbiome feature matrix.

[0052] This application constructs an ecological association map among microorganisms, mapping gut microbial abundance data onto this map to form an enhanced feature matrix containing ecological topological structure information, breaking the limitations of traditional simple vector forms. Secondly, this application employs a multi-channel input strategy, incorporating not only microbial abundance but also multimodal information such as functional and metabolic pathway information. Multiple sets of convolutional kernels extract detailed features from different information layers. Furthermore, multi-scale convolutional kernels are designed to cover receptive fields of varying sizes, capturing complex relationships from neighboring microbiota to larger ecological communities, enhancing the multi-level perception capability of features. This application integrates graph convolutional network concepts, combining adjacency information from the microbial association map and using graph convolution to aggregate and update node features, effectively capturing the network topology of the microbial community and achieving a deep integration of spatial features and topological relationships.

[0053] Step S4: Based on the cross-modal attention mechanism, perform bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features.

[0054] In this embodiment, based on a cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features. Please refer to [link / reference]. Figure 4 The specific content includes: Step 41: Using the clinical feature matrix as the query and the gut microbiome feature matrix as the key and value, a linear transformation is used to map the clinical feature matrix and the gut microbiome feature matrix to different subspaces, obtaining the corresponding query vector, key vector, and value vector. The attention score matrix is ​​then calculated. softmax Function normalization yields the attention weight matrix. Based on this matrix, the value vectors are weighted and summed to obtain the microbiome fusion feature guided by the clinical feature matrix, i.e., the first fusion feature. Assume the clinical feature matrix is... The gut microbiome feature matrix is Clinical feature matrix As a query, gut microbiome feature matrix Using these as keys and values, linear transformations are applied to map the clinical feature matrix and the gut microbiome feature matrix to different subspaces, resulting in... ,in , , Let be the learnable weight matrix. The attention score matrix is ​​calculated as follows: ,in Key vector Dimensions Representation of clinical characteristics within a specific subspace. (Through...) softmax Normalization of the function yields the attention weight matrix. Based on attention weight matrix value vector Weighted summation is performed to obtain the microbiome fusion features guided by the clinical feature matrix. .

[0055] Step 42: Using the gut microbiome feature matrix as the query and the clinical feature matrix as the key and value, a linear transformation is used to map the clinical feature matrix and the gut microbiome feature matrix to different subspaces, obtaining the corresponding query vector, key vector, and value vector. The attention score matrix is ​​then calculated. softmax Function normalization yields the attention weight matrix. Based on this matrix, the value vectors are weighted and summed to obtain the clinical fusion features guided by microbiome features, i.e., the second fusion feature. For example, the gut microbiome feature matrix... As a query, clinical feature matrix As keys and values, the clinical feature matrix and gut microbiome feature matrix are mapped to different subspaces through linear transformations to obtain the corresponding query vector, key vector, and value vector, i.e.: ,in 、 、 The weight matrix is ​​a learnable matrix. This represents the characteristics of the gut microbiome within a specific subspace. The attention score matrix is ​​calculated as follows: ,in Key vector The dimension. (Pass) softmax Normalization of the function yields the attention weight matrix. 。 Based on attention weight matrix value vector Weighted summation is performed to obtain the clinical feature matrix guided by microbiome fusion features. .

[0056] Step 43: The first and second fusion features acquired through bidirectional attention are fused using a weighted summation method to obtain the final fusion feature. The first fusion feature acquired through bidirectional attention... and the second fusion feature Perform weighted fusion to obtain the final fusion features. , ,in This is a hyperparameter used to balance the contributions of the modal characteristics of both modalities.

[0057] Step S5: Calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterology diseases.

[0058] In this embodiment, the SHAP value of the fused feature is calculated, and the fused feature is used as input to a preset risk prediction module to obtain a risk prediction value for gastroenterological diseases, including: Based on the SHAP theory of game theory, a computational framework suitable for fusion features is built. Each element in the fusion feature is regarded as a participant in the game. A fast approximation algorithm is used to calculate the SHAP value of each element in the fusion feature. The fusion features are ranked by importance, and key features whose SHAP values ​​meet the preset threshold are retained. The key fusion features filtered by SHAP values ​​are input into the preset risk prediction module. The preset risk prediction module integrates multiple prediction results, performs weighted or voting decisions, and finally outputs the risk prediction value of gastroenterology diseases in the form of probability.

[0059] This application calculates SHAP values ​​and retains key features whose SHAP values ​​meet a preset threshold. This reduces the dimensionality of the data, effectively avoids noise interference, and improves the interpretability and prediction accuracy of the model.

[0060] This application will focus on clinical and gut microbiome-related information that has a significant impact on the risk of gastroenterological diseases, providing precise data support for risk prediction.

[0061] It should be noted that the preset risk prediction module in this application can use logistic regression, random forest, support vector machine, etc. This application uses the risk prediction value of gastroenterological diseases to represent the probability of the target object suffering from a certain gastroenterological disease. At the same time, combined with SHAP value analysis, it can clearly show which key features have the main impact on the prediction results, providing clinicians with intuitive and interpretable decision-making basis and assisting doctors in formulating disease prevention and treatment plans.

[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0063] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0064] Continued reference Figure 5 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a risk prediction device based on gastrointestinal diseases. This system embodiment is similar to... Figure 2 Corresponding to the illustrated method embodiment, this system can be specifically applied to various electronic devices, including: a data acquisition module 501, a clinical feature acquisition module 502, a gut microbiome feature acquisition module 503, a fusion feature acquisition module 504, and a risk prediction value acquisition module 505. Wherein: The data acquisition module 501 is used to acquire clinical data and gut microbiome data of the target object, and to preprocess and standardize the clinical dataset and gut microbiome data respectively. The clinical feature acquisition module 502 is used to input the preprocessed clinical data into the first encoder to obtain the clinical feature matrix. The gut microbiome feature acquisition module 503 is used to input the preprocessed gut microbiome into the second encoder to obtain the gut microbiome feature matrix. The fusion feature acquisition module 504 is used to perform bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix based on the cross-modal attention mechanism to acquire fusion features; The risk prediction value acquisition module 505 is used to calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterological diseases.

[0065] This application acquires clinical data and gut microbiome data of the target subjects, preprocesses and standardizes the clinical dataset and gut microbiome data respectively; inputs the preprocessed clinical data into a first encoder to obtain a clinical feature matrix; inputs the preprocessed gut microbiome data into a second encoder to obtain a gut microbiome feature matrix; based on a cross-modal attention mechanism, performs bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features; calculates the SHAP value of the fused features, and uses the fused features as input to a preset risk prediction module to obtain risk prediction values ​​for gastroenterological diseases. This application, through bidirectional interaction and fusion of clinical data and microbiome data of the target subjects, combined with SHAP value analysis, can assist doctors in interpreting which key clinical indicators or gut microbiome features have a greater impact on risk prediction, assisting clinical decision-making, and further improving the interpretability of risk prediction for gastroenterological diseases.

[0066] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.

[0067] The computer device 6 includes a memory 6a, a processor 6b, and a network interface 6c that are interconnected via a system bus. It should be noted that only the computer device 6 with components 6a-6c is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0068] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0069] The memory 6a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 6a may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 6a may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 6a may also include both the internal storage unit and its external storage device of the computer device 6. In this embodiment, the memory 6a is typically used to store the operating system and various application software installed on the computer device 6, such as program code based on risk prediction methods for gastrointestinal diseases. In addition, the memory 6a can also be used to temporarily store various types of data that have been output or will be output.

[0070] In some embodiments, the processor 6b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 6b is typically used to control the overall operation of the computer device 6. In this embodiment, the processor 6b is used to run program code stored in the memory 6a or process data, for example, to run the program code for the risk prediction method based on gastrointestinal diseases.

[0071] The network interface 6c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 6 and other electronic devices.

[0072] This application also provides another embodiment, namely, providing a non-volatile computer-readable storage medium storing a program for a risk prediction method based on gastrointestinal diseases, the risk prediction based on gastrointestinal diseases being executable by at least one processor to cause the at least one processor to perform the steps of the risk prediction method based on gastrointestinal diseases as described above.

[0073] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0074] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A risk prediction method based on gastrointestinal diseases, characterized in that, include: Acquire clinical data and gut microbiome data of the target subjects, and preprocess and standardize the clinical dataset and gut microbiome data respectively; The preprocessed clinical data is input into the first encoder to obtain the clinical feature matrix; The preprocessed gut microbiome is input into the second encoder to obtain the gut microbiome feature matrix; Based on the cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features; Calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterology diseases.

2. The risk prediction method based on gastroenterological diseases according to claim 1, characterized in that, Before inputting the preprocessed clinical data into the first encoder, the following steps are included: Anatomical features of the intestine / abdomen are obtained from imaging data and transformed into shape description operators; Calculate the deviation between the biochemical indicators in laboratory test data and the corresponding normal reference range to obtain the abnormal offset vector; Keywords from historical diagnoses in clinical data are extracted and converted into disease association strength scores using a pre-defined medical knowledge graph.

3. The risk prediction method based on gastroenterological diseases according to claim 2, characterized in that, Anatomical features of the intestine / abdomen are obtained from imaging data and converted into shape description operators, including: Semantic detection is performed on text fields in clinical diagnosis and treatment data. When a keyword is detected, the image data corresponding to the keyword is retrieved from the image. Geometric analysis is performed on the image data to obtain geometric deformation parameters that reflect tissue morphology. The geometric deformation parameters are used as shape description operators.

4. The risk prediction method based on gastroenterological diseases according to claim 1, characterized in that, Calculate the deviation values ​​of biochemical indicators in laboratory test data from their corresponding normal reference ranges, and obtain the abnormal offset vector. This includes: calculating the direction and magnitude of the deviation of each biochemical indicator in the laboratory test data from its corresponding normal reference value, and obtaining the abnormal offset vector reflecting the degree of abnormality of the biochemical indicator based on the comprehensive expression of the deviation direction and magnitude.

5. The risk prediction method based on gastroenterological diseases according to claim 1, characterized in that, The preprocessed clinical data is input into the first encoder to obtain the clinical feature matrix, including: The preprocessed shape descriptor, anomaly offset vector and disease association strength score are input into the first encoder, which maps the shape descriptor into a spatial attention mask. The disease association strength score is adjusted based on the offset direction of the abnormal offset vector. Logical judgment is performed on the offset direction of the abnormal offset vector in the offset direction of a certain biochemical indicator. Based on the result of the logical judgment, the weight coefficient of the corresponding scoring path is adjusted. The clinical data is filtered by element-wise multiplication of the spatial attention mask with the disease association strength score and offset magnitude adjusted by the weight coefficients, and then the filtered data is uniformly scaled by linear projection to complete feature alignment and fusion, generating a clinical feature matrix.

6. The risk prediction method based on gastroenterological diseases according to claim 1, characterized in that, The preprocessed gut microbiome is input into the second encoder to obtain the gut microbiome feature matrix, including: Based on correlation analysis among microbial species, an ecological network of microbial communities was constructed. For each target object's gut microbiome sample, based on the constructed microbial community ecological network, the microbial abundance information is mapped to the network nodes to form a node feature matrix, while the adjacency matrix of the network is preserved. The sample is represented as a graph data structure containing the node feature matrix and the adjacency matrix. The input graph data structure is reconstructed based on the topology of the microbial community ecological network, forming graph structure data with multiple channels; In the convolutional layer, multi-scale convolutional kernels are used to perform parallel operations on the graph structure data of multiple channels in each channel to obtain multi-level and multi-dimensional features of the graph structure data of multiple channels, and to obtain multi-scale convolutional features. The graph convolutional features are further fused by GCN graph convolution to obtain graph convolutional features; The gut microbiome feature matrix is ​​obtained by weighted averaging and fusing multi-scale convolutional features with graph convolutional features.

7. The risk prediction method based on gastroenterological diseases according to claim 1, characterized in that, Based on a cross-modal attention mechanism, bidirectional attention calculation is performed on the clinical feature matrix and the gut microbiome feature matrix to obtain fused features, including: Obtain the microbiome fusion features guided by the clinical feature matrix, i.e., the first fusion feature; To obtain clinical fusion features guided by microbiome characteristics, i.e., second fusion features; The first and second fusion features obtained by bidirectional attention are fused together using a weighted summation method to obtain the final fusion feature.

8. A risk prediction device based on digestive diseases, used to implement the risk prediction method based on digestive diseases according to claims 1-7, characterized in that, include: The data acquisition module is used to acquire clinical data and gut microbiome data of the target subjects, and to preprocess and standardize the clinical dataset and gut microbiome data respectively. The clinical feature acquisition module is used to input the preprocessed clinical data into the first encoder to obtain the clinical feature matrix. The gut microbiome feature acquisition module is used to input the preprocessed gut microbiome into the second encoder to obtain the gut microbiome feature matrix. The fusion feature acquisition module is used to perform bidirectional attention calculation on the clinical feature matrix and the gut microbiome feature matrix based on the cross-modal attention mechanism to obtain fusion features; The risk prediction value acquisition module is used to calculate the SHAP value of the fusion feature and use the fusion feature as input to the preset risk prediction module to obtain the risk prediction value of gastroenterological diseases.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.