The invention provides a biomedical causal relationship inference method and system based on a Bayesian knowledge graph, and relates to the technical field of biomedical data mining and artificial intelligence, and the method comprises the steps: carrying out the multi-source evidence fusion of biomedical data, and obtaining a structured triple containing Bayesian confidence; analyzing the triple by using priori knowledge and obtaining a conditional probability table through parameterized filling; carrying out posteriori updating by using the Bayesian theorem; and analyzing the updated knowledge graph state by using a graph neural network model to obtain an inference result. Wherein the Bayesian inference module is combined with the graph neural network model, the former provides priori knowledge with confidence, the latter provides a fine path dependency relationship, and the accuracy and robustness of inference are remarkably improved. According to the method, the problems of evidence isomerism fragmentation, causal inference subjectization and knowledge discovery inefficiency are solved, and intelligent and automatic inference of the causal relationship is realized.
The invention discloses a multi-modalbiomedical data security fusion query treatment method and system. The method comprises the following steps: receiving a fusion query request; legality verification is carried out through the block chain smart contract; decomposing the ontology model based on the multi-modalmetadata into sub-query tasks and distributing the sub-query tasks; each data holder node generates an intermediate result identified by a unified pseudonym identifier in a local privacy protection computing environment; executing data alignment and aggregation operation of privacy protection; and returning a fusion result and recording the key event in the block chain smart contract. According to the method, integrated treatment with data availability and invisibility, flexible query, process auditing and contribution incentive is achieved, and the method is suitable for safety collaborative analysis of multi-mode biomedical data such as genomes, images and electronic medical records.
This invention discloses a method, system, device, and storage medium for predicting drug-disease associations based on cross-propagation fusion and diffusion-guided multi-scale Transformer. The method acquires multi-source drug similarity, disease similarity, and known drug-disease association data; fuses multi-source similarity networks through local neighborhood sparsification and bidirectional cross-propagation; introduces noise based on the potential diffusion process and learns denoising representations; models the local-global interaction relationship between drugs and diseases using a multi-scale Transformerencoder; and finally outputs a drug-disease association probability score for ranking candidate treatment associations. This invention can improve the robustness and predictive performance of drug-disease association prediction under multi-source biomedical data and can be used for prioritizing drug relocation candidates.
The invention belongs to the technical field of data collaborative computing, and discloses a federated learning-based cross-institution biological medicine data collaborative computing method and system. Comprising the steps of performing statistical modeling on biological medicine data of each medical institution, and constructing a data portrait of each medical institution; local training hyper-parameters of the medical institutions are generated in a self-adaptive mode according to the data portraits, local dynamic training is executed on a pre-constructed global model, and local model parameter increments of the medical institutions are calculated; quantitatively evaluating the contribution index of each medical institution, and carrying out differentiated distribution on the computing resource quota and the training weight quota of each medical institution; fairly aggregating local model parameter increments of the medical institutions, and optimizing a global model; according to the method, a federal learning ecosystem which has competitive vitality and keeps cooperative balance is constructed, a feasible technical scheme is provided for safety sharing and value mining of cross-institution biological medicine data, and the dilemma of data islands is effectively broken through.
Some embodiments relate to methods, systems, and frameworks for data analytics using machine learning, such as methods and systems for preprocessing of biomedical data, using machine learning, for input to a predictive model. The method may include receiving data from a data source, using at least one machine learning (ML) algorithm from a plurality of ML algorithms to obtain at least one combination of preprocessing steps, and computing an accuracy score for each of the at least one combination based on accuracy of prediction of the predictive model. The method may further include using at least one ML algorithm to optimize the feature selection of the predictive model, combining a plurality of datasets into a single dataset, and using a parallel computing network to provide a framework for executing such predictive model.
This application relates to the field of natural languageprocessing technology, and in particular to a biomedical information extraction method based on a large language model. The method includes: acquiring the biomedical dataset to be processed and performing standardized preprocessing; converting the relation extraction data into high-dimensional vectors and constructing a local vector library; acquiring the medical information text to be processed as the query text, and performing a two-stage example retrieval and filtering in the local vector library to obtain a high-quality example set; generating context examples and performing context learning to understand the current task requirements and generate the information extraction results of the query text; and parsing the information extraction results. Based on the reordering capability of the cross-encoder model, this application designs a two-stage retrieval and reordering mechanism. By ensuring that the ICL examples provided to the large language model have both high semantic relevance and high task guidance, it significantly improves the accuracy and robustness of the model in biomedical named entity recognition and relation extraction tasks.
This application discloses a method and system for secure fusion query governance of multimodal biomedical data. The method includes: receiving a fusion query request; verifying its legitimacy through a blockchainsmart contract; decomposing it into sub-query tasks based on a multimodal metadata ontology model and distributing them; each data holder node generating intermediate results identified by a unified pseudonym in its local privacy-preserving computing environment; performing privacy-preserving data alignment and aggregation operations; returning the fusion result and recording key events in the blockchainsmart contract. This application achieves integrated governance that ensures data is available but not visible, queries are flexible, processes are auditable, and contributions are incentivized. It is applicable to the secure collaborative analysis of multimodal biomedical data such as genomics, imaging, and electronic medical records.
The invention discloses a cardiovascular multi-modal data feature processing and description method, device and equipment and a medium, and relates to the field of artificial intelligence and biomedical data analysis. The method comprises the following steps: mapping cardiovascular multi-modal data of a target user to a unified semantic embedding space by adopting a multi-modal pre-training model to obtain a panoramic feature vector; optimizing the panoramic feature vector by adopting a context dynamic pruning mechanism; inputting the optimized simplified context vector into the trained multi-agent hierarchical collaborative system to obtain the feature description of the target user and the corresponding confidence coefficient; wherein a signal analysis agent in the multi-agent layered collaborative system outputs an electrophysiological feature vector and waveform classification; the morphological analysis agent outputs an anatomical structure feature vector and an image anomaly mark; and the global evaluation agent performs cross-modal consistency verification on the outputs of the first two agents. According to the method, the modal barrier is broken, and the feature processing and feature description of the cardiovascular multi-modal data are accurately and efficiently realized.
The invention provides a nerve cell regeneration drug target delineation method based on a large model, and the method comprises the steps: obtaining and fusing at least two multi-modalbiomedical data reflecting a nerve cell regeneration process, and constructing a dynamic knowledge graph; based on the dynamic knowledge graph, constructing a digital twinborn model for simulating target nerve cell regeneration; identifying a potential drug target based on a digital twin model, performing virtual intervention, and generating a prediction result of nerve cell regeneration response after intervention; based on the prediction result, experimental intervention is applied to potential drug targets in an observable biological model, quantifiable nerve cell regeneration related signals are collected, and experimental feedback data are obtained; and feeding back experimental feedback data to the digital twinborn model for iteration. According to the method, the target is dynamically simulated and optimized by constructing the digital twinborn model, and the model is continuously corrected through experimental data, so that the target screening accuracy and research and development efficiency are improved.
Embodiments of the present application provide a biomedical candidate gene discovery method based on text-graph fusion and related equipment. The method comprises: constructing a literature-derived semantic predicate graph (LDSPG) centered on a target disease; based on search enhancement generation, using a large language model to perform chain thinking reasoning on the evidence retrieved from the LDSPG, and curating high-quality training sample pairs; constructing a double-encoder model comprising a text encoder and a graph encoder, and aligning the text semantic space and the graph topological space to a unified embedding space through a hybrid contrastive loss function; merging the graph diffusionranking and the double-encoder ranking through a reciprocal ranking fusion algorithm to generate a fusion ranking result; and based on an external biomedical database, performing deterministic evidence grading on the candidate genes and outputting an auditable candidate genelist. Through high-quality data curation, cross-modalinformation fusion, complementary ranking fusion and traceable evidence grading, the accuracy and explainability of candidate gene discovery are effectively improved.
The application discloses a biomedical storage system and method based on big data, and relates to the technical field of data processing.The method comprises the following steps: collecting a timestamp and a data block identifier of a biomedical data access operation, calculating a co-occurrence access frequency of a data block pair, constructing a data correlation graph and calculating an affinity value between data blocks, performing dynamic redistribution of the data blocks according to the affinity matrix, and storing the data blocks with high affinity in a centralized manner.A multi-level index structure comprising a global index layer and a local index layer is constructed, which is used for quickly locating data.When an access request is received, a candidate data block set is determined based on conditional access probability, and high-probability data blocks are prefetched to a cache.
The application belongs to the technical field of biomedical data analysis and spatial omicsdata processing, and particularly relates to a disease target network construction method fusing pathological images and spatial transcriptome. Based on automatic or manual definition of a region of interest according to pathological characteristics, single-cell data and spatial transcriptome data are integrated; the spatial enrichment degree of different cell types and genes is quantitatively scored; and finally, a spatial co-localization network of genes and cells is constructed in the region of interest. The method can be compatible with various pathological imaging methods, and is suitable for irregularly shaped and significantly spatially heterogeneous disease tissues, overcoming the limitations of traditional spatial analysis methods guided by transcriptome characteristics in disease region positioning, and realizing systematic analysis from spatial pathology positioning to cell, gene and molecular correlation levels. The method is suitable for spatial mechanism research of complex diseases such as cardiovascular diseases, tumors and neurodegenerative diseases, and has high biological interpretation value and practical application significance.
The present application relates to the field of natural languageprocessing, and particularly relates to a biomedical event trigger extraction method and system based on hypergraph neural network. The method of the present application first preprocesses an unstructured biomedical data set; then obtains feature embeddings of all text information in the biomedical corpus through a pre-trained model to obtain vector representations of each word; then generates a corresponding hypergraph structure for each sentence; then inputs the feature embeddings and the hypergraph structure of each sentence into a hypergraph convolutional neural network to define a cross-entropy loss function to train the model; finally, trigger word detection is performed on an unlabeled test set. The present application differs from existing methods in that a bidirectional LSTM is used to aggregate context information in each sentence, but instead uses a hypergraph structure to aggregate context information, which is highly effective and achieves the purpose of improving the accuracy of trigger word extraction in biomedical text.
This application relates to a drugrelocation prediction method, system, device, and medium based on knowledge graphs. The method includes: acquiring entities from biomedical data and constructing a dynamic unified knowledge graph using entities as nodes and relationships as edges; extracting topological features of source disease subgraphs and target disease subgraphs based on the dynamic unified knowledge graph to obtain a set of source disease topological structure patterns and a set of target disease topological structure patterns; identifying cross-disease topological symmetric pattern pairs based on the source disease topological structure pattern set and the target disease topological structure pattern set to obtain a list of topological symmetric pattern pairs; comprehensively scoring each of the topological symmetric pattern pairs based on the list of topological symmetric pattern pairs to obtain a symmetric pattern scorelist; and mapping the drug derivation corresponding to the source disease to the target disease based on the symmetric pattern score list to obtain a drugrelocation scheme. This method can capture the mirror relationship between functional modules and regulatory logic.
This application discloses a method for determining developmental trajectories based on single-cell multi-omics clustering, belonging to the field of biomedical data mining technology. The method includes: acquiring multi-omics data of single cells from the same tissue; determining the potential representation of each single cell in different omics based on feature encoding technology; constructing a K-nearest neighbor graph for each omics based on the distance between single cells; and determining the corresponding multi-order similarity matrix; using the multi-order similarity matrix to complete the missing potential representations of single cells, obtaining the complete potential representation of each omics; performing cluster analysis on each omics to obtain single-cell clustering results; weighted fusion of the potential representations of each single cell in different omics to obtain a comprehensive potential representation; and analyzing the developmental trajectory of single cells based on the comprehensive potential representation. This application can stably and accurately cluster single cells under conditions of missing single-cell omics or significant differences in omics quality, thereby accurately determining the developmental trajectory of single cells.
Biomedical data can be applied to facilitate a vision test in a virtual reality (VR) environment using an electronic device that includes a head-mounted display (HMD) and a camera. The electronic device can direct the camera to an eye area of a user wearing the electronic device, and displays, on the HMD, a visual stimulus. While displaying the visual stimulus, in real time, the electronic device captures a sequence of eye images using the camera of the electronic device, and each eye image includes a respective region of interest (ROI) corresponding to a subset of the eye area of the user. Biomedical data are extracted from the sequence of eye images. The electronic device obtains a user response to the visual stimulus, and generates an output based on the user response and the biomedical data, the output indicating at least whether the user response satisfies a criterion.