The present disclosure discloses a method and system for constructing a knowledge graph of a standard data element of a biomedical dataset, comprising collecting relevant standard texts of data elements of different types of biomedical datasets and data of a relevant standard of the biomedical dataset; analyzing and summarizing the relevant standard texts of the data elements of the different types of biomedical datasets and the data of the relevant standard of the biomedical dataset; constructing a knowledge model of the knowledge graph of the standard data element of the biomedical dataset; extracting entity type data and attribute data from structured data and an unstructured text in the structured data; and obtaining the knowledge graph of the standard data element of the biomedical dataset by performing knowledge fusion on a plurality of types of data based on an a plurality of types of semantic associative relationships between one or more entity types.
The invention provides a biomedical causal relationship inference method and system based on a Bayesian knowledge graph, and relates to the technical field of biomedical data mining and artificial intelligence, and the method comprises the steps: carrying out the multi-source evidence fusion of biomedical data, and obtaining a structured triple containing Bayesian confidence; analyzing the triple by using priori knowledge and obtaining a conditional probability table through parameterized filling; carrying out posteriori updating by using the Bayesian theorem; and analyzing the updated knowledge graph state by using a graph neural network model to obtain an inference result. Wherein the Bayesian inference module is combined with the graph neural network model, the former provides priori knowledge with confidence, the latter provides a fine path dependency relationship, and the accuracy and robustness of inference are remarkably improved. According to the method, the problems of evidence isomerism fragmentation, causal inference subjectization and knowledge discovery inefficiency are solved, and intelligent and automatic inference of the causal relationship is realized.
This application discloses a method, apparatus, device, and product for constructing a multidimensional knowledge graph, applicable to the field of data processing technology. The method includes: acquiring at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities; standardizing similar entities in the at least two biomedical databases to obtain at least two standardized entities; reconstructing entity relationships between different standardized entities based on the entity relationships between different entities in the at least two biomedical databases; and constructing the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities. This method can construct a multidimensional knowledge graph, primarily based on gene-related entities, by integrating databases.
The invention discloses a multi-modalbiomedical data security fusion query treatment method and system. The method comprises the following steps: receiving a fusion query request; legality verification is carried out through the block chain smart contract; decomposing the ontology model based on the multi-modalmetadata into sub-query tasks and distributing the sub-query tasks; each data holder node generates an intermediate result identified by a unified pseudonym identifier in a local privacy protection computing environment; executing data alignment and aggregation operation of privacy protection; and returning a fusion result and recording the key event in the block chain smart contract. According to the method, integrated treatment with data availability and invisibility, flexible query, process auditing and contribution incentive is achieved, and the method is suitable for safety collaborative analysis of multi-mode biomedical data such as genomes, images and electronic medical records.
This invention discloses a method, system, device, and storage medium for predicting drug-disease associations based on cross-propagation fusion and diffusion-guided multi-scale Transformer. The method acquires multi-source drug similarity, disease similarity, and known drug-disease association data; fuses multi-source similarity networks through local neighborhood sparsification and bidirectional cross-propagation; introduces noise based on the potential diffusion process and learns denoising representations; models the local-global interaction relationship between drugs and diseases using a multi-scale Transformerencoder; and finally outputs a drug-disease association probability score for ranking candidate treatment associations. This invention can improve the robustness and predictive performance of drug-disease association prediction under multi-source biomedical data and can be used for prioritizing drug relocation candidates.
The present invention belongs to the field of three-dimensional visualization technology and discloses a three-dimensional visualization method for a biomedical platform data map. The method includes installing SQLServer in the biomedical platform; constructing a biomedical database through SQL services; collecting biomedical product data and production and sales geographic data through a vertical search engine; transmitting the data to the platform in the form of a data stream; cleaning the collected medical data and storing it in the biomedical database; constructing a biomedical platform map canvas through a map construction program; constructing a map model on the map canvas, and displaying the map model in association with the biomedical data through key geographic information in the biomedical data; and constructing a biomedical knowledge graph, and associating the knowledge graph with the map model through a vector file. The present invention can reduce the number of displayed tables by constructing a biomedical platform map canvas through a map construction program, thereby improving the efficiency of constructing data maps.
Apparatus for identification of abnormal biomedical features within images of biomedical data and methods used therein are described. The apparatus includes an image capture device, a processor connected to the image capture device, a memory connected to the processor, and a display device connected to the processor. The image capture device is configured to capture an image of biomedical data. The memory contains instructions configuring the processor to receive the image, extract a plurality of biomedical features from the biomedical data, receive repository data from a medical repository as a function of the plurality of biomedical features, generate at least a distance metric as a function of the plurality of biomedical features and the repository data, and highlight at least a biomedical feature within the image as a function of the at least a distance metric.
The invention belongs to the technical field of data collaborative computing, and discloses a federated learning-based cross-institution biological medicine data collaborative computing method and system. Comprising the steps of performing statistical modeling on biological medicine data of each medical institution, and constructing a data portrait of each medical institution; local training hyper-parameters of the medical institutions are generated in a self-adaptive mode according to the data portraits, local dynamic training is executed on a pre-constructed global model, and local model parameter increments of the medical institutions are calculated; quantitatively evaluating the contribution index of each medical institution, and carrying out differentiated distribution on the computing resource quota and the training weight quota of each medical institution; fairly aggregating local model parameter increments of the medical institutions, and optimizing a global model; according to the method, a federal learning ecosystem which has competitive vitality and keeps cooperative balance is constructed, a feasible technical scheme is provided for safety sharing and value mining of cross-institution biological medicine data, and the dilemma of data islands is effectively broken through.
Some embodiments relate to methods, systems, and frameworks for data analytics using machine learning, such as methods and systems for preprocessing of biomedical data, using machine learning, for input to a predictive model. The method may include receiving data from a data source, using at least one machine learning (ML) algorithm from a plurality of ML algorithms to obtain at least one combination of preprocessing steps, and computing an accuracy score for each of the at least one combination based on accuracy of prediction of the predictive model. The method may further include using at least one ML algorithm to optimize the feature selection of the predictive model, combining a plurality of datasets into a single dataset, and using a parallel computing network to provide a framework for executing such predictive model.
This application relates to the field of natural languageprocessing technology, and in particular to a biomedical information extraction method based on a large language model. The method includes: acquiring the biomedical dataset to be processed and performing standardized preprocessing; converting the relation extraction data into high-dimensional vectors and constructing a local vector library; acquiring the medical information text to be processed as the query text, and performing a two-stage example retrieval and filtering in the local vector library to obtain a high-quality example set; generating context examples and performing context learning to understand the current task requirements and generate the information extraction results of the query text; and parsing the information extraction results. Based on the reordering capability of the cross-encoder model, this application designs a two-stage retrieval and reordering mechanism. By ensuring that the ICL examples provided to the large language model have both high semantic relevance and high task guidance, it significantly improves the accuracy and robustness of the model in biomedical named entity recognition and relation extraction tasks.
The invention discloses a biomedical knowledge graph construction method, which belongs to the technical field of knowledge graphs, and comprises the following steps: S1, data preprocessing: classifying and grading multi-source biomedical data from a literature database, an electronic medical record, genome data and a medical image, and desensitizing sensitive information through a differential privacy technology; s2, knowledge extraction and annotation: using a BERT-BiLSTM-CRF-based joint model to extract entities, combining an attention mechanism to identify relationships between the entities, and adopting an ontology framework to perform semantic annotation; s3, cross-source alignment and fusion: calculating entity similarity based on a graph embeddingalgorithm, and integrating entities and relationships in different data sources through a probability soft alignment model; and S4, a quality controlsystem: constructing a three-level verification mechanism comprising a rule engine, statistical analysis and expert feedback, and dynamically monitoring the integrity, consistency and timeliness of the knowledge graph.
This application relates to the field of biomedical data analysis technology and discloses a data analysis method, system, and storage medium for genetic diseasegene detection. The method includes the following steps: collecting patient samples and performing high-throughput sequencing to obtain raw data; performing quality control and comparison processing on the data to generate variant detection data and calculate the amount of variant information; automatically screening suspicious variant sites, analyzing the amount of variant information based on information entropy, and screening key variants according to preset thresholds; performing Bayesian inference analysis on key variants, calculating the probability of pathogenicity, and making pathogenicity determinations based on the ACMG standard; calculating the match between key variants and phenotypes based on patient phenotypic information, and screening variants that meet phenotypic characteristics; submitting the screening results to a doctor for review to generate the final screening results; and generating a standardized genetic test report based on the final results. This method efficiently and accurately identifies pathogenic variants, improving the automation level and clinical application value of genetic diseasegene detection.
This application discloses a method and system for secure fusion query governance of multimodal biomedical data. The method includes: receiving a fusion query request; verifying its legitimacy through a blockchainsmart contract; decomposing it into sub-query tasks based on a multimodal metadata ontology model and distributing them; each data holder node generating intermediate results identified by a unified pseudonym in its local privacy-preserving computing environment; performing privacy-preserving data alignment and aggregation operations; returning the fusion result and recording key events in the blockchainsmart contract. This application achieves integrated governance that ensures data is available but not visible, queries are flexible, processes are auditable, and contributions are incentivized. It is applicable to the secure collaborative analysis of multimodal biomedical data such as genomics, imaging, and electronic medical records.
The invention discloses a cardiovascular multi-modal data feature processing and description method, device and equipment and a medium, and relates to the field of artificial intelligence and biomedical data analysis. The method comprises the following steps: mapping cardiovascular multi-modal data of a target user to a unified semantic embedding space by adopting a multi-modal pre-training model to obtain a panoramic feature vector; optimizing the panoramic feature vector by adopting a context dynamic pruning mechanism; inputting the optimized simplified context vector into the trained multi-agent hierarchical collaborative system to obtain the feature description of the target user and the corresponding confidence coefficient; wherein a signal analysis agent in the multi-agent layered collaborative system outputs an electrophysiological feature vector and waveform classification; the morphological analysis agent outputs an anatomical structure feature vector and an image anomaly mark; and the global evaluation agent performs cross-modal consistency verification on the outputs of the first two agents. According to the method, the modal barrier is broken, and the feature processing and feature description of the cardiovascular multi-modal data are accurately and efficiently realized.
The invention provides a nerve cell regeneration drug target delineation method based on a large model, and the method comprises the steps: obtaining and fusing at least two multi-modalbiomedical data reflecting a nerve cell regeneration process, and constructing a dynamic knowledge graph; based on the dynamic knowledge graph, constructing a digital twinborn model for simulating target nerve cell regeneration; identifying a potential drug target based on a digital twin model, performing virtual intervention, and generating a prediction result of nerve cell regeneration response after intervention; based on the prediction result, experimental intervention is applied to potential drug targets in an observable biological model, quantifiable nerve cell regeneration related signals are collected, and experimental feedback data are obtained; and feeding back experimental feedback data to the digital twinborn model for iteration. According to the method, the target is dynamically simulated and optimized by constructing the digital twinborn model, and the model is continuously corrected through experimental data, so that the target screening accuracy and research and development efficiency are improved.
Embodiments of the present application provide a biomedical candidate gene discovery method based on text-graph fusion and related equipment. The method comprises: constructing a literature-derived semantic predicate graph (LDSPG) centered on a target disease; based on search enhancement generation, using a large language model to perform chain thinking reasoning on the evidence retrieved from the LDSPG, and curating high-quality training sample pairs; constructing a double-encoder model comprising a text encoder and a graph encoder, and aligning the text semantic space and the graph topological space to a unified embedding space through a hybrid contrastive loss function; merging the graph diffusionranking and the double-encoder ranking through a reciprocal ranking fusion algorithm to generate a fusion ranking result; and based on an external biomedical database, performing deterministic evidence grading on the candidate genes and outputting an auditable candidate genelist. Through high-quality data curation, cross-modalinformation fusion, complementary ranking fusion and traceable evidence grading, the accuracy and explainability of candidate gene discovery are effectively improved.
The invention provides a drug interaction prediction method and device based on dynamic pairing architecture search, and relates to the field of artificial intelligence. The method comprises the following steps: carrying out feature extraction on a molecular map of each drug by utilizing a motif decoupling encoder to obtain a structure sensing embedded representation of the drug; performing attention interaction on the structure perception embedded representation of each drug, and searching an operation structure from the candidate graph neural network operation set according to an attention interaction result to obtain a graph neural network encoder of each drug; performing layered polymerization on the structural mode in the molecular map of the drug by using the map neural network encoder of each drug to obtain a map-level embedded representation of each drug; the graph-level embedded representation of each medicine, the SMILES sequence of each medicine and the task cue word are input into a large language model to be processed, a prediction result is obtained, and the large language model is obtained by conducting instruction fine adjustment on medicine attribute information retrieved from a biomedicinedatabase.
The present invention discloses a biomedical big data classification method and system based on artificial intelligence. The method includes big data collection, dynamic knowledge base construction, pathway structure weighted constraints, biological constraint loss optimization, two-dimensional threshold decision optimization and biomedical big data classification. The present invention relates to the field of biomedical data classification technology. To address the problems of lack of biological structure constraints and poor interpretability in traditional methods, a dynamic disease-specific knowledge pathway map is constructed, and pathway structure weighted constraints and multi-level biological regularization mechanisms are introduced. By improving the neural network loss function, the fusion of feature learning and biological function network is achieved, thereby improving the model's discriminative ability and biological rationality. A two-dimensional threshold decision optimization strategy is adopted, combined with statistical significance and structural connectivity, to achieve accurate screening of key markers.
The invention is applicable to the field of data processing, and provides a multi-mechanism data processing method and system based on longitudinal federated learning, the method is applied to each medical mechanism client, and the method comprises the following steps: extracting individual features from biomedical data of a plurality of target individuals in a current medical mechanism, determining a first graph structure matrix corresponding to the current medical institution and used for indicating the similarity between the individual features; inputting the plurality of individual features and the first graph structure matrix into a first graph convolutional network model for iterative learning, and generating a local embedded representation and a local graph structure matrix corresponding to the current medical mechanism; and transmitting the local embedded representation and the local graph structure matrix to a server for disease diagnosis. According to the scheme, privacy protection of the biomedical data of the target individual can be realized, the communication cost can be reduced, and the data transmission efficiency can be improved.
The application discloses a biomedical storage system and method based on big data, and relates to the technical field of data processing.The method comprises the following steps: collecting a timestamp and a data block identifier of a biomedical data access operation, calculating a co-occurrence access frequency of a data block pair, constructing a data correlation graph and calculating an affinity value between data blocks, performing dynamic redistribution of the data blocks according to the affinity matrix, and storing the data blocks with high affinity in a centralized manner.A multi-level index structure comprising a global index layer and a local index layer is constructed, which is used for quickly locating data.When an access request is received, a candidate data block set is determined based on conditional access probability, and high-probability data blocks are prefetched to a cache.
The invention relates to a data processing method and device based on multi-modal contrast fusion and a graph neural network, and a medium. The method comprises the following steps: collecting multi-modalbiomedical data set samples of mRNA expression data, DNAmethylation data and miRNA expression data; embedded matrixes of three modes are obtained through an independent feedforward neural networkencoder; adopting a contrast learning mechanism based on NT-Xent to carry out unsupervised contrast alignment on the embedded matrixes of different modalities, and generating aligned embedded matrixes of three modalities; embedding and stacking the aligned modalities into a sequence, and inputting the sequence into a multi-layer Transform encoder to obtain a unified fusion representation; normalizing the fusion representation, and constructing a sample similarity graph structure; and based on the graph structure and the fusion representation, node high-order adjacency features are extracted through a three-layer graph convolutional network, sample classification is completed in combination with two full-connection layers, and a sample classification result is output. Compared with the prior art, the method has the advantages of cross-modal data fusion optimization, medical data discrimination enhancement, high efficiency and the like.