Multi-modal large model-based scientific research data fusion analysis method and system
By jointly embedding and training multimodal large models and dynamically fusing features, the heterogeneity problem in multimodal data processing in the biomedical field is solved, achieving deep correlation and adaptive optimization of cross-modal data, and improving the efficiency and interpretability of scientific data fusion analysis.
Patent Information
- Application Number
- CN202512047696.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies suffer from feature space heterogeneity in multimodal data processing in the biomedical field, making it difficult to quantify cross-modal semantic correlations. Traditional methods are unable to dynamically adjust information integration mechanisms, resulting in systemic defects in cross-modal correlation modeling and dynamic decision optimization.
A multimodal large model is used for joint embedding training to generate a cross-modal unified feature space. Through dynamic feature fusion and task decomposition driven by a large language model, the generation and iterative optimization of cross-modal feature vectors are realized. Combined with knowledge graphs and multi-objective optimization algorithms, a closed-loop feedback mechanism is formed.
It achieves deep correlation mapping and adaptive feature reconstruction of cross-modal data, improves the efficiency and interpretability of scientific research data fusion analysis, and significantly enhances the efficiency and reproducibility of uncovering implicit patterns in biomedical research and development.
Smart Images

Figure CN121483565A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scientific research data analysis technology, and more specifically, to a method and system for scientific research data fusion analysis based on a multimodal large model. Background Technology
[0002] In the current biomedical field, the processing of multimodal data has become an important part of scientific research and development. However, existing technologies usually adopt isolated modeling methods when processing different modal data. The independent training involved in this method leads to the heterogeneity of feature space and makes it difficult to quantify cross-modal semantic correlation, especially in the representation of complex entity relationships in the biomedical field, where there are significant biases.
[0003] Furthermore, traditional multimodal fusion methods often employ static weight allocation or post-feature concatenation strategies, failing to dynamically adjust the information integration mechanism according to specific research task objectives. This results in a lack of adaptability in application scenarios involving experimental constraints and data association rules. Moreover, existing intelligent decision-making systems mostly use fixed-process rule engines for task decomposition, unable to achieve context-aware task planning based on large language models, leading to a mismatch between the inherent correlation between sub-task sequences and multimodal data. In the data optimization stage, traditional techniques often rely on iterative models driven by human experience, lacking a closed-loop feedback mechanism to back-map decision results to the multimodal feature space for adaptive adjustment. This makes it difficult to meet the needs of multi-objective collaborative optimization in biomedical research and development. These problems result in systemic deficiencies in cross-modal correlation modeling and dynamic decision optimization using traditional methods. Summary of the Invention
[0004] In order to at least overcome the above-mentioned shortcomings in the prior art, one of the objectives of the present invention is to provide a scientific research data fusion and analysis method and system based on a multimodal large model.
[0005] This invention provides a method for scientific research data fusion and analysis based on a multimodal large-scale model, comprising: acquiring a multimodal scientific research dataset, the multimodal scientific research dataset including biomedical text data, gene sequence data, chemical structure data, and experimental image data; performing joint embedding training on the multimodal scientific research dataset to generate a cross-modal unified feature space, the joint embedding training including synchronous optimization of parameters of a text encoder, gene encoder, chemical encoder, and image encoder; performing dynamic feature fusion on the cross-modal unified feature space based on task parsing instructions to generate cross-modal feature vectors matching the task parsing instructions; the task parsing instructions including scientific research objective parsing, experimental condition constraints, and data association rules; calling a pre-trained multimodal large-scale language model to perform task decomposition on the cross-modal feature vectors to generate at least one sequence of scientific research sub-tasks to be executed; and iteratively optimizing the multimodal scientific research dataset according to the scientific research sub-task sequence to output scientific research decision parameters that meet preset biomedical research and development indicators.
[0006] This invention also provides a scientific research data fusion and analysis system, including a processor, a memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the above-described scientific research data fusion and analysis method based on a multimodal large model.
[0007] This invention also provides a computer-readable storage medium storing a program that, when executed by a processor, implements the above-described method for scientific data fusion and analysis based on a multimodal large model.
[0008] This invention provides a scientific research data fusion and analysis method and system based on a multimodal large model. First, through a joint embedding training mechanism of a multimodal encoder, a deep cross-domain semantic association mapping is established in a unified feature space. This enables resolvable vectorized alignment of textual semantics, gene expression patterns, chemical spatial topology, and image phenotypic features, breaking through the information barriers of traditional single-modal modeling. Second, a dynamic feature fusion engine is innovatively introduced. Based on the scientific research objectives, experimental constraints, and association rules of the task parsing instructions, the contribution weights and combination paths of different modal features are adaptively adjusted, thereby achieving demand-oriented representation reconstruction while maintaining the inherent correlation of the data. Finally, through a task decomposition architecture driven by a large language model, complex scientific research problems are transformed into interpretable sub-task execution chains, forming a closed-loop optimization mechanism of "global decision-local verification," significantly improving the efficiency and reproducibility of uncovering implicit scientific research patterns.
[0009] In summary, the multi-level collaborative intelligent computing framework provided in this application can realize the entire process from scientific research data fusion to knowledge reasoning, effectively solving the systemic defects of traditional methods in cross-modal association modeling and dynamic decision optimization. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 The flowchart illustrates a scientific data fusion and analysis method based on a multimodal large model, as provided in an embodiment of the present invention.
[0012] Figure 2 This is a block diagram of a scientific research data fusion and analysis system provided in an embodiment of the present invention.
[0013] icon: 100-Scientific Research Data Fusion and Analysis System; 101 - Processor; 102 - Memory; 103 - Bus. Detailed Implementation
[0014] Exemplary embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0015] To better understand the above technical solutions, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0016] Figure 1 The flowchart of a scientific research data fusion and analysis method based on a multimodal large model according to an embodiment of the present invention is applied to a scientific research data fusion and analysis system, including steps 101-105.
[0017] Step 101: Obtain a multimodal scientific research dataset, which includes biomedical text data, gene sequence data, chemical structure data, and experimental image data.
[0018] In this embodiment, a systematic aggregation of biomedical research data can be achieved by constructing a multi-source heterogeneous data integration channel. For example, the scientific research data fusion and analysis system deploys a standardized application programming interface cluster, connecting with mainstream international biomedical databases and local private data warehouses to form a full-dimensional data acquisition network covering text, sequences, structures, and images.
[0019] For biomedical text data, the system is configured with a natural language processing pipeline to extract research results related to disease targets from academic paper databases, including core content such as experimental hypotheses, methodological descriptions, and results. It also integrates pharmacodynamic and safety data from drug clinical trial reports. The acquisition of gene sequence data employs a distributed computing framework, downloading raw sequencing data involving key gene mutations in parallel from multiple omics databases, including various genetic variation types such as single nucleotide polymorphisms, copy number variations, and structural variations. The integration of chemical structure data relies on a professional chemical information platform, using a structure search component to batch acquire small molecule compounds with specific pharmacological activities, simultaneously collecting their physicochemical properties, synthetic routes, and patent information. The processing of experimental image data incorporates a computer vision preprocessing module to standardize the format and resolution of raw files output from microscopic imaging equipment, covering multi-level visual data such as subcellular structure observation, tissue section analysis, and in vivo imaging records. All data undergoes a multi-level quality control system during the data entry stage, including a format verification component to verify data integrity, a metadata annotation system to establish a data traceability index, and an anomaly detection algorithm to eliminate noise interference.
[0020] Ultimately, by constructing knowledge graph components, discrete multimodal data entities are transformed into a network of nodes and edges. Gene mutation nodes are connected to protein function change entities through "cause" relationships, and compound nodes are connected to disease target entities through "target" relationships, forming a biomedical knowledge system that supports cross-modal reasoning.
[0021] Step 102: Perform joint embedding training on the multimodal scientific research dataset to generate a cross-modal unified feature space. The joint embedding training includes synchronous optimization of the parameters of the text encoder, gene encoder, chemical encoder and image encoder.
[0022] In this embodiment, the scientific data fusion and analysis system can construct a modular encoder architecture. The text encoder adopts a hierarchical Transformer structure, initializing parameters through a pre-trained biomedical language model. During the fine-tuning phase, it focuses on learning specialized semantic patterns such as target protein functional descriptions and drug action mechanisms. The gene encoder designs a dual-channel convolutional neural network to process the base-coding features of DNA sequences and epigenetic modification signals, respectively. It captures long-range dependencies through dilated convolution operations, effectively identifying the combined features of regulatory elements. The chemical encoder is developed based on a geometric graph neural network, resolving molecular structures into geometric graph data containing atomic types, bonding types, and spatial coordinates. It uses rotationally equivariant convolutional layers to achieve equivariant feature extraction of three-dimensional conformations. The image encoder integrates a multimodal fusion attention mechanism, adding a cross-channel feature interaction module to the traditional convolutional neural network, enabling synergistic enhancement of cell morphology features and molecular marker distribution features.
[0023] Understandably, the system achieves joint encoder optimization through a cross-modal contrastive learning framework. A triplet loss function is designed to drive the approximation of positive sample pairs (such as a gene mutation and its corresponding pathological image) in the feature space, while simultaneously distancing unrelated sample pairs. During training, a dynamic weight allocation strategy is introduced, automatically adjusting the contribution ratio of the loss function based on the information entropy of each modality's data to ensure semantic balance in the feature space. After sufficient training, the deep features of different modalities form a biologically meaningful topological structure within a unified vector space. For example, the activation state of oncogenes, the abnormal activation of corresponding signaling pathways, and the invasive phenotype of tumor cells exhibit a dense clustered distribution in space, providing an interpretable feature basis for subsequent cross-modal association analysis.
[0024] Step 103: Perform dynamic feature fusion on the cross-modal unified feature space based on the task parsing instructions to generate a cross-modal feature vector that matches the task parsing instructions; the task parsing instructions include scientific research goal parsing, experimental condition constraints, and data association rules.
[0025] In this embodiment, when the scientific research data fusion and analysis system performs an adaptive feature fusion process oriented towards R&D tasks, after the user submits the R&D instruction "develop a novel kinase inhibitor", the task parsing component initiates a multi-dimensional instruction interpretation process: First, the semantic parsing module identifies the core elements in the instruction, including key constraints such as target type (kinase family), expected activity intensity (nanomolar IC50 value), and drug type (oral small molecule); then, the knowledge graph reasoning module is activated to expand the associated entities along the "kinase-signaling pathway-disease phenotype" path and construct a contextual knowledge network covering target structure, compound library, and activity detection methods.
[0026] Furthermore, the feature fusion controller configures a dynamic weight matrix based on the analysis results. Tailored to the characteristics of inhibitor development tasks, it increases the weight coefficients of complementary features in the kinase binding pocket within the chemical structure modality and enhances the feature strength of SNP sites related to kinase expression regulation in the gene modality, while suppressing task-irrelevant cell migration image features. The system employs a gated attention mechanism to enhance cross-modal feature interaction. A learnable parameter gating function controls the fusion depth of different modal features. For example, when predicting compound permeability, it strengthens the cross-modal interaction between chemical structure and cell membrane image features, while when assessing metabolic stability, it emphasizes the correlation analysis between chemical structure and liver enzyme gene expression profiles.
[0027] Based on the above, multi-scale feature aggregation can be performed using a feature pyramid network to generate a comprehensive feature vector containing molecular design guidance information. This vector encodes both the stereoelectronic properties of the kinase catalytic domain and the implicit rules of the lead compound structure optimization direction.
[0028] In step 103, the cross-modal unified feature space is a semantic alignment environment formed by mapping heterogeneous data such as biomedical text, gene sequences, chemical structures, and experimental images into a shared vector space through joint embedding training techniques. The cross-modal unified feature space utilizes the hierarchical Transformer structure of the text encoder to parse protein function descriptions in papers, captures the combined features of regulatory elements in DNA sequences through the dual-channel convolutional network of the gene encoder, extracts molecular three-dimensional conformation information using the geometric graph neural network of the chemical encoder, and integrates cell morphology features from microscopic images using the multimodal fusion attention mechanism of the image encoder. Within the cross-modal contrastive learning framework, the system forces biologically related data pairs (such as a gene mutation and its resulting pathological image) to be close to each other in the vector space, and balances the differences in information density between different modalities through a dynamic weight allocation strategy. Ultimately, this allows complex biological phenomena such as oncogene activation states, abnormal activation of signaling pathways, and tumor cell invasion phenotypes to form interpretable topological structures within the space, providing a mathematical representation basis for cross-modal associative reasoning.
[0029] Cross-modal feature vectors are comprehensive representation vectors generated by dynamically fusing multi-source data in a unified feature space based on task requirements. When the system receives the instruction "develop novel kinase inhibitors," the task parsing component first extracts core elements such as target type, activity intensity, and drug form, activating the knowledge graph to expand contextual associations along the "kinase-signaling pathway-disease" path. The feature fusion controller adjusts the weight matrix through a gated attention mechanism, enhancing the contribution of complementary features of kinase binding pockets in the chemical modality and expression-regulating SNP sites in the gene modality, while suppressing interference from irrelevant image features. A feature pyramid network is used to aggregate cross-scale information such as molecular pharmacophore details and kinase evolutionary conservation. The final generated vector encodes both the stereoelectronic properties of the kinase catalytic domain and the implicit rules of lead compound optimization direction, providing input features for virtual screening and structure-activity relationship analysis.
[0030] In this embodiment, for task parsing instructions, research objective parsing is a crucial process that transforms user-proposed R&D needs into executable computational tasks, involving a dual mechanism of natural language understanding and knowledge graph reasoning. When the instruction "develop an oral EGFR inhibitor" is input, the semantic parsing module identifies elements such as the target entity (EGFR), administration method (oral), and molecular type (small molecule), converting the colloquial description into standard biomedical concepts through ontology mapping. The knowledge graph reasoning engine retrieves existing research results along the path "EGFR → non-small cell lung cancer → tyrosine kinase inhibitor → osimertinib," and constructs the task context by combining related information such as the drug resistance mutation T790M. This process requires coordination between entity recognition, relation extraction, and logical reasoning modules to ensure that vague R&D intentions are transformed into clear molecular design constraints, such as requiring compounds to meet a permeability score ≥5 and avoid hERG channel inhibition.
[0031] Experimental constraints are a set of physicochemical and biological limitations that must be followed during the research and development process. These constraints are integrated into the computational flow through a multi-objective optimization algorithm. In the kinase inhibitor development scenario, hard constraints include drug-likeness (compliance with Lipinski's five rules) and synthetic feasibility (synthetic step complexity score < 5), while soft constraints involve balancing target affinity (IC50 ≤ 10 nM) and metabolic stability (liver microsomal half-life ≥ 30 minutes). The system encodes constraints as penalty terms in a loss function. When a virtual screening compound violates the hERG toxicity threshold (inhibition rate < 50%), the loss value is automatically increased, driving the optimization direction away from high-risk areas. Pareto front analysis is used to identify the trade-offs between indicators such as activity, toxicity, and synthetic difficulty, dynamically adjusting the priority of each constraint. For example, during the lead compound optimization stage, the synthetic complexity can be temporarily relaxed to explore potential for activity enhancement.
[0032] Data association rules are a knowledge representation system that ensures the biological rationality of cross-modal reasoning, built upon a framework that integrates descriptive logic and graph neural networks. The system's built-in structure-function association rules stipulate that if a compound forms at least two hydrogen bonds with the ATP binding pocket of the kinase and has a van der Waals surface complementarity >70%, it is considered to have potential inhibitory activity. Gene-phenotype association rules clarify the logical causal relationship between TP53 gene mutations and abnormal cell cycle regulation. When the gene encoder detects specific splicing variations, it automatically triggers a tumor proliferation phenotype warning. These rules are formally expressed using the OWL ontology language, constructing a causal chain in the knowledge graph of "EGFR L858R mutation → sustained activation of the kinase domain → increased nucleocytoplasmic ratio in pathological images," ensuring that the dynamic feature fusion process conforms to established biomedical principles and avoids generating association hypotheses that violate common sense.
[0033] Step 104: Call the pre-trained multimodal large language model to perform task decomposition on the cross-modal feature vectors and generate at least one sequence of scientific research sub-tasks to be executed.
[0034] In this embodiment, the scientific research data fusion and analysis system can initiate an intelligent R&D task planning and decomposition mechanism. Specifically, the system calls a multimodal large language model pre-trained based on a biomedical knowledge graph. This model, through self-supervised learning from massive amounts of scientific literature and patent data, has established a complete knowledge reasoning chain from molecular structure to clinical efficacy.
[0035] After inputting the cross-modal feature vectors into the model, the task objectives are first deconstructed: identifying the core scientific problems that need to be addressed (such as improving kinase subtype selectivity) and clarifying the necessary boundary conditions (such as avoiding hERG channel inhibition). Then, based on the gold standard workflow of drug development, a task sequence including virtual screening, structure-activity relationship analysis, and ADMET prediction is automatically generated. Each subtask is configured with detailed execution parameters; for example, virtual screening requires Lipinski rule filtering, and structure-activity relationship analysis requires the construction of a three-dimensional pharmacophore model. The task orchestration component uses critical path analysis to optimize the execution process, identifying the strong dependency between lead compound optimization and in vitro activity validation, and establishing a pipeline mechanism for task execution accordingly.
[0036] Simultaneously, the system can deploy a dynamic resource scheduler to allocate heterogeneous computing resources based on the computational complexity and hardware requirements of each subtask. For example, molecular dynamics simulation tasks can be prioritized for scheduling to GPU clusters, while literature data mining tasks can be assigned to CPU nodes for execution, ensuring optimal time efficiency for the overall R&D process.
[0037] Step 105: Iteratively optimize the multimodal scientific research dataset according to the scientific research sub-task sequence, and output scientific research decision parameters that meet the preset biomedical research and development indicators.
[0038] In this embodiment of the application, the scientific data fusion and analysis system can implement a closed-loop feedback-driven iterative optimization strategy in this step. For example, the system can establish a bidirectional feedback channel for cross-modal data: molecular docking results update the compound activity feature vector in real time, in vitro experimental data correct the scoring function of the virtual screening model, and animal experimental conclusions simultaneously optimize the parameter configuration of the ADMET prediction model.
[0039] Understandably, in each iteration, the reinforcement learning controller evaluates the effectiveness of historical decisions and dynamically adjusts the collaboration strategies of each module. For example, when a deviation is found between the drug-likeness prediction and the measured results, the weight coefficient of the compound structure diversity screening is automatically increased. The system uses a multi-objective optimization algorithm to balance various R&D indicators, constructs an evaluation space that includes dimensions such as target affinity, metabolic stability, and synthetic feasibility, and determines the optimal solution set through Pareto front analysis.
[0040] Furthermore, after multiple iterations, the characterization capabilities of the feature space have continuously evolved, enabling it to capture substructure modification patterns that are difficult to detect using traditional methods, such as the nonlinear enhancement effect of specific heterocyclic substituents on kinase selectivity. The final output R&D plan includes validated lead compound structures, recommended synthetic routes, and corresponding in vitro validation experimental designs, forming a complete evidence chain from computational prediction to experimental verification, significantly improving the success rate and efficiency of innovative drug development.
[0041] More specifically, in step 105, iterative optimization refers to a dynamic improvement process based on a closed-loop feedback mechanism, gradually approaching the optimal solution through multiple rounds of data-driven adjustments. In the biopharmaceutical R&D scenario, the system achieves adaptive optimization of the R&D plan through a bidirectional feedback channel of cross-modal data: molecular docking calculation results update the compound activity feature vector in real time, in vitro experimental data reversely correct the scoring function of the virtual screening model, and animal experimental conclusions simultaneously optimize the parameter configuration of the ADMET prediction model. Each iteration uses a reinforcement learning algorithm to evaluate the effectiveness of historical decisions, dynamically adjusts the collaborative strategies of each module (such as increasing the screening weight of compound structural diversity), and uses a multi-objective optimization algorithm to balance indicators such as target affinity, metabolic stability, and synthetic feasibility. This process captures substructure modification patterns that are difficult to discover using traditional methods (such as the nonlinear effect of heterocyclic substituents on kinase selectivity) through continuous evolution of the feature space, ultimately forming a complete evidence chain from computational prediction to experimental verification.
[0042] Furthermore, the pre-defined biopharmaceutical R&D indicators refer to a system of quantitative and qualitative evaluation standards that must be met throughout the entire drug development process. This system constructs a multi-dimensional evaluation space, including core dimensions such as target affinity (e.g., IC50 value ≤ 10 nM), metabolic stability (liver microsomal half-life ≥ 30 minutes), synthetic feasibility (synthetic step complexity score < 5), drug-likeness (compliance with Lipinski's five rules), toxicity risk (hERG inhibition rate < 50%), and intellectual property patentability. These indicators are used to determine their interrelationships through Pareto frontier analysis. For example, when optimizing the activity of kinase inhibitors, their impact on cardiotoxic pathways must be monitored simultaneously. The system automatically triggers optimization path adjustments based on indicator thresholds. For instance, when the membrane permeability score of the lead compound is lower than the pre-defined critical value, the lipophilic modification scheme of the molecular skeleton is prioritized.
[0043] Finally, research decision parameters are the core elements of the executable plan output after multimodal data fusion and iterative optimization. These parameters include validated three-dimensional structural parameters of the lead compound (such as the spatial coordinates of pharmacophore feature points), recommended synthetic routes (involving yield and purity thresholds for key intermediates), in vitro validation experimental designs (including concentration gradient settings for IC50 assays to detect kinase activity), and risk control parameters (such as hepatotoxicity warning thresholds in ADMET prediction). These parameters are generated through knowledge graph reasoning and feature pyramid network aggregation, encoding both the stereoelectronic properties of the kinase catalytic domain and the substructure modification rules implicit in the structure-activity relationship model (such as the contribution coefficient of ortho-substituents on selectivity of the benzene ring), forming a key quantitative basis for guiding drug development.
[0044] In one implementation, step 102, which involves jointly embedding and training the multimodal research dataset to generate a cross-modal unified feature space, includes: Step 1021: Perform semantic segmentation on the biomedical text data to generate a text semantic block sequence, and use a bidirectional attention mechanism to mine the context encoding corresponding to the text semantic block sequence.
[0045] When performing semantic segmentation on biomedical text data, a natural language processing pipeline is first configured to segment the input academic paper into paragraphs, identifying semantic units containing core research information such as disease target descriptions, experimental methods, and results discussions. For example, when processing a paper on the relationship between EGFR gene mutations and lung cancer, the system divides "EGFR exon19 deletion mutation leads to sustained activation of the tyrosine kinase domain" in the abstract into an independent semantic block, and "using CRISPR-Cas9 to construct mutant cell lines" in the methods section into an experimental operation semantic block.
[0046] Subsequently, the system employs a bidirectional attention mechanism to perform contextual encoding on the text semantic block sequence, capturing cross-sentence logical connections through a pre-trained biomedical language model. Specifically, when encoding the semantic block "EGFR mutation promotes downstream MAPK pathway activation," the model identifies a causal relationship between the previously mentioned "elevated EGFR phosphorylation level" and this block through a self-attention layer, ultimately generating a contextual encoding vector containing long-range dependencies.
[0047] Step 1022: Perform base encoding mapping on the gene sequence data to generate gene fragment vectors, and use a sliding window convolution kernel to extract local functional region features from the gene fragment vectors.
[0048] When performing base coding mapping on gene sequence data, the original FASTA format DNA sequence is converted into a four-dimensional one-hot vector representation. For example, when processing sequencing data of the BRAF gene V600E mutation site, the corresponding "T" base is encoded as [0, 0, 0, 1], and the base sequences within a 500bp range upstream and downstream are converted into continuous vectors according to this rule. Next, the system uses a sliding window convolutional kernel to extract local functional region features from the gene fragment vector. A convolutional kernel with a width of 21bp slides along the sequence to capture conserved TATA box features in the promoter region. When detecting the KRAS gene coding region, the three-layer dilated convolutional structure can cross exon junctions, effectively identifying local sequence pattern changes caused by the G12D mutation at codon 12, and generating local feature vectors with functional annotation significance.
[0049] Step 1023: Perform topological graph transformation on the chemical structure data to generate an atomic node graph, and use a graph neural network to aggregate the chemical bond connections in the atomic node graph.
[0050] When performing topological graph transformation on chemical structure data, the two-dimensional structural formulas of small molecule compounds are parsed into graph data structures containing atomic nodes and chemical bond edges. For example, when processing imatinib molecules, the system identifies the single bond connections between nitrogen atom nodes and adjacent carbon atom nodes in the aniline group, and the conjugated double bond edge properties on the pyridine ring. Subsequently, the system uses a graph neural network to aggregate the chemical bond connections in the atomic node graph, updating the feature representation of each atomic node through a message passing mechanism. When processing the cyano substituent of the EGFR inhibitor osimertinib, the graph convolutional layer can capture the electron cloud distribution characteristics between triple-bonded carbon and nitrogen atoms, generating a chemical bond connection feature vector that reflects the overall hydrophobicity and hydrogen bond donor capability of the molecule.
[0051] Step 1024: Perform multi-scale segmentation on the experimental image data to generate a set of regional image patches, and use a spatial attention mechanism to extract biomarker features from the set of regional image patches.
[0052] When performing multi-scale segmentation on experimental image data, the microscopic imaging files are first standardized and preprocessed. For example, when processing tumor tissue slice images, the system segments 50 μm diameter nucleus regions under a 20x objective lens and extracts 2 μm chromatin condensation feature blocks under a 40x objective lens. Subsequently, the system employs a spatial attention mechanism to extract biomarker features from the set of regional image blocks, enhancing the signal intensity of important regions through a learnable attention weight matrix. When processing HER2 immunohistochemical staining images, the attention mechanism focuses on brown deposition regions on the cell membrane surface, generating biomarker feature vectors reflecting the overexpression level of HER2 protein.
[0053] Step 1025: Jointly train the context encoding, the local functional region features, the chemical bond connection relationship and the biomarker features by using a contrastive learning loss function. The joint training is used to indicate that the projection of different modal data in the cross-modal unified feature space satisfies the semantic alignment condition.
[0054] When performing joint training using a contrastive learning loss function, cross-modal semantic alignment constraints can be constructed. The system uses pathology report text, driver gene sequencing data, therapeutic drug molecular structures, and tissue slice images of the same cancer sample as positive sample groups, ensuring that data from different modalities are mapped to adjacent regions in a unified feature space. For example, ERBB2 gene amplification data, Herceptin molecular structure data, and their immunohistochemical images from breast cancer patients form compact clusters in the feature space. Simultaneously, the system randomly replaces gene data in the positive sample group with irrelevant TP53 mutation sequences to generate negative sample groups, optimizing the feature space topology by maximizing the distance between positive and negative samples.
[0055] In one implementation, step 1025, which involves jointly training the context encoding, the local functional region features, the chemical bond connections, and the biomarker features using a contrastive learning loss function, includes: Step 10251: Extract biomedical text data, gene sequence data, chemical structure data and experimental image data of the same biological entity from the multimodal scientific research dataset as a positive sample group, and randomly replace at least one of the modal data to generate a negative sample group.
[0056] When constructing positive and negative sample groups, it is essential to strictly adhere to the principle of biological entity correspondence. For example, the positive sample group can be constructed by selecting the textual description of the α-synuclein gene SNCA associated with Parkinson's disease, genotype data at the rs356219 locus, molecular structure of dopamine analogs, and electron micrographs of Lewy bodies in substantia nigra neurons. When generating the negative sample group, the system retains the textual description, replaces the gene data with the APP gene sequence associated with Alzheimer's disease, replaces the chemical structure with acetylcholinesterase inhibitors, and replaces the images with microscopic images of amyloid plaques, thus creating interference samples across disease types.
[0057] Step 10252: Input the biomedical text data of the positive sample group into the text encoder to generate a first text feature vector, input the gene sequence data of the positive sample group into the gene encoder to generate a first gene feature vector, input the chemical structure data of the positive sample group into the chemical encoder to generate a first chemical feature vector, and input the experimental image data of the positive sample group into the image encoder to generate a first image feature vector.
[0058] When performing positive sample feature extraction, for example, the text data "BRAF V600E mutation leads to continuous activation of the MAPK pathway" of the positive sample group is input into the text encoder to generate the first text feature vector containing the semantics of kinase activity regulation; the "chr7:140753336T>A" mutation site in the corresponding gene data is processed by the gene encoder to output the first gene feature vector reflecting the allosteric effect of the kinase domain; the chemical encoder processes the vemurafenib molecular structure to generate the first chemical feature vector targeting the BRAF kinase pocket; and the image encoder analyzes the melanoma cell proliferation image to output the first image feature vector showing the activation characteristics of the MAPK pathway.
[0059] Step 10253: Calculate the first intermodal similarity between the first text feature vector and the first gene feature vector in the cross-modal unified feature space, and calculate the second intermodal similarity between the first chemical feature vector and the first image feature vector.
[0060] When calculating intermodal similarity, cosine similarity can be used to measure the strength of cross-modal feature association. For example, the intermodal similarity between the semantic "kinase domain allosteric variation" in the first text feature vector and the V600E mutation feature in the first gene feature vector reaches 0.85, reflecting a strong correlation between the two at the level of kinase activity regulation; the intermodal similarity between the DFG-out binding mode of vemurafenib in the first chemical feature vector and the tumor cell proliferation inhibition feature in the first image feature vector is 0.78, indicating the correspondence between drug action mechanism and phenotypic change.
[0061] Step 10254: Input the biomedical text data of the negative sample group into the text encoder to generate a second text feature vector, input the gene sequence data of the negative sample group into the gene encoder to generate a second gene feature vector, input the chemical structure data of the negative sample group into the chemical encoder to generate a second chemical feature vector, and input the experimental image data of the negative sample group into the image encoder to generate a second image feature vector.
[0062] When processing negative sample groups, for example, the text data of the negative sample group retains the description "BRAF mutation treatment," but the gene data is replaced with the NRAS Q61K mutation sequence, the chemical structure is replaced with the MEK inhibitor trametinib, and the image is replaced with images of treatment-resistant tumor metastases. The second text feature vector generated by the text encoder still contains targeted therapy semantics, the second gene feature vector output by the gene encoder is converted into MAPK upstream regulatory features, the second chemical feature vector of the chemical encoder reflects MEK binding properties, and the second image feature vector of the image encoder shows the drug resistance phenotype.
[0063] Step 10255: Calculate the third intermodal similarity between the second text feature vector and the second gene feature vector in the cross-modal unified feature space, and calculate the fourth intermodal similarity between the second chemical feature vector and the second image feature vector.
[0064] For example, the description of "BRAF-specific inhibition" in the second text feature vector and the NRAS mutation feature in the second gene feature vector only have a modal similarity of 0.12, indicating that the gene mutation does not match the treatment mechanism; the similarity between the trametinib MEK inhibition property in the second chemical feature vector and the transfer feature in the second image feature vector is 0.25, reflecting the lack of effective correspondence between drug treatment and imaging phenotype.
[0065] Step 10256: Input the first intermodal similarity and the second intermodal similarity into an exponential function to generate positive sample similarity weights, and input the third intermodal similarity and the fourth intermodal similarity into a logarithmic function to generate negative sample similarity weights.
[0066] When generating similarity weights, a nonlinear transformation can be used to enhance discriminative power: the 0.85 similarity between the text-gene modality and the 0.78 similarity between the chemistry-image modality in the positive sample group are input into an exponential function to obtain a significantly increased positive sample similarity weight; simultaneously, the 0.12 and 0.25 similarity scores in the negative sample group are input into a logarithmic function to obtain a decayed negative sample similarity weight. This process ensures that the model focuses on positive sample features with strong intermodal correlations, while weakening interference signals from erroneous associations.
[0067] Step 10257: Construct a modality alignment loss function based on the difference between the positive sample similarity weight and the negative sample similarity weight, and update the parameters of the text encoder, gene encoder, chemical encoder and image encoder through the backpropagation algorithm, so that the distance between each modality feature vector of the positive sample group in the cross-modal unified feature space is less than the distance between each modality feature vector of the negative sample group.
[0068] When constructing the modality alignment loss function, parameter optimization can be driven by dynamic weight differences. The difference between the similarity weights of positive and negative samples is calculated as the loss signal, and the network parameters of each encoder are updated synchronously using the backpropagation algorithm. For example, when it is found that the gene encoder is insufficient in extracting features of kinase domain mutations, the convolution kernel weights are adjusted to enhance the detection sensitivity of codon mutation regions; when the chemical encoder fails to accurately capture molecular conformational changes, the edge feature aggregation strategy of the graph neural network is optimized. After multiple iterations, the Euclidean distance of the cross-modal feature vectors of the positive sample group in the unified space is reduced to below 0.3, while the feature distance of the negative sample group is expanded to above 1.2, achieving the optimization goal of cross-modal semantic alignment.
[0069] In one implementation, step 103, which involves dynamically fusing features in the cross-modal unified feature space based on task parsing instructions to generate a cross-modal feature vector matching the task parsing instructions, includes: Step 1031: Perform intent recognition on the task parsing instruction and extract key operators and constraints from the task parsing instruction; the key operators include at least one of data filtering, feature association, and experimental verification.
[0070] When performing intent recognition on task parsing instructions, the natural language understanding module can parse the user's input R&D goals. For example, when receiving the instruction "screen for dual kinase inhibitors targeting HER2-positive breast cancer with excellent membrane permeability," the system first uses dependency parsing to extract the core verb phrase "screen" as the key operator, identifying the dual requirements of data screening and feature association. The semantic role labeling module further parses out the constraints: the target type is limited to HER2-positive breast cancer-related kinases, the drug properties require membrane permeability, and the compound type is specified as a dual kinase inhibitor. The system verifies the validity of "dual kinase inhibitor" through the knowledge graph verification module, finding that this term specifically refers to a small molecule category that simultaneously inhibits the activity of HER2 and EGFR kinases, thereby establishing a set of boundary conditions for task execution.
[0071] Step 1032: Based on the type of the key operator, dynamically select feature subspaces of at least two modalities from the cross-modal unified feature space.
[0072] When dynamically selecting feature subspaces based on key operator types, the cross-modal association requirements of the current task can be considered. For the aforementioned dual kinase inhibitor screening task, the system prioritizes activating the feature subspace reflecting kinase binding pocket complementarity in the chemical structure modality, while simultaneously associating the feature subspace of HER2 gene amplification-related variant sites in the gene sequence modality. The feature selection engine simultaneously searches the knowledge graph for multimodal association paths related to "membrane permeability," determining the feature subspaces of cell membrane permeability detection images from the experimental image modality that need to be fused. Through a feature space projection algorithm, the system extracts three subspaces from the unified cross-modal feature space: molecular hydrophobicity distribution features from the chemical structure modality, HER2 expression regulation features from the gene modality, and lipid bilayer permeability features from the image modality, forming the basis for task-oriented feature combination.
[0073] Step 1033: Use a gated attention mechanism to assign weights to the feature subspaces of the at least two modalities to generate a modal weight coefficient matrix.
[0074] When using a gated attention mechanism for weight allocation, a dynamic mapping relationship between task semantics and feature subspaces can be established. The system concatenates the vector generated by semantically encoding the task parsing instructions with the chemical structure feature subspace to form a composite feature containing the semantics of "dual kinase inhibition". The fully connected layer performs a nonlinear transformation on this concatenated feature, outputting the initial attention score of the chemical structure modality for the current task. After Softmax normalization, a weight coefficient of 0.65 is obtained, indicating that molecular structure features play a dominant role in the screening. Similarly, the gene sequence modality and the experimental image modality obtain weight coefficients of 0.25 and 0.10, respectively, reflecting the auxiliary decision-making status of HER2 gene status and membrane permeability verification. After multiplying each modality feature subspace with its weight coefficient element-wise, the weighted chemical structure features highlight the stereocomplementarity of kinase binding sites, the weighted gene features strengthen the HER2 copy number variation signal, and the weighted image features retain information on key regions of cell membrane permeability.
[0075] Step 1034: Based on the modal weight coefficient matrix, linearly superimpose the feature subspaces of the at least two modalities to generate the cross-modal feature vector.
[0076] When performing linear superposition of features, cross-modal information fusion can be implemented according to the modal weighting coefficient matrix. The interatomic interaction feature vector (weighted by chemical structure mode), the HER2 exon variant feature vector (weighted by gene mode), and the membrane penetration trajectory feature vector (weighted by image mode) are aligned in dimension. After transforming the three to the same dimension using a feature space projection matrix, vector addition is performed. The resulting comprehensive feature vector simultaneously encodes the pharmacophore features required for kinase inhibition activity prediction, target-specific gene variant features, and cell membrane interaction features dependent on permeability assessment. This process ensures that the final output cross-modal feature vector retains key information from each modality while highlighting the contribution strength of task-related features.
[0077] Step 1035: Regularize the cross-modal feature vector according to the constraints so that the cross-modal feature vector meets the dimensional requirements of the task parsing instruction.
[0078] During regularization, the distribution characteristics of the feature vectors can be adjusted according to task constraints. For the permeability constraint, the system applies L2 regularization to the lipid-water partition coefficient-related dimensions in the feature space to suppress abnormal feature values caused by excessive modification of polar groups. Simultaneously, to meet the target selectivity requirements of dual kinase inhibition, Dropout technology is applied to the feature dimensions corresponding to the differing sites in the HER2 and EGFR kinase domains, enhancing the model's robustness in identifying key binding sites. The regularized cross-modal feature vectors not only meet the preset 512-dimensional feature space requirement, but their statistical distribution also better matches the expected distribution of input features by the virtual screening model, providing a standardized data foundation for subsequent calculations and predictions.
[0079] In a preferred implementation, step 1033, which involves using a gated attention mechanism to assign weights to the feature subspaces of the at least two modalities to generate a modal weight coefficient matrix, includes: Step 10331: Concatenate the semantic vector of the task parsing instruction with the feature subspace of each modality to generate a concatenated feature vector.
[0080] During semantic vector concatenation in step 10331, the abstract representation of the task instructions can be deeply integrated with the original feature subspace. Taking the task of "developing a third-generation kinase inhibitor to overcome the EGFR T790M resistance mutation" as an example, the system concatenates the "resistance mutation overcoming" topic vector generated by the semantic encoder with the molecular conformation feature subspace of the chemical structure modality to form extended features containing knowledge of the resistance mechanism. This concatenation operation realizes feature selection guided by task semantics in the vector dimension, enabling subsequent attention calculations to focus on molecular features related to the T790M mutation, such as the spatial accessibility of covalent binding sites.
[0081] Step 10332: Map the concatenated feature vectors using a fully connected layer to generate an initial attention score.
[0082] In step 10332, when generating the initial attention score, a multilayer perceptron can be used to mine cross-modal association patterns. When handling the task of developing Alzheimer's disease amyloid inhibitors, a fully connected network analyzes the spliced task semantics and gene sequence feature subspace, identifying a strong correlation between APP gene splicing site mutations and β-amyloid deposition. The network output layer activation function transforms this biological association into an initial attention score for the gene modality, reflecting the decision weight of gene data in target selection. Simultaneously, the features of β-sheet disruptors in the chemical structure modality are processed by the same network to obtain an initial score reflecting the small molecule intervention capability.
[0083] Step 10333: Normalize the initial attention score using the Softmax function to generate modal weight coefficients.
[0084] When performing Softmax normalization in step 10333, it is necessary to ensure that the weight allocation of multimodal features conforms to the probability distribution characteristics. When dealing with tasks involving the regulation of the tumor immune microenvironment, the system normalizes the attention scores of the tumor-infiltrating lymphocyte density features of the experimental image modality, the PD-L1 expression level features of the gene modality, and the checkpoint inhibitor features of the chemical structure modality. After Softmax transformation, the initial scores of the three are converted into modality weight coefficients of 0.50, 0.30, and 0.20, respectively, accurately quantifying the relative importance of each modality feature in the development of immunomodulators.
[0085] Step 10334: Multiply the modal weight coefficients element-wise with the corresponding feature subspace to generate weighted modal features.
[0086] In step 10334, when performing feature weighting, task-oriented feature enhancement needs to be achieved through element-wise multiplication. In the antiviral drug development scenario, the system multiplies the normalized weight coefficient of 0.70 with the protease inhibition feature subspace of the chemical structure modality, significantly improving the feature intensity of key pharmacophores. Simultaneously, the feature subspace of virus replication-related genes in the gene modality is scaled with a weight of 0.20, and the cytopathic effect features of the experimental image modality participate in the fusion with a weight of 0.10, forming a feature combination pattern that highlights the main mechanism of action while also considering auxiliary factors.
[0087] Step 10335: Input the weighted modal features into the residual network for feature enhancement to generate the modal weight coefficient matrix.
[0088] When using a residual network for feature enhancement in step 10335, skip connections can be used to preserve important information of the original features. When processing the task of developing antibiotic resistance reversal agents, the β-lactam ring modification features in the weighted chemical structure modality and the mecA gene expression features in the gene modality are processed by residual blocks. While extracting high-order cross-features, identity mapping preserves the electronic properties of the original molecular structure and the baseline expression level of gene regulation. This mechanism effectively prevents the loss of important features during nonlinear transformation, ensuring that the final generated modality weight coefficient matrix contains both deeply abstract task-related features and maintains the biological interpretability of the original data.
[0089] In one implementation, step 104, which involves calling a pre-trained multimodal large language model to perform task decomposition on the cross-modal feature vectors to generate at least one sequence of scientific research sub-tasks to be executed, includes: Step 1041: Input the cross-modal feature vector into the encoder layer of the multimodal large language model to generate a global context representation.
[0090] When inputting cross-modal feature vectors into the encoder layer of a multimodal large language model, the global dependencies between features are first captured through a multi-head self-attention mechanism. Taking the task of developing HER2-positive breast cancer targeted drugs as an example, the system inputs cross-modal feature vectors containing kinase domain features, HER2 gene amplification patterns, lead compound 3D conformations, and tumor cell proliferation image features into the encoder. The encoder's 12-layer Transformer structure extracts nonlinear relationships between features layer by layer. For example, it identifies the complementary relationship between the benzene ring substituent features of compounds and the shape of the HER2 kinase ATP-binding pocket, and simultaneously associates the causal relationship between gene amplification features and the increased mitotic phase features in cell images. Finally, it generates a global contextual representation vector that integrates multi-dimensional information. This vector encodes the knowledge associations throughout the entire process from molecular design to phenotypic validation.
[0091] Step 1042: Based on the scientific research objectives in the task parsing instruction, construct a task decomposition template; the task decomposition template includes a data acquisition stage, an experimental design stage, and a result verification stage.
[0092] When constructing the task decomposition template, the principles for dividing the research and development stages in the task parsing instructions can be followed. For the research objective of "developing an oral CDK4 / 6 inhibitor for the treatment of HR-positive breast cancer," the system retrieves standard drug development processes from the knowledge graph and constructs a three-stage template including the data acquisition stage (kinase activity compound library screening), the experimental design stage (in vitro kinase inhibitory activity assay), and the results validation stage (animal model efficacy evaluation). Each stage in the template is further refined into executable units. For example, the data acquisition stage is subdivided into three sub-task slots: structure-based virtual screening, drug-likeness prediction, and synthetic feasibility assessment, ensuring that the task decomposition conforms to drug development specifications.
[0093] Step 1043: In the decoder layer of the multimodal large language model, a pointer network is used to extract candidate subtasks that match the task decomposition template from the global context representation.
[0094] When using a pointer network to extract candidate subtasks, the global context representation is dynamically matched with the task decomposition template. In handling the development task of PD-1 / PD-L1 immune checkpoint inhibitors, the pointer network first locates the core mechanism of "blocking protein-protein interactions" in the global context, and then sequentially extracts candidate subtasks such as "monoclonal antibody humanization," "binding affinity assay," and "tumor-infiltrating lymphocyte detection" from the feature vector. Through an attention weight allocation mechanism, the network prioritizes subtasks directly related to the task objective, for example, assigning a higher selection probability to the "epitope affinity maturation" subtask while reducing the weight coefficient of "small molecule compound library screening," ensuring that the task sequence focuses on the biologics development pathway.
[0095] Step 1044: Perform logical relationship verification on the candidate subtasks to generate the research subtask sequence; the logical relationship verification includes time sequence constraints, data dependency constraints, and resource allocation constraints.
[0096] When performing logical relationship verification, a multi-dimensional constraint checking mechanism can be established. Taking the development of EGFR-TKI resistance reversal agents as an example, if the system detects that "construction of drug-resistant cell lines" should precede "compound sensitivity screening" in the candidate subtask sequence, violating the time order constraint, the task order will be automatically adjusted. Simultaneously, if the verification of "kinase activity detection" requires data dependency on "mutant EGFR purified protein," and a missing protein expression task is detected, this subtask will be automatically inserted. Resource allocation constraint checks reveal that "molecular dynamics simulation" requires GPU acceleration resources, but the current computing cluster load is too high; therefore, the system automatically adds a "computing resource reservation" subtask to ensure the executability of the task sequence.
[0097] Step 1045: Perform similarity matching between the research sub-task sequence and the historical task database, remove duplicate sub-tasks and add missing sub-tasks to generate an optimized research sub-task sequence.
[0098] When optimizing task sequences, intelligent corrections based on historical experience can be implemented. When processing PARP inhibitor development tasks, the system compares the generated candidate sub-task sequences with similar projects in the historical task library. It finds that the "BRCA mutation status verification" sub-task exists in 98% of historical projects, but is missing from the current sequence, and is automatically inserted. Simultaneously, it identifies functional overlap between "in vitro permeabilization experiment" and the "brain-like organ model testing" in historical tasks. After confirming an overlap of 85% through similarity calculation, the "in vitro permeabilization experiment" sub-task, which better meets the current task requirements, is retained, redundant items are removed, and a streamlined and optimized task sequence is formed.
[0099] In a preferred implementation, step 1044, which involves verifying the logical relationships between the candidate subtasks and generating the research subtask sequence, includes: Step 10441: Construct a knowledge graph in the field of biomedicine, wherein the knowledge graph includes entity nodes, relation edges and attribute constraints.
[0100] When constructing the biomedical knowledge graph in step 10441, multi-source heterogeneous data can be integrated to establish a semantic network. Taking the field of tumor immunotherapy as an example, the system creates entity nodes containing "PD-1 inhibitors," "T cell activation," and "tumor microenvironment," connected by relational edges such as "enhancement," "inhibition," and "regulation," and attaching attribute constraints such as "clinical trial stage ≥ Phase II" and "objective response rate > 20%." The entity nodes of the knowledge graph cover core elements such as target mechanisms, compound characteristics, and experimental methods. For example, the "CAR-T cell therapy" node is associated with the "cytokine release syndrome" side effect entity and labeled with the operational constraint "CRS prevention required."
[0101] Step 10442: Map the candidate subtasks to entity nodes in the knowledge graph to generate a set of subtask nodes.
[0102] In step 10442, when mapping candidate subtasks to the knowledge graph, semantic-level alignment can be used as an example. When processing the "develop bispecific antibody" task, the system maps the "epitope binding affinity optimization" subtask to the "antibody affinity maturation" node in the knowledge graph, associating it with the "phage display technology" experimental method entity. Simultaneously, the "in vivo efficacy evaluation" subtask is mapped to the "xenograft model" node, associating it with the "tumor volume measurement" data acquisition entity. During the mapping process, the system automatically identifies the "Fc segment engineering modification" subtask in the knowledge graph corresponding to the "antibody Fc receptor binding" node and inherits the "affects half-life" attribute constraint, ensuring the completeness and compliance of task elements.
[0103] Step 10443: Traverse the relation edges in the set of subtask nodes to verify whether the data flow between the candidate subtasks conforms to the attribute constraints; if there are subtask nodes that violate the attribute constraints, perform path replanning on the subtask nodes to generate corrected subtask nodes.
[0104] In step 10443, when verifying the data flow of subtasks, compliance checks can be performed based on the relational paths of the knowledge graph. Taking the development of ADC drugs as an example, the system detected that the critical path of "lysosomal cleavage" was missing between the "toxin small molecule conjugation" subtask node and the "endocytosis efficiency detection" subtask node, violating the mechanism of action constraints of antibody-drug conjugates. The system automatically inserted the "lysosomal enzyme activity assay" subtask node to reconstruct the "toxin release" data flow. At the same time, it was found that there was a circular dependency between the "drug loading assay" subtask and the "plasma stability detection" subtask. The path was adjusted to a sequential execution relationship through path replanning to ensure that the task sequence conforms to the principles of biopharmaceutics.
[0105] Step 10444: Remap the corrected subtask nodes to the task decomposition template and update the scientific research subtask sequence.
[0106] When updating the research sub-task sequence in step 10444, dynamic reconstruction and verification can be considered. When processing the influenza neutralizing antibody development task, the revised sub-task node set adds "pseudovirus neutralization experiment" and "mutant strain cross-protection test" nodes, which the system maps back to the data acquisition stage of the task decomposition template. The updated task sequence adds the "antibody Fc fragment glycosylation analysis" sub-task in the experimental design stage and the "rhesus monkey challenge experiment" node in the result verification stage, forming a complete R&D process that meets biosafety requirements. The system verifies the logical completeness of the updated sequence through knowledge graph reasoning, and after confirming that all critical paths satisfy the attribute constraints, outputs the final optimized research sub-task sequence.
[0107] In one implementation, step 105, which iteratively optimizes the multimodal research dataset based on the research sub-task sequence, includes: Step 1051: Input each subtask in the research subtask sequence into the reinforcement learning algorithm to generate an initial execution strategy.
[0108] In step 1051, when the research sub-task sequence is input into the reinforcement learning algorithm, an initial execution strategy can be generated through the policy network. Taking the development of a third-generation ALK kinase inhibitor as an example, the system constructs a state space for the "virtual screening" sub-task, including molecular descriptors of the compound library, three-dimensional coordinates of the kinase binding pockets, and historical activity data features. The action space is defined as a combination of molecular docking parameters, including the selection of sampling algorithm, scoring function weights, and the number of conformations generated. The policy network output recommends an initial execution strategy that combines a flexible docking mode with MM / GBSA binding free energy calculation. This strategy controls resource consumption while ensuring computational accuracy, laying the foundation for subsequent optimization.
[0109] Step 1052: Execute the initial execution strategy in a simulated experimental environment and collect multimodal feedback data during the strategy execution process.
[0110] When executing strategies in a simulated experimental environment, a multi-dimensional feedback acquisition channel can be constructed. When performing a kinase inhibition activity prediction task, the system simultaneously records the molecular docking time, conformational sampling coverage, and the deviation between the predicted results and the measured IC50 value. The experimental environment simulation module generates a virtual screening scenario containing 200 allosteric pockets for ALK mutants, automatically collecting multimodal feedback data such as computational resource utilization, number of key hydrogen bond formations, and hydrophobic contact area for each compound's docking process, forming a comprehensive evaluation system for coverage efficiency, accuracy, and cost.
[0111] Step 1053: Calculate the reward value for the multimodal feedback data to generate a strategy optimization gradient.
[0112] In step 1053, when calculating the reward value, a multi-factor contribution model can be established as an example. When processing the in vitro permeability prediction task, the system converts prediction accuracy, number of experimental replicates, and cell culture cost into standardized scores. A quality assessment function identifies incomplete data on transmembrane transport rates and applies a quality decay coefficient to data points that do not reach three replicates. A cost normalization function maps the high-throughput screening cost of the brain-like organoid model to a consumption index of 0.8, which, after priority weighting, generates an initial reward value, accurately quantifying the overall benefits of strategy execution.
[0113] Step 1054: Dynamically adjust the parameters of the reinforcement learning algorithm according to the optimization gradient of the strategy to generate an optimized execution strategy.
[0114] When adjusting reinforcement learning parameters in step 1054, a gradient-guided optimization mechanism can be considered. During the optimization of the lead compound toxicity prediction strategy, the strategy gradient algorithm identifies insufficient feature extraction of liver microsomal stability data and dynamically increases the weight coefficients of the metabolic site detection layer in the graph neural network. The parameter update direction calculation module, combined with the second derivative information of the reward value curve, adjusts the balance coefficient between exploration and utilization, ensuring that the optimized execution strategy maintains ADMET prediction accuracy while reducing sample consumption in in vitro hepatotoxicity experiments.
[0115] Step 1055: Compare the optimized execution strategy with the preset biomedical R&D indicators. If the termination condition is met, output the scientific research decision parameters; otherwise, re-execute the strategy optimization process.
[0116] When determining the termination condition in step 1055, multi-indicator joint verification is required. For the EGFR inhibitor synthesis route optimization task, the system compares the yield improvement, chiral purity achievement rate, and equipment downtime of the optimized strategy with preset 75% yield thresholds, 98% ee value standards, and 72-hour reaction cycles. When the improvement of three consecutive strategy iterations is less than 1%, the termination condition is triggered, and the final decision parameters are output, including recommended catalyst combinations, temperature gradient settings, and purification schemes, completing the closed loop from computational optimization to experimental implementation.
[0117] In one optional implementation, step 1053, which involves calculating the reward value from the multimodal feedback data and generating a policy optimization gradient, includes: Step 10531: Separate the experimental timestamp sequence, data integrity label, and resource consumption matrix from the multimodal feedback data.
[0118] For example, in processing tumor cell migration inhibition experiments, the system extracts experimental timestamp sequences from the raw logs and marks them as imaging detection time nodes every 24 hours. Data integrity tags record the cell count integrity status at each time point, and a resource consumption matrix statistically analyzes the time spent using the microscopic imaging equipment and the amount of culture medium consumed. This separation process ensures independent analysis of time-related parameters, quality indicators, and cost factors, providing structured input for subsequent refined reward calculations.
[0119] Step 10532: Input the experimental timestamp sequence into the time decay function to generate the experimental efficiency decay coefficient, input the data integrity label into the quality assessment function to generate the data quality score, and input the resource consumption matrix into the cost normalization function to generate the cost consumption index.
[0120] When generating the decay coefficient in step 10532, a time-value decay model can be applied. Analysis of the timestamp sequence of the kinase activity detection experiment revealed that the first three experimental cycles accounted for 80% of the total time. The time decay function applied an exponential decay weight to this, highlighting the potential for efficiency improvement in later experiments. The quality assessment function detected a lack of integrity in the SDS-PAGE bands of the mutant kinase purified samples, lowering the corresponding data quality score by 40%. The cost normalization function converted the combined cost of using the low-temperature centrifuge and fluorescence detector into a consumption index of 0.65, accurately reflecting the occupancy cost of high-value equipment.
[0121] Step 10533: Dynamically weight the experimental efficiency decay coefficient, data quality score, and cost consumption index according to the priority parameters in the task parsing instruction to generate an initial reward value.
[0122] For example, in an antiviral drug development task, the task parsing instruction assigns the highest priority parameter of 0.6 to the activity screening stage. The system sets the weights of the experimental efficiency decay coefficient to 0.4, the data quality score to 0.5, and the cost consumption index to 0.1. After weighted calculation, although a candidate compound exhibits a nanomolar EC50 value, its initial reward value is reduced by 15% due to the excessive time spent on purification steps, thus guiding the strategy optimization towards rapid purification technology.
[0123] Step 10534: Detect instrument error signals, data interruption events, and over-limit operation records in the multimodal feedback data, and generate an anomaly penalty coefficient based on the frequency of the instrument error signals, the duration of the data interruption events, and the severity level of the over-limit operation records.
[0124] When detecting abnormal events based on step 10534, a three-level severity classification system can be established. When analyzing feedback data from high-throughput sequencing experiments, the system identified that the frequency of flow cell bubble error signals exceeded the threshold, and an anomaly penalty coefficient of 0.3 was generated based on a frequency of 5 errors per gigabyte of data. For data interruption events caused by sample labeling errors, a penalty of 0.2 was applied based on the interruption duration of 3 consecutive experimental cycles. Events involving excessive centrifugation speed in the operation log were classified as Level 2 severity, with an additional penalty coefficient of 0.15, ensuring that the strategy optimization process proactively avoids experimental risks.
[0125] Step 10535: Subtract the abnormal penalty coefficient from the initial reward value to generate a deducted reward value, and input the deducted reward value into a reward smoothing function to eliminate noise fluctuations and generate a target reward value.
[0126] In step 10535, when calculating the deducted reward value, a noise filtering mechanism can serve as an auxiliary mechanism. The initial reward value of 0.85 for a certain antibody affinity maturation task fluctuates due to occasional ELISA plate edge effects. The reward smoothing function uses a moving average algorithm to eliminate random fluctuations, generating a stable target reward value of 0.82. This processing effectively distinguishes between the real benefit improvement brought about by strategy improvement and experimental random errors, preventing the reinforcement learning algorithm from overfitting to accidental optimization paths.
[0127] Step 10536: Arrange the target reward values in chronological order to generate a reward value sequence, and input the reward value sequence into a moving average window to generate the final reward value.
[0128] In step 10536, when generating the final reward value, a dynamic window adjustment technique can be used. When analyzing the reward value sequence of a kinase-selective optimization strategy over 10 consecutive rounds, the moving average window is automatically expanded to 5 periods to smooth short-term fluctuations and preserve long-term trends. If a certain generation of the strategy experiences a temporary decrease in reward value due to the introduction of a conformational entropy penalty term, the reward value will still maintain a positive gradient after window averaging, avoiding premature termination of potential optimization directions and ensuring the global optimization capability of the strategy evolution process.
[0129] Step 10537: Input the final reward value into the policy gradient algorithm to calculate the parameter update direction of the reinforcement learning algorithm and generate the policy optimization gradient.
[0130] In step 10537, when calculating the policy optimization gradient, a trust region constraint mechanism can be implemented. For parameter updates in the drug-likeness prediction model, the policy gradient algorithm, combined with the Fisher information matrix, determines the parameter update step size, improving prediction sensitivity while controlling the risk of overfitting. The gradient direction calculation module identifies insufficient contribution from features predicting bleeding-brain barrier penetration, and directionally enhances the extraction intensity of molecular polar surface area and hydrogen bond donor number features. This allows the optimized execution strategy to maintain Lipinski rule compliance while improving the success rate of central nervous system drug design.
[0131] In a non-restrictive implementation, the method further includes: Step 201: Extract the feature dimensions related to the scientific research objective analysis in the task parsing instruction from the cross-modal feature vector, and generate a feature dimension filtering mask.
[0132] In step 201, when the scientific research data fusion analysis system extracts relevant feature dimensions from cross-modal feature vectors, it uses an attention weight ranking mechanism to generate a feature dimension filtering mask. Taking the development of MET kinase inhibitors as an example, the system, based on the "selective inhibition of MET over VEGFR2" objective in the task parsing instruction, identifies dimensions 128-135 related to hydrogen bond formation in the kinase hinge region from the 1024 features of the chemical structure modality, and simultaneously filters dimensions 512-520 corresponding to MET exon 14 skipping mutations in the gene sequence modality. The feature dimension filtering mask generation module calculates the cosine similarity between each dimension and the task objective, marking dimensions with similarity exceeding a threshold as 1 and the rest as 0, forming a binary filtering template to accurately capture key features affecting kinase selectivity.
[0133] Step 202: Dimensionally filter the cross-modal feature vector according to the feature dimension filtering mask to generate a subset of contributing features.
[0134] In step 202, when the scientific data fusion and analysis system performs dimensional filtering, it retains a subset of contributing features through Hadamard product operations. When processing PD-1 / PD-L1 interaction inhibitor development tasks, the system applies a feature dimension filtering mask to the cross-modal feature vector, retaining key dimensions such as antibody complementarity-determining region conformation features, immune checkpoint gene expression features, and immune synapse image features. The filtered subset of contributing features eliminates molecular weight distribution features, gene intron region variation features, and cytoplasmic background image features unrelated to binding affinity, focusing the feature space on core dimensions related to antigen epitope recognition and immune regulation mechanisms.
[0135] Step 203: Perform reverse mapping between the contribution feature subset and the original data fragments in the multimodal scientific research dataset to locate the biomedical text data paragraphs, gene sequence data fragments, chemical structure data sub-graphs, and experimental image data regions corresponding to the contribution feature subset.
[0136] In step 203, when the scientific data fusion and analysis system performs reverse mapping, it uses a feature tracing algorithm to locate the original data fragments. In the task of optimizing EGFR-TKI resistance reversal agents, the system traces the 256th dimension feature in the contributing feature subset back to the conclusion paragraph in the biomedical text data that "T790M mutation leads to steric hindrance of the ATP binding pocket," and simultaneously associates it with the C>T mutation fragment at position chr7:55249071 in the gene sequence data. The chemical structure data sub-map locates the acrylamide covalently bound group in the molecular structure, and the experimental image data region is used to select the abnormal proliferation lesions of drug-resistant cell lines, forming a cross-modal data evidence chain.
[0137] Step 204: Tag the biomedical text data paragraphs with keywords, annotate the gene sequence data fragments with functions, label the chemical structure data sub-graphs with functional groups, and select the experimental image data regions with biomarkers.
[0138] In step 204, when the scientific data fusion analysis system performs multimodal data annotation, it can implement a domain knowledge-driven annotation strategy. When processing the ALK fusion gene detection task, the system annotates the biomedical text data segment with the mutation type keyword "EML4-ALK variant 3" and annotates the gene sequence data chr2:29415645-29416732 fragment with the function of "kinase domain breakpoint". The chemical structure data sub-graph marks the π-π stacking interaction region between the diaminopyridine ring and the ALK binding pocket in the crizotinib molecule, and the experimental image data box selects the breakpoint region where orange and green signals are separated in fluorescence in situ hybridization detection, establishing a traceable biomarker association system.
[0139] Step 205: Combine the labeled biomedical text data paragraphs, annotated gene sequence data fragments, labeled chemical structure data sub-graphs, and selected experimental image data regions according to modal type to generate a visual decision-making basis map.
[0140] In step 205, when the scientific research data fusion and analysis system constructs a visualized decision-making basis map, it can employ multi-view collaborative layout technology. For BRAF V600E mutant melanoma treatment decisions, the system arranges the annotated text paragraph "BRAF inhibitor resistance mechanism," the annotated gene fragment "V600E missense mutation," the labeled molecular structure "vemurafenib α-helix binding pattern," and the selected image "tumor-infiltrating lymphocyte distribution" according to modal type. The map achieves cross-modal association through spatiotemporal synchronization technology. When the user focuses on a certain chemical substructure, the corresponding gene mutation site and pathological image region are automatically highlighted, forming a three-dimensional evidence display network.
[0141] Step 206: Connect the visualization decision basis map with the key operators in the task parsing instructions to generate an interactive explanatory report.
[0142] In step 206, when the scientific data fusion and analysis system generates an interactive and interpretable report, it can establish dynamic semantic association channels. When optimizing neoadjuvant therapy for HER2-positive breast cancer, the system connects the "trastuzumab response prediction" key operator with the HER2 immunohistochemical scoring region, ERBB2 gene copy number fragment, and antibody Fc glycosylation characteristics in the visualized decision-making atlas. When the user clicks on the "pathological complete response rate" indicator, the report automatically expands with corresponding molecular subtyping text annotations, gene amplification heatmaps, and pre- and post-treatment MRI image comparisons, providing multi-level interpretations of the decision-making basis.
[0143] Step 207: Embed a parameter adjustment interface in the interactive interpretable report, and perform online calibration of the projection weights of the cross-modal unified feature space according to the feedback instructions input by the user through the parameter adjustment interface.
[0144] In step 207, when the scientific data fusion and analysis system embeds a parameter adjustment interface, it can be implemented based on a feedback-based online learning mechanism. When a user reviews the CDK4 / 6 inhibitor permeability prediction report, the projection weight of the lipophilic features in the chemical structure modality is increased by a slider. The system adjusts the contribution of the clogP-related dimension in the cross-modal unified feature space in real time. The adjusted feature space immediately recalculates the ranking of candidate compounds, improving palbociclib's predicted ranking by three places. Simultaneously, the corresponding molecular polarity surface area annotation and cell membrane permeability image bounding area in the interpretation report are updated, forming a dynamically optimized decision support closed loop.
[0145] Applying the above-described technical solution of this application's embodiments, firstly, through the joint embedding training mechanism of a multimodal encoder, a deep association mapping of cross-domain semantics is established in a unified feature space, enabling analyzable vectorized alignment of text semantics, gene expression patterns, chemical spatial topology, and image phenotypic features, breaking through the information barrier of traditional single-modal modeling. Secondly, a dynamic feature fusion engine is innovatively introduced, adaptively adjusting the contribution weights and combination paths of different modal features according to the research objectives, experimental constraints, and association rules of the task parsing instructions, thereby achieving demand-oriented representation reconstruction while maintaining the inherent correlation of data. Finally, through a task decomposition architecture driven by a large language model, complex research problems are transformed into interpretable sub-task execution chains, forming a closed-loop optimization mechanism of "global decision-local verification," significantly improving the efficiency and reproducibility of mining implicit research patterns.
[0146] In summary, the multi-level collaborative intelligent computing framework provided in this application can realize the entire process from scientific research data fusion to knowledge reasoning, effectively solving the systemic defects of traditional methods in cross-modal association modeling and dynamic decision optimization.
[0147] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements the scientific data fusion and analysis method based on a multimodal large model.
[0148] This invention provides a processor for running a program, wherein the program executes the scientific data fusion and analysis method based on a multimodal large model.
[0149] In embodiments of the present invention, such as Figure 2 As shown, the scientific research data fusion and analysis system 100 includes at least one processor 101, and at least one memory 102 and bus 103 connected to the processor 101; wherein, the processor 101 and the memory 102 communicate with each other through the bus 103; the processor 101 is used to call the program instructions in the memory 102 to execute the above-mentioned scientific research data fusion and analysis method based on multimodal large model.
[0150] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, scientific data fusion and analysis systems (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0151] In a typical configuration, a scientific research data fusion and analysis system includes one or more processors (CPUs), memory, and a bus. The system may also include input / output interfaces, network interfaces, etc.
[0152] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.
[0153] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage. Computer-readable storage media or any other non-transferable media can be used to store information that can be accessed by a scientific data fusion and analysis system. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0154] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or computer-readable storage medium that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or computer-readable storage medium. Unless otherwise specified, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or computer-readable storage medium that includes that element.
[0155] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A scientific research data fusion and analysis method based on a multimodal large model, characterized in that, include: Acquire a multimodal scientific research dataset, which includes biomedical text data, gene sequence data, chemical structure data, and experimental image data; Joint embedding training is performed on the multimodal scientific research dataset to generate a cross-modal unified feature space. The joint embedding training includes synchronous optimization of parameters of text encoder, gene encoder, chemical encoder and image encoder. Dynamic feature fusion is performed on the cross-modal unified feature space based on task parsing instructions to generate cross-modal feature vectors that match the task parsing instructions; the task parsing instructions include scientific research goal parsing, experimental condition constraints, and data association rules; The pre-trained multimodal large language model is invoked to perform task decomposition on the cross-modal feature vectors, generating at least one sequence of scientific research sub-tasks to be executed; The multimodal research dataset is iteratively optimized based on the research sub-task sequence to output research decision parameters that meet preset biomedical research and development indicators.
2. The method as described in claim 1, characterized in that, The step of jointly embedding and training the multimodal research dataset to generate a cross-modal unified feature space includes: The biomedical text data is semantically segmented to generate a text semantic block sequence, and a bidirectional attention mechanism is used to mine the context encoding corresponding to the text semantic block sequence. The gene sequence data is base-coded and mapped to generate gene fragment vectors, and local functional region features in the gene fragment vectors are extracted using a sliding window convolution kernel. The chemical structure data is transformed into a topological graph to generate an atomic node graph, and a graph neural network is used to aggregate the chemical bond connections in the atomic node graph. The experimental image data is segmented at multiple scales to generate a set of regional image patches, and a spatial attention mechanism is used to extract biomarker features from the set of regional image patches. The context encoding, local functional region features, chemical bond connections, and biomarker features are jointly trained by a contrastive learning loss function. This joint training is used to indicate that the projections of different modal data into the cross-modal unified feature space satisfy semantic alignment conditions.
3. The method as described in claim 2, characterized in that, The step of jointly training the context encoding, the local functional region features, the chemical bond connections, and the biomarker features using a contrastive learning loss function includes: Biomedical text data, gene sequence data, chemical structure data, and experimental image data of the same biological entity are extracted from the multimodal scientific research dataset as positive sample groups, and at least one modality of data is randomly replaced to generate negative sample groups. The biomedical text data of the positive sample group is input into the text encoder to generate a first text feature vector; the gene sequence data of the positive sample group is input into the gene encoder to generate a first gene feature vector; the chemical structure data of the positive sample group is input into the chemical encoder to generate a first chemical feature vector; and the experimental image data of the positive sample group is input into the image encoder to generate a first image feature vector. In the cross-modal unified feature space, the first intermodal similarity between the first text feature vector and the first gene feature vector is calculated, and the second intermodal similarity between the first chemical feature vector and the first image feature vector is calculated. The biomedical text data of the negative sample group is input into the text encoder to generate a second text feature vector; the gene sequence data of the negative sample group is input into the gene encoder to generate a second gene feature vector; the chemical structure data of the negative sample group is input into the chemical encoder to generate a second chemical feature vector; and the experimental image data of the negative sample group is input into the image encoder to generate a second image feature vector. In the cross-modal unified feature space, the third intermodal similarity between the second text feature vector and the second gene feature vector is calculated, and the fourth intermodal similarity between the second chemical feature vector and the second image feature vector is calculated. The first intermodal similarity and the second intermodal similarity are input into an exponential function to generate positive sample similarity weights, and the third intermodal similarity and the fourth intermodal similarity are input into a logarithmic function to generate negative sample similarity weights. A modal alignment loss function is constructed based on the difference between the positive sample similarity weight and the negative sample similarity weight. The parameters of the text encoder, gene encoder, chemical encoder and image encoder are updated through the backpropagation algorithm, so that the distance between each modal feature vector of the positive sample group in the cross-modal unified feature space is less than the distance between each modal feature vector of the negative sample group.
4. The method as described in claim 1, characterized in that, The step of dynamically fusing features in the cross-modal unified feature space based on task parsing instructions to generate a cross-modal feature vector matching the task parsing instructions includes: The task parsing instructions are subjected to intent recognition, and key operators and constraints are extracted from the task parsing instructions; the key operators include at least one of data filtering, feature association, and experimental verification. Based on the type of the key operator, dynamically select feature subspaces of at least two modalities from the cross-modal unified feature space; A gated attention mechanism is used to assign weights to the feature subspaces of the at least two modalities to generate a modal weight coefficient matrix. Based on the modality weight coefficient matrix, the feature subspaces of the at least two modes are linearly superimposed to generate the cross-modality feature vector; The cross-modal feature vector is regularized according to the constraints to ensure that the cross-modal feature vector meets the dimensional requirements of the task parsing instructions.
5. The method as described in claim 4, characterized in that, The step of employing a gated attention mechanism to assign weights to the feature subspaces of the at least two modalities, generating a modality weight coefficient matrix, includes: The semantic vector of the task parsing instruction is concatenated with the feature subspace of each modality to generate a concatenated feature vector; The concatenated feature vectors are mapped using a fully connected layer to generate an initial attention score; The initial attention score is normalized using the Softmax function to generate modal weight coefficients; The modal weight coefficients are multiplied element-wise with the corresponding feature subspace to generate weighted modal features; The weighted modal features are input into the residual network for feature enhancement to generate the modal weight coefficient matrix.
6. The method as described in claim 1, characterized in that, The pre-trained multimodal large language model is invoked to perform task decomposition on the cross-modal feature vectors, generating at least one sequence of scientific research sub-tasks to be executed, including: The cross-modal feature vectors are input into the encoder layer of the multimodal large language model to generate a global context representation; Based on the scientific research objectives in the task parsing instructions, a task decomposition template is constructed; the task decomposition template includes a data acquisition stage, an experimental design stage, and a result verification stage. In the decoder layer of the multimodal large language model, a pointer network is used to extract candidate subtasks that match the task decomposition template from the global context representation; Logical relationship verification is performed on the candidate subtasks to generate the research subtask sequence; the logical relationship verification includes time sequence constraints, data dependency constraints, and resource allocation constraints. The research subtask sequence is matched with the historical task database for similarity, duplicate subtasks are removed and missing subtasks are added to generate an optimized research subtask sequence.
7. The method as described in claim 6, characterized in that, The step of verifying the logical relationships between the candidate subtasks and generating the sequence of scientific research subtasks includes: Construct a knowledge graph in the field of biomedicine, wherein the knowledge graph includes entity nodes, relation edges and attribute constraints; The candidate subtasks are mapped to entity nodes in the knowledge graph to generate a set of subtask nodes; Traverse the relation edges in the set of subtask nodes to verify whether the data flow between the candidate subtasks conforms to the attribute constraints; If there are sub-task nodes that violate the attribute constraints, then the path of the sub-task node is replanned to generate a corrected sub-task node. The corrected subtask nodes are remapped to the task decomposition template, and the scientific research subtask sequence is updated.
8. The method as described in claim 1, characterized in that, The iterative optimization of the multimodal scientific research dataset based on the scientific research sub-task sequence includes: Each subtask in the research subtask sequence is input into a reinforcement learning algorithm to generate an initial execution strategy; The initial execution strategy is executed in a simulated experimental environment, and multimodal feedback data is collected during the strategy execution process; The reward value is calculated from the multimodal feedback data to generate a strategy optimization gradient; The parameters of the reinforcement learning algorithm are dynamically adjusted according to the optimization gradient of the strategy to generate an optimized execution strategy; The optimized execution strategy is compared with the preset biomedical research and development indicators. If the termination condition is met, the scientific research decision parameters are output; otherwise, the strategy optimization process is re-executed. The step of calculating the reward value from the multimodal feedback data and generating the policy optimization gradient includes: Separate the experimental timestamp sequence, data integrity label, and resource consumption matrix from the multimodal feedback data; The experimental timestamp sequence is input into a time decay function to generate an experimental efficiency decay coefficient, the data integrity label is input into a quality assessment function to generate a data quality score, and the resource consumption matrix is input into a cost normalization function to generate a cost consumption index. The experimental efficiency decay coefficient, data quality score, and cost consumption index are dynamically weighted according to the priority parameters in the task parsing instruction to generate an initial reward value. The instrument error signal, data interruption event, and over-limit operation record are detected in the multimodal feedback data. An anomaly penalty coefficient is generated based on the frequency of the instrument error signal, the duration of the data interruption event, and the severity level of the over-limit operation record. The abnormal penalty coefficient is subtracted from the initial reward value to generate a deducted reward value. The deducted reward value is then input into a reward smoothing function to eliminate noise fluctuations and generate a target reward value. The target reward values are arranged in chronological order to generate a reward value sequence, and the reward value sequence is input into a moving average window to generate the final reward value. The final reward value is input into the policy gradient algorithm to calculate the parameter update direction of the reinforcement learning algorithm, thereby generating the policy optimization gradient.
9. A scientific research data fusion and analysis system, characterized in that, It includes a processor and a memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the scientific research data fusion and analysis method based on a multimodal large model as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the scientific data fusion and analysis method based on a multimodal large model as described in any one of claims 1-8.