Birth defect full-period intelligent management system and method based on multi-modal large model
The multimodal large-scale intelligent management system for the entire life cycle of birth defects has solved the problems of data fragmentation and knowledge lag in birth defect prevention and control, and has achieved precise prevention and control throughout the entire life cycle, improving screening accuracy and prevention and control effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIVERSITY FIRST HOSPITAL (PEKING UNIVERSITY FIRST CLINICAL MEDICAL COLLEGE)
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot achieve full-cycle, precise prevention and control of birth defects, and suffer from problems such as data fragmentation, lagging knowledge updates, and limited screening accuracy.
The system employs a multimodal large-scale model-based intelligent management system for the entire lifecycle of birth defects. Through multimodal data acquisition and preprocessing, medical knowledge graph construction, and multimodal large-scale model analysis, it achieves unified data integration, real-time knowledge updates, and early and accurate identification of complex causes.
It has achieved unified integration of medical data at all stages of pregnancy, childbirth, and postpartum, and incorporated the latest clinical research findings in a timely manner, thereby improving the accuracy of screening and prevention and control, and generating quantitative risk level information and personalized prevention and control guidance.
Smart Images

Figure CN122050804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of artificial intelligence, medical informatics, and public health, and in particular to an intelligent management system and method for the entire life cycle of birth defects based on a multimodal large model. Background Technology
[0002] Birth defects are a major cause of developmental disorders, chronic diseases, disabilities, and even death in children, and have become a significant global public health issue. The current technical challenge is how to achieve full-cycle prevention and control of birth defects, namely, continuous monitoring and intervention at all stages from preconception, prenatal, to postnatal, in order to reduce incidence and improve health outcomes.
[0003] To address these issues, current technologies primarily rely on manual screening and decentralized, single-disease detection tools, such as clinical assessment methods based on traditional guidelines and manuals, and independent detection systems for specific birth defects. These technologies attempt to identify risk factors and provide basic prevention and control recommendations through the manual collection and analysis of medical data; some systems also utilize electronic health records or imaging data for preliminary analysis.
[0004] However, these existing technologies have significant drawbacks: data fragmentation makes it difficult to unify and integrate medical data from preconception, prenatal, and postnatal stages, resulting in insufficient cross-institutional collaboration; knowledge updates lag behind, with traditional guidelines and manuals unable to quickly incorporate the latest clinical research findings; and screening accuracy is limited, with single data sources and manual methods failing to achieve early and accurate identification of complex causes. These shortcomings limit the efficiency and accuracy of birth defect prevention and control.
[0005] Therefore, there is an urgent need for an intelligent system that integrates multimodal data, possesses medical knowledge understanding and real-time decision support capabilities, to achieve precise prevention and control of birth defects throughout the entire life cycle, in order to overcome the shortcomings of existing technologies and improve the level of public health management. Summary of the Invention
[0006] The technical problem to be solved by this invention is to address the shortcomings of existing technologies, specifically by providing a multimodal large-scale model-based intelligent management system and method for the entire life cycle of birth defects, as detailed below: 1) In a first aspect, the present invention provides an intelligent management system for the entire life cycle of birth defects based on a multimodal large model, the specific technical solution of which is as follows: It includes a multimodal data acquisition and preprocessing module, a medical knowledge graph construction module, and a multimodal large model analysis module; The multimodal data acquisition and preprocessing module is used to: acquire multimodal data related to birth defects of preset users from multiple data sources, and preprocess the acquired multimodal data; The medical knowledge graph construction module is used to: construct medical knowledge graphs based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects; The multimodal large model analysis module is used to: process preprocessed multimodal data using a multimodal large model and in conjunction with a medical knowledge graph, generate risk level information to characterize the likelihood of birth defects, and obtain prevention and control guidance suggestions based on the risk level information.
[0007] The beneficial effects of the intelligent management system for the entire life cycle of birth defects based on a multimodal large model provided by this invention are as follows: The multimodal data acquisition and preprocessing module achieves unified integration of medical data across preconception, prenatal, and postnatal stages, effectively solving the data fragmentation problem in existing technologies. The medical knowledge graph construction module, based on etiology, genetics, treatment guidelines, and research findings, constructs a dynamic knowledge graph that can promptly incorporate the latest clinical research results, overcoming the limitations of traditional guidelines' outdated updates. The multimodal large-scale model analysis module utilizes a multimodal large-scale model to fuse and analyze clinical text, medical images, genomics data, and environmental factor data. Through cross-modal feature extraction and joint reasoning, it achieves early and accurate identification of complex etiologies, significantly improving screening accuracy. The risk level information generated by the system provides a quantitative basis for birth defect prevention and control. The prevention and control guidance recommendations generated based on this information ensure both assessment accuracy and support for personalized intervention. The overall system, through the collaborative work of multiple modules, establishes a complete technical path from data acquisition to risk assessment and intervention recommendations, providing an effective intelligent solution for precise prevention and control of birth defects throughout the entire lifecycle.
[0008] Based on the above scheme, the intelligent management system for the entire life cycle of birth defects based on a multimodal large model of the present invention can be further improved as follows.
[0009] Furthermore, it also includes an intelligent interaction and decision support module, which is used to: generate health management suggestions based on risk level information and prevention and control guidance, and adaptively adjust the expression of health management suggestions according to the user role of the preset user.
[0010] The beneficial effects of adopting the above-mentioned further solutions are: health management recommendations are generated based on risk level information and prevention and control guidance, ensuring that prevention and control measures can be transformed into concrete and actionable management plans. The function of adaptively adjusting the expression format according to user roles allows doctors to receive professional and detailed treatment advice, while patients receive easily understandable health guidance, significantly improving the acceptability and operability of the information. This role-adaptive interaction method enhances the system's universality and professionalism, meeting the needs of medical professionals for accurate information while also considering the comprehension abilities of ordinary users. By optimizing the information presentation method, this module promotes the effective implementation of prevention and control measures, strengthens communication between doctors and patients, and provides more humanized and efficient technical support for the whole-cycle prevention and control of birth defects.
[0011] Furthermore, it also includes a closed-loop management and continuous evolution module, which is used to analyze health management recommendations and subsequently collected user multimodal data, and feed the analysis results back to the multimodal large model analysis module and the medical knowledge graph construction module to update the multimodal large model and the medical knowledge graph.
[0012] The beneficial effects of adopting the above-mentioned further approach are as follows: By analyzing health management recommendations with subsequently collected multimodal user data, the system can assess the actual effectiveness of prevention and control measures. The analysis results are fed back to the multimodal large-scale model analysis module, enabling the model to continuously optimize parameters based on real-world application data, thus improving the accuracy of risk assessment. Simultaneously, feedback is also provided to the medical knowledge graph construction module, promoting the timely updating and improvement of the knowledge graph content. This continuous evolution mechanism ensures that the system can adapt to new medical discoveries and clinical practices, maintaining the timeliness and scientific rigor of prevention and control guidance recommendations. Through continuous iteration and updates, the system can continuously improve its performance during long-term use, providing more accurate and reliable technical support for birth defect prevention and control.
[0013] Furthermore, multimodal data associated with birth defects include clinical texts, medical images, genomic data, and environmental factor data.
[0014] The beneficial effects of adopting the above-mentioned further approach are as follows: by defining the multimodal data associated with birth defects as clinical text, medical imaging, genomics data, and environmental factor data, comprehensive coverage of risk factors is achieved. This multi-source data integration can comprehensively utilize multi-dimensional information such as text descriptions, image features, genetic information, and environmental exposure, effectively compensating for the limitations of a single data source. By fusing different modalities of data, the system enhances the comprehensiveness and depth of etiological analysis, supports more accurate risk assessment and early identification, thereby improving the accuracy and reliability of birth defect prevention and control, and providing a more scientific data foundation for full-cycle management.
[0015] 2) Secondly, the present invention also provides a method for intelligent management of birth defects throughout the entire life cycle based on a multimodal large model, the specific technical solution of which is as follows: Multimodal data on birth defects associated with a preset number of users are collected from multiple data sources, and the collected multimodal data is preprocessed. A medical knowledge graph was constructed based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects. By using a multimodal large model and combining it with a medical knowledge graph, the preprocessed multimodal data is processed to generate risk level information to characterize the likelihood of birth defects, and prevention and control guidance suggestions are obtained based on the risk level information.
[0016] Based on the above scheme, the intelligent management method for the entire life cycle of birth defects based on a multimodal large model of the present invention can be further improved as follows.
[0017] Furthermore, it also includes: generating health management suggestions based on risk level information and prevention and control guidance, and adaptively adjusting the expression of health management suggestions according to the user role of the preset user.
[0018] Furthermore, it also includes: analyzing health management recommendations with subsequently collected user multimodal data, and updating the multimodal big data model and medical knowledge graph based on the analysis results.
[0019] Furthermore, multimodal data associated with birth defects include clinical texts, medical images, genomic data, and environmental factor data.
[0020] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so that the electronic device realizes any of the above-mentioned intelligent management method for the whole life cycle of birth defects based on a multimodal large model.
[0021] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned intelligent management methods for the entire life cycle of birth defects based on a multimodal large model.
[0022] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below: Figure 1 This is a schematic diagram of the structure of a multimodal large model-based intelligent management system for the entire life cycle of birth defects according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for intelligent management of birth defects throughout the entire lifecycle based on a multimodal large model, according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0024] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0025] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0026] like Figure 1 As shown in the figure, an intelligent management system for the entire life cycle of birth defects based on a multimodal large model according to an embodiment of the present invention includes: It includes a multimodal data acquisition and preprocessing module, a medical knowledge graph construction module, and a multimodal large model analysis module; The multimodal data acquisition and preprocessing module is used to: acquire multimodal data related to birth defects of preset users from multiple data sources, and preprocess the acquired multimodal data; Multimodal data associated with birth defects include clinical text, medical imaging, genomics data, and environmental factors. Specifically, clinical text refers to written records generated during the medical process, including diagnostic descriptions, medical orders, surgical reports, nursing notes, and follow-up summaries in electronic medical records. This textual data is usually in unstructured or semi-structured form, recording patients' symptoms, signs, treatment history, and efficacy assessments, and is an important source of information for birth defect risk assessment. Medical imaging refers to images of internal human structures obtained through medical imaging equipment, such as ultrasound images, magnetic resonance imaging (MRI) images, computed tomography (CT) scans, and X-rays. This image data is stored in digital format and can visually display fetal development, organ morphological abnormalities, or structural defects, providing visual evidence for the early identification of birth defects. Genomics data refers to genetic information obtained through gene sequencing technology, including DNA sequences, single nucleotide polymorphisms (SNPs), copy number variations, and gene expression profiles. This data reflects an individual's genetic background and potential mutation sites and can be used to analyze the genetic causes and familial risk of birth defects. Environmental data refers to physicochemical parameters related to the pregnant woman's living environment, such as air pollutant concentrations, water quality indicators, radiation exposure levels, occupational hazard exposure records, and nutritional status information. This data is collected through environmental monitoring networks or questionnaires to help assess the impact of the external environment on fetal development and its association with birth defects.
[0027] The process involves collecting multimodal data related to birth defects of a preset user from multiple data sources and preprocessing the collected multimodal data. The specific implementation process is as follows: 1) The multimodal data acquisition and preprocessing module systematically collects multimodal data related to birth defects from multiple data sources for predefined users through standardized interfaces. Data sources include hospital information systems, electronic medical record systems, medical imaging archive systems, gene testing institution databases, and environmental monitoring platforms. The acquisition process uses standard medical information exchange interfaces such as HL7 or FHIR to ensure seamless cross-platform data access. For clinical text data, the system extracts text information such as diagnostic records, medical notes, and laboratory reports from electronic medical records; for medical imaging data, the system obtains image files such as ultrasound and MRI from imaging archives and communication systems; for genomics data, the system downloads gene sequencing results and variant annotation files from gene testing laboratories through a secure application programming interface; for environmental factor data, the system collects records such as air pollution index, radiation dose, and chemical exposure history from public health databases or IoT sensors. All collected data is encrypted using transport layer security protocols during transmission to prevent unauthorized access.
[0028] 2) The preprocessing stage begins with cleaning and standardizing the multimodal data. Clinical text data undergoes word segmentation, entity recognition, and standardized encoding using natural language processing techniques, converting it into a unified structured format, such as mapping diagnostic terms to the International Classification of Diseases (ICD) standard code. Medical imaging data is denoised, normalized, and format-converted using image processing algorithms to ensure all images have consistent resolution and color space. Genomics data undergoes quality control and sequence alignment using bioinformatics tools, removing low-quality reads and converting it to a standard variant calling format. Environmental factor data undergoes unit unification and outlier removal using data integration methods, such as converting pollutant concentrations from different monitoring stations to a unified unit of measurement. Next, the system performs de-identification processing, removing all direct personal identifiers such as names and ID numbers, and using differential privacy or hash encryption techniques to de-identify indirect identifiers. Finally, the preprocessed multimodal data is stored in a secure database and shared across institutions through a federated learning framework, ensuring data privacy while supporting subsequent analysis tasks.
[0029] The medical knowledge graph construction module is used to construct a medical knowledge graph based on the etiology, genetics, treatment guidelines, and research findings in the field of birth defects. The specific implementation process is as follows: 1) In the knowledge and data collection phase, etiological, genetic, diagnostic and treatment guidelines, and research findings related to birth defects are obtained from authoritative sources. Etiological data comes from medical textbooks, academic journals, and clinical research databases, covering causes such as environmental factors, sources of infection, and maternal health conditions. Genetic data is extracted from gene banks and variant databases, including gene mutation sites, inheritance patterns, and information on related syndromes. Diagnostic and treatment guidelines data comes from official documents issued by domestic and international public health institutions, such as prenatal screening guidelines and neonatal disease diagnostic criteria. Research findings data are collected through academic paper databases and clinical trial registration platforms, involving the latest research findings and evidence-based medicine. All data is downloaded in batches via application programming interfaces or web crawling technology and stored in a unified format in a temporary knowledge base.
[0030] 2) In the knowledge extraction and cleaning stage, natural language processing (NLP) techniques are used to structure the raw data. For text-based etiological descriptions and treatment guidelines, named entity recognition (NID) algorithms are used to identify medical entities, such as disease names, symptoms, drugs, and examination items. Relationship extraction models extract associations between entities from the literature, such as the causal relationship between a gene mutation and a specific birth defect. For genetic data, bioinformatics tools parse gene sequence annotation files, converting variation information into standardized genotype-phenotype association records. The data cleaning process removes duplicate entries and contradictory information, and standardizes terminology mapping, such as unifying synonyms to Medical Subject Headings (MTBG) codes.
[0031] 3) In the knowledge fusion and modeling phase, the extracted entities and relationships are integrated into a graph structure. Entities serve as nodes, including diseases, genes, environmental factors, diagnostic methods, and treatments; relationships serve as edges, including causal relationships, associative relationships, treatment responses, and genetic patterns. Graph neural network technology is used to learn node embedding representations, capturing complex semantic relationships between entities. The knowledge fusion process resolves conflicts between multi-source data, for example, by fusing recommendations from different guidelines using confidence weighting. After modeling is complete, the knowledge graph is stored in a resource description framework format, defining a unified semantic pattern to ensure logical consistency.
[0032] 4) In the knowledge storage and indexing phase, the constructed knowledge graph is imported into the graph database system, and an efficient query interface is established. The graph database uses a native graph storage engine, supporting complex path queries and real-time traversal. To improve access performance, the system creates indexes for commonly used query patterns, such as quickly retrieving relevant nodes by disease name or gene symbol. Simultaneously, the knowledge graph is exposed to the multimodal large-scale model analysis module through a web service interface, supporting graph-based reasoning requests.
[0033] 5) The dynamic update and optimization phase ensures the continuous evolution of the knowledge graph. The system regularly collects incremental data from newly released research findings and guideline updates, automatically extracting entities and relationships using natural language processing technology. The graph neural network model incrementally learns new knowledge, adjusting node embeddings and topology. The knowledge graph version management mechanism records historical changes and maintains data integrity through consistency checking algorithms. The optimization process includes query performance tuning and semantic enrichment, such as adding time attributes to track the history of knowledge evolution.
[0034] In another feasible approach, the process of acquiring a medical knowledge graph is as follows: 1) Construct a multi-source heterogeneous data acquisition framework for the field of birth defects. This framework integrates multimodal data sources such as electronic medical record systems, genomic databases, environmental exposure registries, and international treatment guidelines. A distributed crawler architecture is employed to achieve parallel data acquisition, while blockchain technology is introduced to ensure the traceability and integrity of data sources. Time series analysis is used to identify the optimal time window for data acquisition, ensuring coverage of the entire data flow from preconception counseling and prenatal screening to postpartum follow-up. This data acquisition framework adopts a layered architecture design, including four main components: a data source access layer, a distributed acquisition layer, a quality control layer, and a temporary storage layer, to achieve systematic acquisition and integration of multimodal data. Specifically: At the data source access layer, the framework connects to various data sources through standardized interface protocols. For hospital electronic medical record systems, the HL7FHIR standard interface is used to extract clinical text data, including diagnostic records, laboratory reports, and medication history. Genomic databases are accessed through bioinformatics toolkits such as BioPython to obtain gene variation data and expression profile data. Environmental exposure registries collect environmental factor data such as air quality and water quality monitoring through IoT device interfaces. International clinical practice guidelines databases synchronize the latest versions of clinical practice guidelines and expert consensus through API interfaces. Each data source is configured with a dedicated adapter to handle the conversion of different data formats and transmission protocols. The distributed acquisition layer adopts a master-slave crawler architecture, consisting of a scheduling node and multiple acquisition nodes. The scheduling node allocates acquisition tasks using a consistent hashing algorithm to achieve load balancing. Acquisition nodes are deployed on servers in different geographical locations to execute data crawling tasks in parallel. For dynamically updated data such as clinical practice guidelines, an incremental acquisition strategy is adopted, identifying new content based on version numbers and timestamps. For large-scale genomic data, a fragmented transmission mechanism is implemented, dividing the large dataset into multiple data blocks for parallel transmission. The acquisition process adopts an asynchronous communication mode, using message queues to buffer data transmission and avoid system overload. The framework incorporates blockchain technology to build a data traceability system. Each data collection transaction generates a blockchain record containing a data hash value, timestamp, data source identifier, and collection node information. The hash value is calculated using the SHA-256 algorithm. in, Represents data content, Represents a timestamp. These records represent data source identifiers. They form a distributed ledger deployed across multiple verification nodes to ensure the transparency and immutability of the data collection process.
[0035] The quality control layer implements data integrity verification and optimizes the timing of data collection. The integrity verification module calculates the checksum of the collected data and compares it with the hash value stored on the blockchain to ensure data transmission integrity. The time series analysis module uses Fourier transform to detect the periodicity of data updates based on historical collected data, identifying the optimal collection time window. Differentiated collection strategies are established for the data characteristics of different stages: pre-pregnancy, prenatal, and postpartum. For example, genetic screening data is prioritized in early pregnancy, ultrasound imaging is strengthened in mid-pregnancy, and newborn follow-up data is emphasized in the postpartum stage. The temporary storage layer uses a distributed file system to organize the collected raw data. Data is partitioned and stored according to source type and time dimension, and multi-level indexes are established to support fast retrieval. The storage system retains the original data format and metadata information, providing a complete data foundation for subsequent preprocessing stages. Access control mechanisms are also implemented to ensure the security of sensitive medical data during storage. The entire collection framework displays the connection status, collection progress, and data quality indicators of each data source in real time through a monitoring panel, supporting comprehensive monitoring and scheduling optimization of the collection process by administrators. The framework also provides a configuration interface, allowing adjustment of the collection frequency and priority of each data source according to actual needs, ensuring data collection efficiency and coverage.
[0036] 2) Establish a preprocessing pipeline based on multimodal data fusion. Utilize natural language processing (NLP) technology for deep semantic analysis of text data and employ computer vision algorithms to process key features in medical image data. Develop a dedicated terminology standardization engine to map medical terms from different sources to a unified ontology system. Apply differential privacy technology to de-identify sensitive genetic information, ensuring data usability while meeting privacy protection requirements. This preprocessing pipeline adopts a modular design, comprising four core components: a text processing module, an image processing module, a terminology standardization module, and a privacy protection module, enabling systematic preprocessing of multimodal data. Specifically: In the text processing module, clinical text data first undergoes word segmentation and part-of-speech tagging, using a word segmentation algorithm enhanced with a medical dictionary to accurately segment medical terms. Subsequently, a named entity recognition model is employed to identify medical entities such as disease names, drug names, anatomical locations, and examination items within the text. Entity linking technology links the identified entities to corresponding concepts in a standard medical knowledge base. For deep semantic parsing, a pre-trained language model is used to generate semantic vector representations of the text, capturing implicit information and contextual relationships in clinical descriptions. For long medical record documents, text structuring techniques are used to automatically extract key information fragments and populate them into standardized medical record templates. The image processing module specifically handles medical image data, including ultrasound, MRI, and CT images. Image quality enhancement is performed first, using an adaptive histogram equalization algorithm to improve image contrast and a non-local means denoising algorithm to reduce image noise. Then, a convolutional neural network is used for key feature extraction. This network, pre-trained on a multi-center medical image dataset, is capable of identifying birth defect-related features such as fetal structural abnormalities and organ development indicators. Dedicated feature extraction network branches are designed for different image modalities to ensure effective feature extraction for various medical images. The extracted features include morphological features, texture features, and deep learning features, which are encoded into fixed-dimensional feature vectors for subsequent multimodal fusion. The terminology standardization module constructs a unified medical terminology ontology system. This module first establishes a basic terminology library, integrating standard medical terminology systems such as the International Classification of Diseases, Medical System Nomenclature, and Medical Subject Headings. A terminology mapping algorithm is developed, based on word embedding technology and edit distance calculation, to automatically map medical terms from different data sources to standard terms. For synonyms and polysemous words, a context-aware disambiguation algorithm is used to determine the most appropriate standard term based on the context in which the term appears. The mapping process preserves the hierarchical structure and relationships of terms to ensure semantic integrity. The mapping results are manually reviewed and validated using machine learning models to continuously improve mapping accuracy. The privacy protection module focuses on desensitizing sensitive genetic information. Differential privacy technology is used to add precisely calibrated noise to the gene sequence data. The noise addition follows a strict mathematical definition: in, This indicates query results that satisfy differential privacy. This represents the original query function. Indicates Gaussian noise. Carefully designed according to a privacy budget ε, sensitive variant sites in genomic data are replaced with frequency ranges using generalization techniques. A k-anonymity-based population privacy protection method was also developed to ensure that any single record is indistinguishable from the other k-1 records in the dataset. Privacy protection strength is evaluated using quantitative indicators, establishing an adjustable balance between data availability and privacy protection. The processed data from each module is ultimately integrated into a unified data representation layer. Text semantic vectors, image feature vectors, standardized terms, and desensitized genetic data are organized into structured data objects, preserving the semantic information and cross-modal associations of the original data. A data quality assessment component verifies the integrity, consistency, and accuracy of the preprocessed data to ensure it meets the input requirements for subsequent multimodal large-scale model analysis. The entire preprocessing pipeline coordinates the execution order and data flow of each module through a workflow engine, achieving efficient and reliable multimodal data preprocessing.
[0037] 3) Design a multi-task joint learning model to simultaneously perform medical entity recognition, relation extraction, and confidence assessment. This model integrates graph convolutional networks and attention mechanisms to capture complex semantic relationships between entities, such as gene-environment interactions and drug-phenotype associations. A reinforcement learning strategy is introduced to dynamically optimize the decision-making process for entity boundary recognition and relation classification, improving the accuracy and recall of knowledge extraction. The multi-task joint learning model adopts a unified architecture to simultaneously perform the three core tasks of medical entity recognition, relation extraction, and confidence assessment. The model is designed based on a shared encoder and task-specific decoder, achieving parameter sharing and cross-task information interaction, effectively capturing complex semantic relationships in medical data. Specifically: The model inputs preprocessed multimodal data, including clinical text, genomics data, and environmental factor data. A shared encoder uses multi-layer Transformer modules to extract contextual representations of the text sequences; each Transformer layer contains a self-attention mechanism and a feedforward neural network. The self-attention mechanism calculates the association weights between the query vector, key vector, and value vector, expressed by the following formula: in, , , These represent the query, key, and value matrices, respectively. The dimension is vector. For non-textual data, such as gene sequences and environmental parameters, an embedding layer is used to convert them into dense vectors, which are then concatenated with textual features to form a unified input representation. The encoder outputs a shared feature vector, capturing the deep semantic information of the input data.
[0038] The medical entity recognition task uses Conditional Random Field (CRF) layers to perform sequence labeling on shared features, identifying medical entities such as disease names, gene loci, and environmental factors. The entity recognition decoder receives the feature sequence output from the encoder, captures contextual dependencies through a bidirectional Long Short-Term Memory (LSTM) network, and outputs an entity type label for each location. Simultaneously, a graph convolutional network processes the structural relationships between entities; the graph convolution operation is defined as: in, This indicates adding a self-join to the adjacency matrix. For degree matrix, It is the first Layer node features This is a trainable weight matrix. An attention mechanism weights important entity features and calculates correlation scores between entities, improving boundary recognition accuracy.
[0039] The relation extraction task constructs entity pairs based on entity recognition results and classifies relation types. The relation extraction decoder uses a graph attention network to model the entity graph structure, where nodes represent entities and edges represent potential relations. The graph attention layer calculates the attention coefficients between node pairs. in, and For node features, This is the weight matrix. This serves as the attention vector. For complex relationships such as gene-environment interactions and drug-phenotype associations, a multi-head attention mechanism is used to capture multi-dimensional interaction patterns. The relationship classifier outputs the relationship type and its corresponding probability distribution.
[0040] The confidence assessment task predicts a confidence score for each entity and relation. The confidence module receives intermediate features from entity recognition and relation extraction, and outputs a confidence value between 0 and 1 through a fully connected layer and a sigmoid function. in, For feature vectors, and These are the training parameters. Confidence scores are calculated by fusing model uncertainty and data consistency features to filter for highly reliable knowledge units.
[0041] Reinforcement learning strategies dynamically optimize the decision-making process for entity boundary recognition and relation classification. The reinforcement learning agent takes the current decoding state as input and outputs either an entity boundary adjustment or relation classification action. The state space includes the encoder's hidden state and partial decoding results, while the action space covers entity boundary movement and relation type changes. The reward function is designed based on the final accuracy and recall. in, and To balance hyperparameters, the policy gradient method updates model parameters: in, For policy networks, The model uses cumulative rewards. Through iterative training, it adaptively adjusts its decision-making strategy, reducing error propagation and improving the accuracy and recall of knowledge extraction.
[0042] The model training employs a multi-task loss function to jointly optimize entity recognition, relation extraction, and confidence assessment objectives. in, For sequence labeling cross-entropy loss, For relation classification, use binary cross-entropy loss. To assess the mean squared error loss for confidence level, These represent task weight coefficients. The training process utilizes backpropagation and the Adam optimizer, with model performance monitored through a validation set to achieve end-to-end learning. The entire implementation ensures the model can efficiently process multimodal medical data and accurately extract structured knowledge units.
[0043] 4) Create a self-evolving medical knowledge graph architecture, employing a multi-level graph structure to store basic medical knowledge, clinical practice evidence, and the latest research findings. Develop a distributed graph update mechanism based on federated learning to support collaborative improvement of the knowledge graph by multiple medical centers without data leaving their domains. Introduce a graph neural network inference engine to automatically discover potential knowledge relationships and verify the logical consistency of new knowledge, achieving continuous intelligent evolution of the knowledge graph. This medical knowledge graph architecture adopts a multi-level graph structure design, dividing medical knowledge into three independent but interconnected levels: the basic medical knowledge layer, the clinical practice evidence layer, and the latest research findings layer. The basic medical knowledge layer stores verified medical theories and standard terminology, including basic concepts and their relationships such as disease classification, anatomical structures, and physiological processes. The clinical practice evidence layer includes actual case data, treatment plans, and efficacy evaluations from multiple medical centers, with each evidence node accompanied by a credibility weight and timestamp information. The latest research findings layer integrates cutting-edge findings from academic journals, clinical trials, and conference papers, maintaining the timeliness of knowledge through a dynamic update mechanism. Each level is linked through cross-level connections to form a complete knowledge system, specifically: A distributed graph update mechanism based on federated learning enables collaborative knowledge evolution among multiple centers. Each participating medical center deploys a copy of the knowledge graph and a model training environment locally, using local data to train the graph neural network model. The federated learning server periodically collects model parameter updates from each center and aggregates the global model using a federated averaging algorithm. in, This represents the model parameters of the k-th center in round t. For the amount of data in this center, The total data volume is [amount to be filled in]. The aggregated global model is distributed to various centers to achieve knowledge sharing without exposing the original data. Differential privacy technology is used to add noise during the update process to ensure privacy and security. A version control system is also established to record the evolution history of the knowledge graph, supporting backtracking and verification.
[0044] The graph neural network inference engine achieves knowledge discovery and verification through a multi-step process. The engine first performs representation learning on the knowledge graph, using a graph attention network to generate embedding vectors for nodes and edges: in, Let i be the feature representation of node i. For attention weights, These are trainable parameters. Based on these embedding representations, the engine executes a path reasoning algorithm to explore potential relationship paths between nodes. For newly added knowledge units, the engine verifies their consistency with existing knowledge through semantic similarity calculation and logical rule checks. When conflicts or contradictions are found, an evidence weight evaluation process is initiated, arbitrating based on source credibility and the amount of supporting evidence.
[0045] The continuous evolution of the knowledge graph is achieved through an automated workflow. The system periodically scans predefined knowledge sources, including academic databases, updated clinical guidelines, and new case data from participating institutions. Newly acquired knowledge is added to a temporary knowledge base after entity linking, relation extraction, and confidence assessment. The graph neural network inference engine verifies the logical consistency and sufficiency of evidence for these candidate knowledge sources. Verified knowledge is assigned to the appropriate level based on its type and maturity. Simultaneously, a knowledge decay model is established to downgrade or archive outdated knowledge or knowledge refuted by new evidence.
[0046] The entire architecture uses a monitoring dashboard to display the health status and evolution indicators of the knowledge graph in real time, including knowledge coverage, update frequency, and consistency ratio. Administrators can configure knowledge acquisition sources, set update strategies, and adjust verification thresholds to ensure that the knowledge graph maintains stability while achieving intelligent evolution. Ultimately, this system provides accurate, comprehensive, and continuously updated knowledge support for birth defect prevention and control.
[0047] The multimodal large model analysis module is used to: process preprocessed multimodal data using a multimodal large model and in conjunction with a medical knowledge graph, generate risk level information to characterize the likelihood of birth defects, and obtain prevention and control guidance suggestions based on the risk level information.
[0048] The specific implementation process for generating risk level information to characterize the likelihood of birth defects is as follows: 1) The multimodal large-scale model analysis module first receives structured multimodal data from the preprocessing module, including clinical text, medical images, genomics data, and environmental factor data. This data has been cleaned, normalized, and de-identified to ensure a consistent format and compliance with privacy and security standards. The module employs a Transformer-based multimodal large-scale model, which includes a text encoder, image encoder, sequence encoder, and numerical encoder to process input data from different modalities. For clinical text data, the text encoder uses a pre-trained language model to extract semantic features, converting diagnostic records and symptom descriptions into high-dimensional vector representations. For medical image data, the image encoder uses a convolutional neural network to extract image features, such as fetal structural abnormalities or organ morphology features in ultrasound images. For genomics data, the sequence encoder uses a recurrent neural network or a Transformer variant to process gene sequence information, capturing variant sites and expression patterns. For environmental factor data, the numerical encoder converts numerical values such as pollutant concentrations and exposure history into embedded vectors through fully connected layers.
[0049] 2) The module performs cross-modal feature fusion. The multimodal large model calculates the correlation weights between features from different modalities through a multi-head attention mechanism, achieving feature alignment and interaction. Specifically, the model concatenates the feature vectors of each modality into a joint representation and applies a cross-modal attention layer for information integration. The formula is expressed as: in, Represents the feature vector of clinical text. Represents the feature vector of a medical image. Represents the feature vector of genomics data. Represents the feature vector of environmental factor data. This is the fused multimodal feature matrix. The attention weights are calculated as follows: in, and These represent the query and key matrices, respectively. The feature dimension is used as the feature matrix. The fused feature matrix is further compressed into a unified risk representation vector through a feedforward network.
[0050] 3) The module integrates with a medical knowledge graph for enhanced reasoning. The medical knowledge graph stores etiological, genetic, and diagnostic knowledge related to birth defects in a graph structure, including disease nodes, gene nodes, environmental factor nodes, and their relational edges. The multimodal large model accesses the knowledge graph through a graph attention network, interacting with the unified risk representation vector and entity embeddings within the knowledge graph. Specifically, the model queries the knowledge graph for entities and paths related to the input data, such as matching the association between gene mutations and known birth defects, or the causal chain between environmental factors and disease risk. The knowledge graph embedding vector is updated through a graph neural network and weightedly fused with multimodal features to generate a knowledge-enhanced risk feature vector.
[0051] 4) Based on knowledge-enhanced risk feature vectors, the module performs risk assessment and risk classification. The output layer of the multimodal large model uses a fully connected network and a softmax function to calculate the probability score of the likelihood of birth defects. The probability score ranges from 0 to 1, representing the individual's risk level. Subsequently, the system maps the probability score to discrete risk levels according to preset thresholds. For example, a probability score below 0.3 is classified as low risk, 0.3 to 0.7 as medium risk, and above 0.7 as high risk. Risk level information is output in a structured format, including level labels, confidence scores, and a list of key influencing factors.
[0052] Risk level information, which characterizes the likelihood of birth defects, refers to a graded assessment of the probability of birth defects through quantitative analysis. This information represents the degree of risk in discrete levels, such as low risk, medium risk, or high risk, with each level corresponding to a specific probability range and clinical significance. Risk level information is generated based on multimodal data and medical knowledge and is used to assist in medical decision-making and the formulation of prevention and control strategies.
[0053] The specific process for obtaining prevention and control guidance and recommendations based on risk level information is as follows: 1) The system performs deep matching of risk level information with a medical knowledge graph. For each risk level, the system queries the knowledge graph for corresponding standard prevention and control pathways and intervention measures. For example, low-risk levels correspond to routine monitoring and basic prevention recommendations; medium-risk levels correspond to enhanced screening and early intervention programs; and high-risk levels correspond to specialized diagnosis and comprehensive management strategies. Nodes in the knowledge graph contain specific examination items, treatment plans, follow-up plans, and health guidance content, and edge relationships define the applicability of these measures to specific risk conditions.
[0054] 2) The matching process employs a graph database-based query language to retrieve all effective intervention pathways related to the current risk level. The system also considers the temporal characteristics of the risk, differentiating between prevention and control priorities at different stages: preconception, prenatal, and postpartum. The query results form a basic set of prevention and control measures, including diverse content such as medical examination recommendations, lifestyle adjustments, nutritional supplementation plans, and specialist referral guidelines.
[0055] 3) The system performs personalized screening and prioritization. By analyzing individual characteristics in multimodal data, such as gestational age, age, genetic markers, and past medical history, the system selects the most suitable specific recommendations for the current user from the set of basic prevention and control measures. For example, for users carrying specific gene variants, the system will prioritize recommending relevant genetic counseling and specialized testing; for users at risk of environmental exposure, the system will emphasize avoidance recommendations.
[0056] Priority calculation uses a weighted scoring algorithm: in, This indicates the priority score for the i-th suggestion. This indicates the strength of the association between the recommendation and the risk level. Indicates the level of evidence. Indicates the appropriateness of the timing. These are the weighting coefficients for each dimension. The system sorts the suggestions in descending order of scores, ensuring that the most critical interventions are listed first.
[0057] 4) The system uses natural language generation technology to convert structured prevention and control recommendations into easily understandable textual expressions. The system uses a predefined template library, selecting appropriate expression methods based on the recommendation type while maintaining the accuracy and accessibility of medical terminology. For example, for high-risk users, the system generates detailed guidance including specific examination items, execution times, and precautions; for low-risk users, it generates concise health maintenance reminders.
[0058] 5) The system formats and prepares the generated prevention and control guidance suggestions for output. The suggestions are organized into different chapters according to clinical importance, including urgent action items, short-term plans, and long-term strategies. The system also generates an implementation timeline and precautions to ensure the actionability of the suggestions. All suggestions undergo consistency verification to avoid content conflicts or duplication, and are finally output in the form of a structured document for direct use by doctors and users.
[0059] Optionally, the above technical solution also includes an intelligent interaction and decision support module, which is used to: generate health management suggestions based on risk level information and prevention and control guidance suggestions, and adaptively adjust the expression form of the health management suggestions according to the user role of the preset user.
[0060] Among them, the health management recommendations are specific implementation plans formed based on risk level information and prevention and control guidance. The content includes risk status description, personalized screening plan, treatment plan recommendations, lifestyle guidance, nutrition and exercise recommendations, follow-up schedule, and early warning reminders for abnormal indicators. The system will present this content in an appropriate form according to the user role to ensure a balance between medical professionalism and user understandability.
[0061] The process of generating health management recommendations includes: 1) The system receives risk level information and prevention and control guidance suggestions from the multimodal large model analysis module. Risk level information includes risk level classification, risk probability values, and analysis of major risk factors; prevention and control guidance suggestions contain structured data such as specific medical examination items, treatment plans, and lifestyle intervention measures. The system uses this information as the basic input for generating health management suggestions.
[0062] 2) The system initiates a user role identification process. By analyzing user login information, access permissions, and historical interaction records, the system automatically identifies the current user as either a doctor or a patient. For doctor users, the system extracts detailed information such as their specialty department and professional title; for patient users, the system obtains attributes such as their educational background and health literacy level. This identification process provides a personalized basis for subsequent content generation.
[0063] 3) The system organizes and adjusts the content and depth of health management recommendations based on user roles. The system's built-in health management knowledge base contains templates and rules for various health management scenarios. The system maps prevention and control guidance recommendations to corresponding health management activity templates, generating a complete health management recommendation framework that includes goal setting, action plans, implementation reminders, and effect evaluation. For physician users, the health management recommendation generation process emphasizes professionalism and completeness. The system retains all professional medical terminology, details the medical basis of each recommendation, and provides relevant references and clinical guidelines. The recommendation content is organized according to clinical pathways, including specific implementation standards, expected effects, and risk response plans. The system also generates accompanying patient education materials and communication points to facilitate doctors' communication with patients. For patient users, the health management recommendation generation process emphasizes comprehensibility and operability. The system converts professional medical terminology into easily understandable everyday language, using specific numerical indicators and vivid metaphors to explain health goals. The recommendation content is broken down into simple and clear action steps, each with specific implementation methods and timelines. The system also provides positive incentives and success story sharing to enhance patients' confidence in implementation.
[0064] 4) In the language style adaptation stage, the system uses natural language generation technology to automatically adjust the content expression. The system has multiple built-in language style models, automatically selecting the appropriate vocabulary, sentence structure, and expression method based on the user's role. The doctor's version uses objective and precise academic language, while the patient's version uses friendly and encouraging conversational language. Simultaneously, the system adjusts the level of detail and presentation order of the content according to the user's reading habits.
[0065] 5) In the formatted output phase of health management recommendations, the system organizes the generated content according to a standard template. The recommendation document includes a basic information area, a risk overview area, a management objective area, a specific measures area, and a follow-up feedback area. The system will adjust the level of detail and expression of each part according to different user roles to ensure that the output health management recommendations are both professional and accurate, as well as easy to understand and implement.
[0066] 6) The system pushes the generated health management suggestions to the corresponding user terminals. Doctors receive detailed analysis reports and management plans through a professional medical workstation; patients receive concise and easy-to-understand health guidance and lifestyle reminders through a mobile application. The system also has a feedback collection mechanism to continuously optimize the quality of the generated health management suggestions.
[0067] Let's take a pregnant woman in her second trimester with a moderate risk assessment as an example. Specifically, when the doctor logs into the system to view this case, the system-generated health management recommendations include the following: The pregnant woman is currently 18 weeks pregnant, with a comprehensive risk assessment of moderate. Key risk factors include: age 35 years and abnormal serum screening indicators. It is recommended to immediately schedule amniocentesis for chromosomal karyotype analysis, while simultaneously strengthening fetal echocardiography monitoring. Special attention needs to be paid to screening for Down syndrome and congenital heart disease. Please complete the above examinations before 20 weeks of pregnancy and have a follow-up examination at 24 weeks. Referral to a prenatal diagnostic center for specialist evaluation is recommended. When the pregnant woman herself logs into the mobile application, the system-generated health management recommendations are presented as follows: Dear expectant mother, based on your recent test results, the doctor has developed a special care plan for you. You need to complete a detailed fetal heart examination and non-invasive prenatal testing within the next two weeks. These tests will help understand the baby's health status. Please remember to book an appointment with a specialist at the prenatal diagnostic center and get plenty of rest before the examination. Record fetal movements daily, and contact your doctor immediately if any abnormalities are observed. Your active cooperation is very important for your baby's health.
[0068] The presentation of health management suggestions is adaptively adjusted based on the user's predefined role. The specific implementation process is as follows: 1) In the user role identification phase, the system determines user identity through multi-source information. The system accesses user registration information to obtain basic identity data, including occupation type, professional qualifications, and permission level. For medical professionals, the system further verifies their scope of practice and professional title; for ordinary users, the system assesses their health literacy level and medical knowledge background. Simultaneously, the system analyzes users' historical query records and interaction behaviors to construct a complete user profile. This identification process provides a basis for subsequent content adjustments.
[0069] 2) Entering the content depth adjustment stage, the system adjusts the information density and professional level of health management suggestions based on the identified user role. For doctors, the system retains complete medical terminology and professional technical details, including disease classification codes, drug dosage parameters, and examination indicator thresholds. Suggestions will include detailed explanations of pathophysiological mechanisms and evidence-based medicine support, providing multiple alternatives and comparisons of their advantages and disadvantages. For patients, the system simplifies professional content, converting complex medical terminology into everyday language, highlighting specific operational steps and precautions. The system filters out overly technical theoretical explanations, retaining practical information directly relevant to the user.
[0070] 3) During the language style conversion stage, the system applies natural language generation technology to differentiate expression methods. The system has two built-in language models: a professional language model employs a rigorous academic writing style, using passive voice and standardized terminology; and a colloquial language model adopts a friendly conversational style, using active voice and everyday expressions. The system calls upon the appropriate language model based on the user's role to provide differentiated expressions for the same health management advice. Simultaneously, the system fine-tunes the language difficulty based on the user's age, education level, and other attributes to ensure the effectiveness of information delivery.
[0071] 4) During the presentation adaptation phase, the system optimizes content organization and display methods based on the usage scenarios of different user roles. Suggestions for doctors adopt a standard medical document format, including clear chapter divisions and hierarchical structures, supporting quick browsing and key point extraction. Content is organized according to clinical reasoning logic, prioritizing the presentation of diagnostic evidence and treatment plans. Suggestions for patients adopt a task-oriented format, breaking down complex management plans into specific to-do items, accompanied by intuitive icons and progress prompts. Important information is highlighted with color and repeated emphasis to enhance attention.
[0072] 5) The system continuously optimizes and adjusts its effectiveness through a real-time feedback mechanism. The system records user interaction data on suggested content, including reading time, click hotspots, and subsequent operation completion status. This data is used to evaluate the suitability of the expression format and dynamically adjust the generation strategy for subsequent suggestions.
[0073] Taking the management of gestational diabetes as an example: When the obstetrician logged into the system, the health management advice received was as follows: The patient is currently 28 weeks pregnant. The oral glucose tolerance test results were: fasting blood glucose 5.8 mmol / L, 1-hour postprandial blood glucose 11.2 mmol / L, and 2-hour postprandial blood glucose 9.1 mmol / L, meeting the diagnostic criteria for gestational diabetes. It is recommended to immediately initiate medical nutrition therapy, controlling the daily total calorie intake to 1800-2000 kcal, with carbohydrates accounting for 45%-50%. The blood glucose monitoring plan is seven blood glucose profiles daily (before each of the three meals + 2 hours after each of the three meals + before bedtime), with target blood glucose levels of ≤5.3 mmol / L for fasting and ≤6.7 mmol / L for 2-hour postprandial blood glucose. If blood glucose control is not achieved after one week, insulin therapy should be considered, with an initial dose of 0.7 U / kg / day.
[0074] When a pregnant woman logs into the system, the health management advice she receives is as follows: Based on your recent glucose tolerance test results, you need to start paying attention to your diet and blood sugar management. Please follow the meal plan provided by your nutritionist, measure your blood sugar daily, and record the results. The ideal blood sugar level is: no more than 5.3 before meals and no more than 6.7 two hours after meals. If your blood sugar is still high after a week, your doctor may recommend using insulin to help control your blood sugar. Please maintain a relaxed attitude; most pregnant women can successfully navigate their pregnancy with proper management.
[0075] Optionally, the above technical solution also includes a closed-loop management and continuous evolution module. This module is used to: analyze health management recommendations and subsequently collected user multimodal data, and feed the analysis results back to the multimodal large-scale model analysis module and the medical knowledge graph construction module to update the multimodal large-scale model and the medical knowledge graph. The specific implementation method is as follows: 1) The system continuously collects multimodal data generated by users during the implementation of health management recommendations. This subsequent data enters the system through the multimodal data acquisition and preprocessing module, including updated clinical texts (such as new symptom descriptions and follow-up medical records), medical images (such as ultrasound images from follow-up examinations), genomic data (such as subsequent genetic testing reports), and environmental factor data (such as information on changed living environments). All data undergoes the same preprocessing process as the initial stage, including standardization, de-identification, and feature extraction, forming a structured dataset that can be used for analysis.
[0076] 2) System-initiated health management recommendation effectiveness analysis. This analysis is performed by comparing changes in user status before and after the implementation of health management recommendations. The system longitudinally compares subsequently collected multimodal data with historical baseline data to calculate the degree of improvement in key health indicators. For example, for a blood glucose control recommendation, the system analyzes the trend of subsequent blood glucose test values relative to the baseline; for a nutritional intervention recommendation, the system tracks changes in relevant biochemical indicators and body composition data. The analysis process employs statistical models and causal inference methods to quantify the actual effect of each health management recommendation and identify user characteristics and environmental factors affecting the effect.
[0077] 3) The system encapsulates the analysis results into structured feedback data and sends them to the multimodal large-scale model analysis module and the medical knowledge graph construction module, respectively. Specifically, For updates to the multimodal large-scale model analysis module, the system employs online learning technology to incorporate feedback data into the model training process. Specifically, the system combines initial user data, health management suggestions, and post-implementation effect data into new training samples, and updates the parameters of the multimodal large-scale model through an incremental learning algorithm. The loss function considers both prediction accuracy and suggestion effectiveness. in, It is a loss item in the risk assessment. This is the loss item in terms of the suggested effect. and It is a hyperparameter that balances the two weights. This update enables the multimodal large model to learn from the effects of real-world applications, continuously improving the accuracy of risk assessment and prevention guidance.
[0078] For updates to the medical knowledge graph construction module, the system integrates validated and effective interventions and newly discovered relationships into the knowledge graph. The system uses entity linking technology to match interventions in health management recommendations with corresponding nodes in the knowledge graph, and then updates node attributes and edge relationships based on the results of effect analysis. For example, when an intervention consistently shows good results in a specific user group, the system strengthens the association between the intervention and the corresponding conditions; when new risk factors are discovered to be associated with birth defects, the system creates new entities and relationships in the knowledge graph. The knowledge graph updates employ a version control mechanism to ensure the traceability of knowledge evolution.
[0079] 4) The system utilizes a federated learning framework to enable cross-institutional data collaboration and model sharing. Each participating institution updates its model and enhances its knowledge locally, then uploads only the incremental updates to the model parameters and knowledge graph to the central server for aggregation. This approach protects user data privacy while achieving multi-center experience sharing and overall system performance improvement.
[0080] Take, for example, a pregnant woman who was assessed as being at high risk for birth defects or neural tube defects and received advice on folic acid supplementation for health management: After the system recommended the supplementation, it continuously collected the pregnant woman's subsequent prenatal data, including serum folic acid concentration measurements, new ultrasound images, and daily nutritional records. Effect analysis showed that four weeks after the folic acid supplementation recommendation, the pregnant woman's serum folic acid concentration increased from a baseline of 8 nmol / L to 28 nmol / L, and subsequent ultrasound examinations revealed no neural tube abnormalities. This analysis result was fed back into the system update process. The multimodal large model analysis module added this case as a positive sample to the training set, reinforcing the association between folic acid supplementation and reduced neural tube defect risk. The medical knowledge graph construction module updated the attributes of the folic acid supplementation node, adding specific information on the effective dosage and intervention duration for the pregnant woman's population (e.g., specific age range, genetic background). Simultaneously, the system discovered that the pregnant woman's vitamin B12 levels also increased, and further analysis revealed that she had also adjusted her dietary structure. This new finding was added to the knowledge graph, establishing a complementary relationship between a balanced diet and neural tube defect prevention, enriching the existing prevention and control knowledge system.
[0081] The present invention will be further described in detail through the following embodiments: This invention's system integrates clinical text, medical imaging, genomics data, and environmental data to construct a sustainably evolving intelligent decision-making platform. This system achieves comprehensive prevention and control coverage across the entire chain—preconception, prenatal, and postnatal—improving early screening accuracy, assisting in diagnosis and treatment decisions, and providing a scientific basis for public health policy formulation. The technical solution comprises two parts: the overall system architecture and the methodological steps, ensuring a complete process from data collection to decision optimization. The overall system architecture consists of multiple core modules that work collaboratively to achieve data integration, knowledge construction, intelligent analysis, and decision support. Specifically: The multimodal data acquisition and preprocessing module is responsible for collecting multimodal data related to birth defects from multiple data sources. Data sources include electronic medical record systems, medical imaging equipment such as ultrasound and MRI, databases of gene testing institutions, records of pregnant women's environmental exposures, and follow-up records. This module uses standardized interfaces such as HL7 and FHIR for data access, ensuring cross-platform compatibility. Privacy-preserving computation techniques such as differential privacy and federated learning are used to achieve data anonymization, encryption, and secure sharing, protecting patient privacy. Preprocessing includes data cleaning, de-identification, format standardization, and feature extraction, providing high-quality input for subsequent analysis. The medical knowledge graph construction module constructs a multi-level medical knowledge graph based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects. This module uses graph neural networks and natural language processing (NLP) techniques to achieve semantic association and dynamic knowledge updates. Graph neural networks learn complex relationships between entities, such as the association between gene mutations and diseases, while NLP extracts entities and relationships from text data. The knowledge graph provides interpretable knowledge support for large models and achieves automatic evolution by continuously monitoring new research findings. The multimodal large-scale model analysis module adopts a Transformer architecture that integrates text, images, and gene sequences to establish a unified representation space, enabling cross-modal feature extraction and joint inference. This module calculates the association weights of different modalities through a multi-head attention mechanism, generating a unified risk prediction vector. The large model supports etiology prediction, risk assessment, treatment plan recommendations, and individualized intervention strategy output, providing accurate decision-making basis based on fused features. The intelligent interaction and decision support module provides customized services for different user roles. For physicians, it offers evidence-based decision-making for difficult cases, screening strategy suggestions, and multidisciplinary collaborative management solutions to assist clinical workflows. For patients, it provides health education, screening guidance, and home care advice, with language style adaptively adjusted according to user role to ensure comprehensibility and usability. This module dynamically generates content through natural language generation technology to enhance user experience. The closed-loop management and continuous evolution module feeds screening, diagnosis, treatment, and follow-up data back into the model in real time, utilizing online learning and reinforcement learning to achieve model self-iteration. This module supports cross-regional and cross-institutional multi-center collaboration, sharing model updates through a federated learning framework without data leaving the domain, enabling regional public health trend analysis and intelligent policy recommendations. The continuous evolution mechanism ensures that the system's performance is constantly optimized as new data accumulates. The specific workflow is as follows: The data collection and cleaning process gathers multimodal data from pregnant women and newborns through interfaces from multiple sources, including hospital HIS / EMR systems, genetic testing institutions, and imaging centers. The collected data includes clinical texts, medical images, genomic data, and environmental factor data. The cleaning process involves de-identification, removing personal identifiers, and standardizing the data format to ensure consistency and computability. Privacy protection technologies such as encryption and anonymization are applied during data transmission and storage.
[0082] The knowledge graph construction process involves entity recognition, relationship extraction, and graph modeling of medical information such as etiological knowledge, treatment guidelines, and case literature. Natural language processing techniques are used to identify entities such as diseases, genes, and environmental factors, while graph neural networks model the relationships between entities, forming a structured knowledge base. The knowledge graph is dynamically updated, integrating the latest research findings to provide real-time knowledge support for large-scale models.
[0083] The multimodal feature fusion step extracts features from different modalities of data, including text, images, and genes, and achieves cross-modal fusion through a multi-head attention mechanism. Text features are extracted using a pre-trained language model, image features are obtained through a convolutional neural network, and gene features are processed using a sequence model. The attention mechanism calculates cross-modal associations and generates a unified risk prediction vector, expressed by the following formula: in, Represents the feature vector of clinical text. Represents the feature vector of a medical image. Represents the feature vector of genomics data. Represents the feature vector of environmental factor data. This is the fused multimodal feature matrix.
[0084] The risk assessment and treatment plan recommendation process outputs individualized risk assessment results based on fusion features, and matches them with a knowledge graph to generate intervention suggestions and treatment pathways for preconception, prenatal, and postnatal stages. A large-scale model calculates the risk probability and classifies risk levels (e.g., low, moderate, or high) based on thresholds. Treatment plan recommendations include specific examinations, treatments, and lifestyle adjustments, ensuring personalized and actionable advice.
[0085] The closed-loop management and feedback iteration process transmits treatment results and follow-up data back to the system, updating model parameters and the knowledge graph in real time. Online learning technology adjusts the weights of the large model, reinforces learning to optimize decision-making strategies, and achieves adaptive evolution and decision optimization. The feedback loop ensures that the system learns from practical applications, continuously improving accuracy and reliability.
[0086] This invention demonstrates the actual workflow of a multimodal large-scale model-based intelligent prevention and control system for birth defects throughout the entire lifecycle through specific application scenarios. These embodiments detail how the system integrates multimodal data, utilizes medical knowledge graphs, performs intelligent analysis, and provides decision support, showcasing its full-cycle prevention and control capabilities and continuous evolution.
[0087] In the scenario of rare disease diagnosis by primary care physicians, these physicians upload multimodal data, including genetic testing reports and symptom descriptions of affected children, through the system's intelligent interaction and decision support module. The multimodal data acquisition and preprocessing module receives this data, including symptom descriptions in clinical text form and genetic testing reports in genomics data form, and performs standardized cleaning and de-identification processing. The medical knowledge graph construction module provides a dynamically updated rare disease knowledge base, and the system automatically matches disease entries in the knowledge base to identify potential rare disease types. The multimodal large-scale model analysis module performs cross-modal feature fusion on genetic data and symptom text, calculates association weights through a multi-head attention mechanism, and generates a unified risk prediction vector. Based on the analysis results, the system generates a multidisciplinary collaborative treatment plan including neurology and genetics, detailing examination items, treatment recommendations, and follow-up plans. The intelligent interaction and decision support module simultaneously pushes the treatment plan to the expert portal of a higher-level hospital, where experts remotely review and provide feedback. The closed-loop management and continuous evolution module collects expert feedback and subsequent treatment results, updates the multimodal large-scale model parameters and knowledge graph content, and optimizes future diagnostic accuracy.
[0088] In the context of pregnancy-related patient health management, pregnant women input gestational age data and prenatal examination indicators through a mobile intelligent interaction and decision support module. The multimodal data acquisition and preprocessing module collects this clinical text and environmental factor data, performing format standardization and privacy protection. The medical knowledge graph construction module provides prenatal and postnatal care knowledge and screening guidelines, automatically matching the screening items corresponding to the current gestational week. The multimodal large-scale model analysis module performs risk assessment on the input data, calculates the probability of birth defects, and generates personalized prevention and control suggestions. The intelligent interaction and decision support module answers specific questions raised by users, such as the difference between Down syndrome screening and non-invasive prenatal testing (NIPT), providing easy-to-understand comparisons using natural language processing technology. Based on the risk assessment results and knowledge graph recommendations, the system generates personalized nutrition suggestions and exercise plans, considering the individual differences and preferences of pregnant women. The closed-loop management and continuous evolution module sets up follow-up reminders for abnormal indicators, automatically pushing notifications when deviations occur in prenatal examination data and collecting user feedback for model iteration. All interactive data is securely shared through federated learning technology, participating in multi-center model optimization to improve overall system performance.
[0089] These embodiments fully demonstrate the entire process of the system, from data input, knowledge retrieval, intelligent analysis to decision output and feedback optimization, showcasing the technical characteristics of multimodal fusion, knowledge-driven approach, and adaptive evolution. The system verifies the feasibility and effectiveness of the technical solution through practical applications, providing a practical tool for birth defect prevention and control.
[0090] like Figure 2 As shown in the figure, an intelligent management method for the entire life cycle of birth defects based on a multimodal large model according to an embodiment of the present invention includes the following steps: S1. Collect multimodal data on birth defects associated with preset users from multiple data sources, and preprocess the collected multimodal data; S2. Construct a medical knowledge graph based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects; S3. The preprocessed multimodal data is processed using a multimodal large model and combined with a medical knowledge graph to generate risk level information that characterizes the likelihood of birth defects, and prevention and control guidance suggestions are obtained based on the risk level information.
[0091] Optionally, the above technical solution also includes: generating health management suggestions based on risk level information and prevention and control guidance, and adaptively adjusting the expression of the health management suggestions according to the user role of the preset user.
[0092] Optionally, the above technical solution also includes: analyzing health management recommendations with subsequently collected user multimodal data, and updating the multimodal big model and medical knowledge graph based on the analysis results.
[0093] Optionally, in the above technical solutions, the multimodal data associated with birth defects include clinical texts, medical images, genomic data, and environmental factor data.
[0094] It should be noted that the beneficial effects of the intelligent management method for the entire life cycle of birth defects based on a multimodal large model provided in the above embodiments are the same as the beneficial effects of the intelligent management system for the entire life cycle of birth defects based on a multimodal large model. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0095] The intelligent management system for the entire life cycle of birth defects based on a multimodal large model of the present invention can be a computer program (including program code) running on a computer device. For example, the intelligent management system for the entire life cycle of birth defects based on a multimodal large model of the present invention is an application software that can be used to execute the corresponding steps in the intelligent management method for the entire life cycle of birth defects based on a multimodal large model of the present invention.
[0096] In some embodiments, the intelligent management system for the entire life cycle of birth defects based on a multimodal large model of the present invention can be implemented in a combination of hardware and software. As an example, the intelligent management system for the entire life cycle of birth defects based on a multimodal large model of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the intelligent management method for the entire life cycle of birth defects based on a multimodal large model of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0097] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.
[0098] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned intelligent management methods for the entire life cycle of birth defects based on a multimodal large model. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the intelligent management method for the entire life cycle of birth defects based on a multimodal large model shown in any embodiment of the present invention by calling the computer program.
[0099] In one alternative embodiment, an electronic device is provided, such as Figure 3 As shown, Figure 3 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.
[0100] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0101] Bus 4002 may include a path for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus 4002 is represented by only one thick line, but this does not mean that there is only one bus or one type of bus.
[0102] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0103] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0104] Among them, electronic devices can also be terminal devices, which can be any device that can install applications, including at least one of smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, smart TVs, and smart in-vehicle devices.
[0105] It should be noted that, Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0106] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned intelligent management methods for the entire life cycle of birth defects based on a multimodal large model.
[0107] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0108] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the above-described intelligent management methods for the entire lifecycle of birth defects based on a multimodal large model.
[0109] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0111] The computer-readable storage medium provided in this invention can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EEPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0112] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0113] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
[0114] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.
[0115] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a circuit, module, or system. Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0116] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A birth defect full-cycle intelligent management system based on a multimodal large model, characterized in that, It includes a multimodal data acquisition and preprocessing module, a medical knowledge graph construction module, and a multimodal large model analysis module; The multimodal data acquisition and preprocessing module is used to: acquire multimodal data related to birth defects of a preset user from multiple data sources, and preprocess the acquired multimodal data; The medical knowledge graph construction module is used to: construct a medical knowledge graph based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects; The multimodal large model analysis module is used to: process the preprocessed multimodal data using a multimodal large model and in conjunction with the medical knowledge graph, generate risk level information to characterize the likelihood of birth defects, and obtain prevention and control guidance suggestions based on the risk level information.
2. The intelligent management system for the entire life cycle of birth defects based on a multimodal large model according to claim 1, characterized in that, It also includes an intelligent interaction and decision support module, which is used to: generate health management suggestions based on the risk level information and the prevention and control guidance suggestions, and adaptively adjust the expression of the health management suggestions according to the user role of the preset user.
3. The intelligent management system for the entire life cycle of birth defects based on a multimodal large model according to claim 2, characterized in that, It also includes a closed-loop management and continuous evolution module, which is used to: analyze the health management recommendations and the subsequently collected user multimodal data, and feed the analysis results back to the multimodal large model analysis module and the medical knowledge graph construction module to update the multimodal large model and the medical knowledge graph.
4. A birth defect full-cycle intelligent management system based on a multimodal large model according to any one of claims 1 to 3, characterized in that, The multimodal data associated with birth defects include clinical texts, medical images, genomic data, and environmental factor data.
5. A method for intelligent management of birth defects throughout the entire lifecycle based on a multimodal large model, characterized in that, include: Multimodal data related to birth defects of a preset user are collected from multiple data sources, and the collected multimodal data is preprocessed. A medical knowledge graph was constructed based on etiology, genetics, treatment guidelines, and research findings in the field of birth defects. The preprocessed multimodal data is processed using a multimodal large model and the medical knowledge graph to generate risk level information that characterizes the likelihood of birth defects, and prevention and control guidance suggestions are obtained based on the risk level information.
6. The intelligent management method for the entire life cycle of birth defects based on a multimodal large model according to claim 5, characterized in that, Also includes: Based on the risk level information and the prevention and control guidance, health management suggestions are generated, and the expression of the health management suggestions is adaptively adjusted according to the user's preset user role.
7. The intelligent management method for the entire life cycle of birth defects based on a multimodal large model according to claim 6, characterized in that, Also includes: The health management recommendations are analyzed in conjunction with subsequently collected user multimodal data, and the multimodal big data model and the medical knowledge graph are updated based on the analysis results.
8. A method for intelligent management of birth defects throughout the entire life cycle based on a multimodal large model according to any one of claims 5 to 7, characterized in that, The multimodal data associated with birth defects include clinical texts, medical images, genomic data, and environmental factor data.
9. An electronic device, characterized in that, The invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent management method for the entire life cycle of birth defects based on a multimodal large model as described in any one of claims 5 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the intelligent management method for the entire life cycle of birth defects based on a multimodal large model as described in any one of claims 5 to 8.