Artificial intelligence corpus construction method, device and equipment based on multi-modal data

By constructing an artificial intelligence corpus based on multimodal data, the difficulties in data integration and utilization in the diabetes corpus were solved, efficient personalized treatment support and dynamic updates were achieved, and data utilization efficiency was improved.

CN120748762APending Publication Date: 2025-10-03CHINA JAPAN FRIENDSHIP HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510887835.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The lack of unified standards and structured processing methods in existing technologies leads to difficulties in data integration and multimodal data utilization of the diabetes artificial intelligence corpus, making it impossible to efficiently meet personalized needs and dynamic updates.

Method used

By acquiring multimodal health data corresponding to disease types, generating multimodal features, extracting medical entities and associations, constructing knowledge entries, building a knowledge graph and generating medical text, integrating information from different data modalities, and performing corpus construction and optimization.

Benefits of technology

It achieves efficient integration and utilization of multimodal data, improves the storage and retrieval efficiency of medical knowledge, adapts to individual differences, and supports personalized treatment and dynamic updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748762A_ABST
    Figure CN120748762A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence corpus construction method, device and equipment based on multi-modal data, and relates to the technical field of data processing. The specific implementation scheme is as follows: acquiring health data corresponding to a disease type; according to the health data, multi-modal features corresponding to the patient codes are generated; extracting medical entities and association relationships according to the multi-modal features, and generating knowledge entries; and constructing an artificial intelligence corpus according to the knowledge entries. According to the scheme, information of different data modes can be integrated, and limitation of single-mode data is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, device and equipment for constructing an artificial intelligence corpus based on multimodal data. Background Art

[0002] In recent years, the rapid development of artificial intelligence technology has provided new solutions for the intelligent management and personalized treatment of various diseases. Especially in the field of medical data analysis, artificial intelligence can process complex medical data and explore potential health patterns, thereby improving the accuracy of diabetes diagnosis, personalizing treatment plans, and predicting and preventing complications. To achieve this goal, building a comprehensive and accurate diabetes artificial intelligence corpus has become the basis for the application of artificial intelligence in diabetes management. This corpus needs to include not only traditional clinical data, but also integrate multimodal data such as patients' living habits, exercise status, and wearable device monitoring indicators to provide sufficient data support for model training. The lack of unified standards and structured processing methods in existing technologies has led to difficulties in data integration during the corpus construction process and the inability to efficiently utilize multiple data sources. Summary of the Invention

[0003] The present invention provides a method for constructing an artificial intelligence corpus based on multimodal data.

[0004] According to a first aspect of the present invention, a method for constructing an artificial intelligence corpus based on multimodal data is provided, comprising: acquiring health data corresponding to disease types; generating multimodal features corresponding to patient codes based on the health data; extracting medical entities and association relationships based on the multimodal features to generate knowledge entries; and constructing an artificial intelligence corpus based on the knowledge entries.

[0005] Optionally, obtaining health data corresponding to the disease type includes: screening and obtaining available data sources according to the disease type; and obtaining health data according to the available data sources.

[0006] Optionally, based on the health data, multimodal features corresponding to the patient code are generated, including: determining the structured data, unstructured data and image data corresponding to the patient code from the health data; determining the structured numerical value based on the structured data; determining the unstructured feature vector based on the unstructured data; determining the image feature vector based on the image data using a pre-trained model; and performing feature fusion on the structured numerical value, unstructured feature vector and image feature vector to obtain multimodal features.

[0007] Optionally, medical entities and association relationships are extracted based on multimodal features to generate knowledge entries, including: identifying medical entities and association relationships corresponding to medical entities based on multimodal features to generate initial knowledge entries; and performing knowledge fusion on medical entities and association relationships based on the initial knowledge entries to obtain knowledge entries.

[0008] Optionally, an artificial intelligence corpus is constructed based on the knowledge items, including: establishing a knowledge graph based on the knowledge items, using medical entities as nodes and association relationships as relationships; generating medical text based on the knowledge graph using a pre-trained large language model; and performing corpus enhancement on the medical text to obtain an artificial intelligence corpus.

[0009] Optionally, the method for constructing an artificial intelligence corpus based on multimodal data also includes: obtaining incremental health data corresponding to the patient code; generating incremental multimodal features corresponding to the patient code based on the incremental health data; generating incremental knowledge entries based on the incremental multimodal features; and updating the artificial intelligence corpus based on the incremental knowledge entries.

[0010] Optionally, the method for constructing an artificial intelligence corpus based on multimodal data further includes: performing quality assessment on the artificial intelligence corpus in a preset manner to generate a corpus quality assessment result; identifying high-uncertainty data in the artificial intelligence corpus based on the corpus quality assessment result, and optimizing the corpus using an active learning method to obtain an optimized artificial intelligence corpus; performing expert annotation correction on the optimized artificial intelligence corpus, and updating the artificial intelligence corpus based on the correction results.

[0011] According to a second aspect of the present invention, there is provided an artificial intelligence corpus construction device based on multimodal data, comprising: a data acquisition module for acquiring health data corresponding to disease types; a feature extraction module for generating multimodal features corresponding to patient codes based on health data; an entry generation module for extracting medical entities and association relationships based on multimodal features to generate knowledge entries; and a corpus construction module for constructing an artificial intelligence corpus based on the knowledge entries.

[0012] According to a third aspect of the present invention, there is provided an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention and any optional embodiment of the first aspect.

[0013] According to a fourth aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method in the first aspect of the present invention and any optional embodiment of the first aspect.

[0014] The technical solution of the present invention can integrate information of different data modalities, thus avoiding the limitations of single-modal data.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. Among them: Figure 1 1 is a flow chart of a method for constructing an artificial intelligence corpus based on multimodal data according to an embodiment of the present invention; Figure 2 2 is a schematic structural diagram of an apparatus for constructing an artificial intelligence corpus based on multimodal data according to an embodiment of the present invention; Figure 3 1 is a schematic diagram of a scenario of a method for constructing an artificial intelligence corpus based on multimodal data according to an embodiment of the present invention; Figure 4 3 is a structural diagram of an electronic device used to implement the method for constructing an artificial intelligence corpus based on multimodal data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, and various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0019] In addition, numerous specific details are provided in the following detailed description to better illustrate the present invention. Those skilled in the art will appreciate that the present invention can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of the present invention.

[0020] In the related art, there have been some technical solutions that attempt to study the construction of artificial intelligence corpora for diabetes, but most of the solutions still have problems such as a single data source and a lack of customized data sets for different application scenarios. The following are several existing technical solutions that are most similar to the present invention: (1) Construction of artificial intelligence corpora for single modal data: In many existing diabetes-related studies, artificial intelligence models usually rely on a single source of clinical data, such as blood glucose test values, blood lipid levels, etc. These methods usually use traditional machine learning models (such as decision trees, support vector machines, etc.) and use a single biochemical indicator as input to perform diabetes risk prediction, treatment response analysis, etc. A typical solution is to use medical record data in an electronic medical record system to train a diabetes prediction model. Although this method can provide certain support for clinical diagnosis, due to the lack of comprehensive and diverse data sources, its prediction accuracy and personalization are low, and it is difficult to adapt to individual differences in patients. Application exploration of multimodal data: In recent years, with the advancement of artificial intelligence technology, some studies have begun to explore the use of multimodal data to construct diabetes artificial intelligence corpora. For example, some studies have attempted to combine basic information, clinical laboratory data, and imaging data of patients to predict diabetic complications. However, existing multimodal data fusion methods still have some problems, such as the difficulty of standardizing data integration and the heterogeneity caused by diverse data sources, which makes it difficult to effectively mine the correlations between data. In addition, some methods ignore how to build targeted corpora based on different clinical application scenarios during data processing, resulting in datasets that are too general and cannot meet the needs of practical applications. Application of dynamic update and real-time data synchronization technology: In artificial intelligence management systems for diabetes, it is crucial to update patient data in a timely manner to reflect changes in the condition. Some existing systems have attempted to achieve real-time monitoring and updating of patient health data based on health monitoring devices. However, the application of these systems is often limited to the recording of static data and does not fully realize the intelligent analysis and updating of data. For example, some systems based on smart devices can only track blood sugar levels or other physiological parameters in real time and cannot realize dynamic updating of real-time data and intelligent decision support within the platform.

[0021] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the present invention proposes a method for constructing an artificial intelligence corpus based on multimodal data, which can integrate information from different data modalities and avoid the limitations of single-modal data.

[0022] The present invention provides a method for constructing an artificial intelligence corpus based on multimodal data. Figure 1It is a flow chart of a method for constructing an artificial intelligence corpus based on multimodal data, and the method for constructing an artificial intelligence corpus based on multimodal data can be applied to an artificial intelligence corpus construction device based on multimodal data. The artificial intelligence corpus construction device based on multimodal data is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to corpus construction terminals, and corpus construction terminals can be mobile phones, tablet computers, smart wearable devices, etc. In some possible implementations, the method for constructing an artificial intelligence corpus based on multimodal data can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1 As shown, the method for constructing an artificial intelligence corpus based on multimodal data includes: S101. Obtain health data corresponding to the disease type; S102. Generate multimodal features corresponding to the patient code based on the health data; S103. Extract medical entities and association relationships based on multimodal features to generate knowledge items; S104. Construct an artificial intelligence corpus based on knowledge items.

[0023] The disease type refers to a systematic classification of diseases based on medical standards. In the embodiment of the present invention, the disease type can be pre-set according to actual conditions or obtained according to the classification standards of the International Classification of Diseases.

[0024] Health data refers to various physiological and pathological information related to disease types. In embodiments of the present invention, health data may include electronic health record (EHR) data, medical imaging data, genomic data, laboratory test data, patient self-reported data, wearable device monitoring data, medical literature and guidelines, etc.

[0025] In an embodiment of the present invention, the health data to be obtained can be first determined based on the type of disease. Specifically, data related to the cause of the disease and / or the predisposing factors can be determined as the health data to be obtained based on the cause of the disease and / or the predisposing factors. For example, for influenza, the cause of the disease may be the invasion of the influenza virus into the human body. In this case, EHR data, medical imaging data, laboratory test data, patient self-reported data, wearable device monitoring data, medical literature and guidelines, etc. can be selected as the health data to be obtained; for diabetes, the cause may include genetic factors and environmental factors. In this case, EHR data, medical imaging data, genomic data, laboratory test data, patient self-reported data, wearable device monitoring data, medical literature and guidelines, etc. can be selected as the health data to be obtained. Subsequently, health data can be obtained from the source corresponding to the health data to be obtained. In particular, data from different sources can be integrated to form a complete health data set. The above is only an exemplary description and is not intended to limit all possible situations for obtaining health data. It is just not exhaustive here.

[0026] A patient code is a standardized code used in the medical field to uniquely identify an individual patient. In embodiments of the present invention, a patient code accurately links all health information about a single patient in medical data management, electronic medical record systems, and scientific research databases, preventing data confusion or duplication. Specifically, a patient code can be a string of numbers, letters, or symbols.

[0027] Multimodal features refer to feature representations generated by integrating different modalities of a patient's health data. In embodiments of the present invention, multimodal features can be used to accurately describe a patient's health status. Specifically, multimodal features can include multiple dimensions, with each dimension corresponding to a data feature.

[0028] In an embodiment of the present invention, corresponding features can be first extracted from health data. Subsequently, a multimodal feature fusion model can be established to convert data features of different modalities into a unified representation. Exemplarily, text and image data can be represented as feature vectors using embedding technology or deep learning models. Then, the extracted features can be grouped by patient based on the patient code to generate multimodal features corresponding to each patient. Finally, the generated multimodal features can be stored in a database. The above is only an exemplary explanation and is not intended to limit all possible situations for generating multimodal features, but it is not exhaustive here.

[0029] A medical entity is a concept or entity object with independent medical semantics in the medical field. In the embodiments of the present invention, a medical entity is a concept or object with independent medical semantics extracted from multimodal features, and may include disease names, symptoms, examination items, drug terms, etc.

[0030] The association relationship refers to the semantic connection between medical entities based on medical logic. In the embodiment of the present invention, the association relationship can be used to describe the interaction, attributes or clinical association between medical entities.

[0031] Wherein, knowledge items refer to structured knowledge records generated by extracting medical entities and their associations. In the embodiment of the present invention, knowledge items can be used to store and express medical knowledge.

[0032] In an embodiment of the present invention, medical entities can be first extracted from multimodal features. For example, for text features, natural language processing technology can be used to extract medical entities; for image features, image processing technology can be used to extract lesion features. In particular, a knowledge base can also be used to perform entity standardization mapping. Subsequently, association relationships can be extracted from multimodal features. For example, a relationship extraction model can be used to identify association relationships between medical entities. Finally, structured knowledge items can be generated based on the extracted entities and relationships. The above is only an exemplary explanation and is not intended to limit all possible situations for generating knowledge items, but it is not exhaustive here.

[0033] The artificial intelligence corpus is a knowledge collection built based on knowledge items, with semantic association and structured storage. In the embodiment of the present invention, the artificial intelligence corpus is a computable, searchable, and scalable knowledge system built using knowledge items.

[0034] In an embodiment of the present invention, knowledge items can first be stored according to corpus standards. In particular, an indexing mechanism can be established. Furthermore, duplicate removal can be performed on the knowledge items to ensure the quality of the generated corpus. The above is merely an example and does not limit all possible scenarios for constructing an AI corpus. This is simply not an exhaustive list.

[0035] The technical solution of the embodiment of the present invention can effectively screen out relevant health data by acquiring health data corresponding to disease types and utilizing the correspondence between data and disease types. By generating multimodal features corresponding to patient codes, information from different data modalities can be integrated, avoiding the limitations of single-modal data. By generating knowledge items, complex multimodal data can be converted into structured knowledge. By constructing an artificial intelligence corpus, the storage and retrieval of knowledge items can be facilitated, improving the efficiency of medical knowledge utilization. At the same time, it can serve as a basic data source for the development of applications such as question-answering systems and recommendation systems.

[0036] In some embodiments, obtaining health data corresponding to the disease type includes: screening and obtaining available data sources according to the disease type; and obtaining health data according to the available data sources.

[0037] The available data sources refer to specific data storage or collection systems that can provide health data related to the target disease type. In embodiments of the present invention, available data sources may include electronic medical record systems, laboratory testing systems, medical imaging databases, wearable device data platforms, medical databases, and patient self-reports.

[0038] In an embodiment of the present invention, the type of disease to be treated can be first determined based on the research or application requirements, and then the health data related to the target disease can be determined based on the target disease type, and then the available data sources can be filtered through the required data. For example, for diabetes, EHR data, medical imaging data, genomic data, laboratory test data, patient self-reported data, wearable device monitoring data, medical literature and guidelines, etc. are selected as the health data to be obtained. At this time, the source selection can be carried out in sequence according to the source of the health data, for example, the electronic medical record system database can be used as the available data source for EHR data, the imaging department database can be used as the source of medical imaging data, the gene database can be used as the source of genomic data, the laboratory department database can be used as the source of laboratory test data, and so on. The above is only an illustrative explanation and is not intended to limit all possible situations for obtaining the source of available data, but it is not exhaustive here.

[0039] In embodiments of the present invention, an interface or direct data extraction tool can be used to access the filtered data sources. Subsequently, health data related to the target disease can be extracted from the data sources based on the disease type and the screening results. In particular, health data from different sources can be converted into a unified data format. The above is merely an example and does not limit all possible scenarios for obtaining health data. This is not intended to be exhaustive.

[0040] By filtering data sources by disease type, we can avoid processing irrelevant data, improve data processing efficiency, and ensure that the data sources are highly relevant to the target disease, thereby obtaining more valuable health data. By obtaining data from the selected available data sources, we can quickly locate and extract the target health data while avoiding data errors or erroneous sources.

[0041] In some embodiments, multimodal features corresponding to patient codes are generated based on health data, including: determining structured data, unstructured data, and image data corresponding to patient codes from health data; determining structured numerical values ​​based on structured data; determining unstructured features based on unstructured data; determining image features based on image data using a pre-trained model; and fusing structured numerical values, unstructured features, and image features to obtain multimodal features.

[0042] Structured data refers to health data stored in a fixed format. In embodiments of the present invention, structured data may be in the form of a table or database, and may include numerical data such as blood sugar and blood pressure values, as well as categorical data such as gender and age.

[0043] Unstructured data refers to free text or semantic data without a fixed format. In the embodiment of the present invention, unstructured data may include diagnostic records, text descriptions in medical records, and patient self-descriptions.

[0044] The image data refers to medical imaging information. In the embodiment of the present invention, the image data may include computed tomography (CT) image data, magnetic resonance imaging (MRI) image data, X-rays, etc.

[0045] In an embodiment of the present invention, the patient's health data can first be located based on the unique patient code. Subsequently, structured data can be filtered from the patient's health data. For example, numerical data and categorical data can be extracted directly from a spreadsheet or database. Next, unstructured data can be filtered from the health data. For example, medical records and doctor's diagnosis records can be extracted. Finally, image data can be filtered from the health data. For example, medical imaging files can be extracted. The above is only an exemplary description and is not intended to limit all possible situations for determining structured data, unstructured data, and image data. It is just not exhaustive here.

[0046] In embodiments of the present invention, the structured data can first be cleaned and preprocessed. For example, outliers can be removed and the data normalized. Subsequently, important numerical features related to the target disease can be selected from the structured data. Finally, the index values ​​corresponding to the selected important numerical features can be used as structured values. The above is merely illustrative and does not limit all possible scenarios for determining structured values; however, this is not intended to be exhaustive.

[0047] In an embodiment of the present invention, the unstructured data can be preprocessed first. For example, the text data can be cleaned and word segmentation can be performed. In particular, speech recognition can also be performed on the audio data to convert the speech input into text data. Subsequently, natural language processing tools can be used to extract text features. For example, high-frequency keywords can be extracted or semantic embeddings can be generated. In particular, the symptoms and disease names in the text can also be mapped to standardized medical terms. Finally, a deep learning model can be used to generate semantic features, and the obtained semantic features can be used as unstructured features. The above is only an exemplary description and is not intended to limit all possible situations for determining unstructured features, but it is not exhaustive here.

[0048] In an embodiment of the present invention, the image data may first be preprocessed. For example, the image format may be converted and image enhancement may be performed to improve feature detection. Subsequently, image features may be extracted using a pre-trained model. For example, an edge detection algorithm or a computer vision model may be used to extract lesion boundary information. Finally, the extracted image features may be converted into a fixed-length feature vector. The above is merely an illustrative description and does not constitute a complete list of all possible scenarios for determining image features. This is simply not an exhaustive list.

[0049] In an embodiment of the present invention, the structured numerical values, unstructured features, and image features can first be normalized or standardized so that they have a uniform scale. Subsequently, feature fusion technology can be used to combine features of different modalities to form a unified feature representation, which is used as a multimodal feature. Exemplarily, feature fusion can be performed by splicing, weighted averaging, deep learning models, etc. The above is only an exemplary description and is not intended to limit all possible situations for obtaining multimodal features. It is just not exhaustive here.

[0050] In this way, by determining structured data, unstructured data, and image data, health data can be divided according to modality, so that the most appropriate processing method can be adopted for different data types. At the same time, data is filtered through patient coding to ensure the consistency between the extracted data and the specific patient. By extracting structured data, the structured numerical values ​​obtained can accurately describe the patient's health status, and at the same time can be used as part of multimodal features to improve the overall analysis effect together with other modal data. By extracting features from unstructured data, semantic information that is difficult to capture in structured data can be supplemented, and at the same time, it can provide semantic-level supplements as part of multimodal features to enhance the expressive power of the model. By extracting features from image data, it can provide an intuitive basis for the diagnosis of the disease, and at the same time, it can be used as part of multimodal features to enhance the overall expressiveness of the data.

[0051] In some embodiments, when diabetes is the target disease, the relevant data typically includes EHR data, medical imaging data, genomic data, laboratory test data, wearable device monitoring data, patient self-reported data, and medical literature and guidelines. Among them, EHR data, genomic data, and laboratory test data are typically structured data, patient self-reported data and medical literature and guidelines are typically unstructured data, and medical imaging data are typically image data.

[0052] Specifically for diabetes, it's also possible to collect and process data on patients' lifestyle habits and exercise patterns. This type of data is considered unstructured or semi-structured behavioral data. For example, devices such as smart bracelets and smartwatches can automatically record step count, heart rate, sleep duration, exercise frequency, and activity level, with data automatically uploaded via Bluetooth or mobile apps. Applications and mini-programs also offer self-service reporting modules for diabetic patients, such as daily diet records, exercise diaries, and daily routines. Data can be collected using standardized questionnaires such as the International Physical Activity Questionnaire and the Sleep Quality Scale, and uploaded to a medical data platform. Clinicians can also record lifestyle information such as dietary frequency, tobacco and alcohol consumption, daily routines, and exercise types during patient follow-up visits. Furthermore, quantitative modeling of exercise behavior can be performed. For example, the intensity of physical activity can be expressed using metabolic equivalents of task (METs), and METs can be calculated to quantify exercise intensity and categorize activity intensity.

[0053] Furthermore, lifestyle and exercise data can be used as behavioral data and aligned with structured medical data, text records, and imaging data. For example, unified patient codes can be used to construct time series data windows, such as daily or weekly granularity, and then integrated through feature fusion or deep models such as multimodal transformers.

[0054] In some implementations, due to the inconsistency of formats among different data sources and the possible presence of missing values, noisy data, and redundant information, the data needs to be cleaned and standardized.

[0055] For example, missing data may come from experimental equipment failure, data collection problems, etc. Optional filling methods include mean filling, median filling, and K-nearest filling.

[0056] The mean filling process can be expressed by the following formula:

[0057] Where, Indicates the data points where missing values ​​need to be filled; represents the number of samples with non-missing data; Numerical value representing all non-missing data points in this feature.

[0058] The process of median filling can be expressed by the following formula:

[0059] Where, Indicates median filling.

[0060] Among them, the process of K proximity filling can be expressed by the following formula:

[0061] Where, Represents the data points after filling; Indicates the number of nearest neighbor data points used for filling; The most similar The values ​​of adjacent samples.

[0062] For example, there may be abnormal data that needs to be removed. The interquartile range method can be used to identify abnormal data.

[0063] The interquartile range method can be expressed by the following formula:

[0064] Where, represents the interquartile range; Indicates the upper quartile (75% position) of the data; Represents the lower quartile (25% position) of the data.

[0065] Specifically, a data point x is considered an outlier if it is smaller than Q1 − 1.5 × IQR or larger than Q3 + 1.5 × IQR.

[0066] For example, different data have different dimensions, and to improve the comparability of the data, normalization is required. Optional normalization methods include maximum and minimum normalization and standard score normalization.

[0067] Among them, the maximum and minimum normalization method can be expressed by the following formula:

[0068] Where, represents the normalized data points; represents the original data points; Indicates the minimum value of the feature; Indicates the maximum value of this feature.

[0069] Among them, the standard score normalization method can be expressed by the following formula:

[0070] Where, represents the normalized data points; represents the original data points; represents the mean value of the feature; Represents the standard deviation of the feature.

[0071] In particular, the standard deviation It can be expressed by the following formula:

[0072] Where, represents the number of samples; Represents each sample point.

[0073] In some embodiments, medical entities and association relationships are extracted based on multimodal features to generate knowledge entries, including: identifying medical entities and association relationships corresponding to the medical entities based on multimodal features to generate initial knowledge entries; and performing knowledge fusion on the medical entities and association relationships based on the initial knowledge entries to obtain knowledge entries.

[0074] Among them, the initial knowledge entry refers to the knowledge record generated after preliminary processing of multimodal features and identifying medical entities and their associations.

[0075] In an embodiment of the present invention, medical entities can be first extracted from multimodal features. For example, for structured data, indicator-type entities can be extracted from structured data through rule matching or data mapping. Similarly, for unstructured data, natural language processing technology can be used to identify medical entities from text features. Similarly, for image data, computer vision technology can be used to identify entities from image features. Subsequently, the association relationships between entities can be identified based on the medical entities in the multimodal features. For example, for structured data, numerical relationships can be directly extracted. Similarly, for unstructured data, semantic relationships can be extracted. Similarly, for image data, spatial relationships can be identified. Finally, the extracted medical entities and association relationships can be combined into initial knowledge entries and formatted for storage. The above is only an exemplary description and is not intended to limit all possible situations for generating initial knowledge entries, but it is not exhaustive here.

[0076] In an embodiment of the present invention, the medical entities in the initial knowledge entries can first be normalized and their associations optimized. Subsequently, knowledge graph technology can be used to construct a semantically consistent knowledge structure from the normalized medical entities and optimized associations, which can then be integrated with knowledge entries from cross-modal data. Finally, the optimized knowledge entries can be stored.

[0077] By generating initial knowledge entries, we can integrate structured data, unstructured data, and image data, enabling knowledge extraction based on multimodal features. This also clarifies medical entities and their relationships, providing a foundation for subsequent optimization and application. Knowledge fusion ensures semantic consistency between medical entities and relationships, facilitating the construction of systematic medical knowledge.

[0078] In some real-time approaches, knowledge can be directly extracted from structured data using database query and pattern matching techniques. For example, database fields can be mapped to medical concepts first, followed by data cleaning to remove noisy data, and finally entity standardization to link different names to the same concept.

[0079] In some embodiments, a deep learning model can be used to perform medical named entity recognition on unstructured data. The loss function of the deep learning model can be expressed as follows:

[0080] Where, Represents input features; Represents a sequence of entity tags; represents the parameter weight; represents the characteristic function; Represents the set of all possible category labels; is the number of the characteristic function.

[0081] In some implementations, for unstructured data, a bidirectional encoder representation model (BERT) combined with an attention mechanism can be used to extract medical relationships. The extraction process can be expressed as follows:

[0082] Where, An unstructured data vector representing the input; represents the dimension scaling factor; Represents the transpose operation of a matrix.

[0083] In some embodiments, for unstructured data, medical text needs to be automatically classified, such as medical record summaries, diabetes health guides, medication instructions, etc., and machine learning classifiers and deep learning models can be used for classification.

[0084] In some embodiments, different data sources may contain the same medical entities and relationships, but with different representations, thus requiring knowledge alignment. The alignment process can be divided into string matching and vector matching.

[0085] Specifically, for the string matching process, the Jaccard similarity can be calculated by using the time-weighted Jaccard algorithm that dynamically optimizes the time series of medical data to perform matching, which can be expressed by the following formula:

[0086] Where, It represents the time-weighted Jaccard similarity, which can measure the similarity between two medical entity strings and has a value range of [0,1]. represents the set of all medical events; Indicates a current medical event; Representing an event The difference between the occurrence time and the current time; Indicates the time attenuation coefficient, usually 0.01-0.1; Represents an indicator function, which is 1 when the condition is met; A set of words representing two medical entity strings to be compared; Indicates the number of words in common between two strings; Represents the total number of unique words in both strings.

[0087] Furthermore, for the vector matching process, matching can be performed by calculating cosine similarity, which can be expressed by the following formula:

[0088] Where, Represents the vector similarity between two medical entities, ranging from [-1, 1]. The closer the value is to 1, the more similar the entities are. Word vectors representing two medical entities; Represents the dot product of two vectors; Represents the magnitude product of two vectors.

[0089] In some embodiments, when different descriptions exist for the same entity, conflict resolution is required.

[0090] Specifically, the conflict resolution process can be performed using a confidence voting method, which can be expressed by the following formula:

[0091] Where, represents the final selected entity representation; Represents possible entity candidates; Indicates the number of data sources; Indicates the The confidence level of each data source; Indicates the Data sources about entities probability; Indicates the Detailed description or information of a data source.

[0092] In some embodiments, an artificial intelligence corpus is constructed based on knowledge items, including: establishing a knowledge graph based on the knowledge items, using medical entities as nodes and association relationships as entity relationships; generating medical text based on the knowledge graph using a pre-trained large language model; and performing corpus enhancement on the medical text to obtain an artificial intelligence corpus.

[0093] A knowledge graph is a semantic network structure that uses nodes to represent entities and edges to represent relationships between entities. In an embodiment of the present invention, nodes in a knowledge graph may represent medical entities, and edges may represent semantic relationships between entities.

[0094] In an embodiment of the present invention, medical entities can first be extracted from knowledge items as nodes, and association relationships can be extracted from knowledge items as edges. Subsequently, a graph database or knowledge graph construction tool can be used to organize the nodes and edges into a graph structure. In particular, attributes can also be added to each node and edge. Subsequently, processes such as deduplication and semantic completion can be performed to optimize the knowledge graph. The above is only an illustrative description and does not limit all possible situations for building a knowledge graph. It is just not an exhaustive list here.

[0095] In an embodiment of the present invention, a template can first be generated based on the structural definition of the knowledge graph. Subsequently, a pre-trained large language model can be used to generate medical text. For example, the node and relationship information of the knowledge graph can be input into the large language model, so that the large language model can be combined with the generated template to generate a natural language description. In particular, the language model can be fine-tuned according to the needs of the specific medical field so that the text it generates is more in line with the language style of the medical profession. The above is only an exemplary explanation and is not intended to limit all possible situations for generating medical text. It is just not exhaustive here.

[0096] In embodiments of the present invention, medical text can be enhanced. For example, data augmentation techniques can be used to generate diverse texts. Language models can also be used to further refine the generated texts. Language models can also be used to generate medical texts in other languages. The optimized medical texts are then stored in a corpus. The above is merely illustrative and does not constitute a complete list of all possible scenarios for generating an AI corpus. This is simply not an exhaustive list.

[0097] By building a knowledge graph, we can systematically store medical entities and relationships, enabling complex semantic reasoning and visualizing medical knowledge. Generating medical text can convert structured knowledge graphs into natural language for easier reading and understanding. Corpus enhancement enables the generation of more fluent and rich medical text, while also expanding the language range of the corpus, allowing for application in different regions and language environments, broadening its scope and helping to broaden its use cases.

[0098] In some embodiments, a Neo4j graph database can be used to store a knowledge graph, where the data structure includes nodes and edges. Nodes represent medical entities, such as diabetes, insulin, and hyperglycemia. Edges connect medical entities and represent relationships between them, such as treatment and triggering.

[0099] In some embodiments, a pre-trained large language model can be used to generate medical text. The process of generating medical text can be expressed by the following formula:

[0100] Where, Represents a given knowledge graph and input Generating text probability; Indicates the generated words; Represents the generated word sequence.

[0101] In some embodiments, the process of corpus enhancement for medical texts can be performed by using synonym replacement, semantic transformation, and adversarial generation.

[0102] The synonym replacement process can be expanded through the medical vocabulary, which can be expressed by the following formula:

[0103] Where, Represents the expanded text; Representing words A set of synonyms for ; Represents the original text; Representing words Synonyms or expansions of .

[0104] Among them, the process of semantic transformation can use the BERT model to generate different expressions.

[0105] Among them, the adversarial generation process can use the generative adversarial network to generate diverse corpus, which can be expressed by the following formula:

[0106] Where, Represents a generator, which is used to generate medical text corpus; Represents the discriminator, which is used to distinguish between real and synthetic corpus; Represents the distribution of real data The expected value of the sample; Represents a random noise distribution The expected value of the sample; Represents the probability that the real data is identified as true by the discriminator; represents the probability that the generated data is identified as fake by the discriminator.

[0107] In some embodiments, the method for constructing an artificial intelligence corpus based on multimodal data further includes: obtaining incremental health data corresponding to the patient code; generating incremental multimodal features corresponding to the patient code based on the incremental health data; generating incremental knowledge entries based on the incremental multimodal features; and updating the artificial intelligence corpus based on the incremental knowledge entries.

[0108] Incremental health data refers to the patient's newly added or updated health data after the artificial intelligence corpus is built. In embodiments of the present invention, incremental health data may include new examination results, new imaging data, or new pathology records.

[0109] In embodiments of the present invention, new health data for patients can be obtained periodically or in real time through linkage with data sources. In particular, the new data can be preprocessed, such as cleaning outliers, standardizing formats, and parsing unstructured text. The above is merely illustrative and does not limit all possible scenarios for obtaining incremental health data, but this is not intended to be exhaustive.

[0110] In an embodiment of the present invention, features of structured data, unstructured data, and image data can be extracted from incremental health data, and then the data of different modalities can be standardized and spliced ​​or fused into a unified incremental multimodal feature. Subsequently, the original multimodal features of the patient can be merged or replaced and fused with the newly added incremental feature vector to ensure dynamic updating of the features. The above is only an illustrative explanation and is not intended to limit all possible situations for generating incremental multimodal features. It is just that this is not an exhaustive list.

[0111] In an embodiment of the present invention, newly added medical entities and relationships can be identified based on incremental multimodal features, and then incremental knowledge entries can be generated based on the newly added medical entities and relationships. In particular, the incremental knowledge entries can be filtered to remove redundant or duplicate knowledge entries, ensuring that the newly added knowledge entries do not conflict with or duplicate existing knowledge entries in the corpus. The above is only an illustrative description and does not limit all possible situations for generating incremental knowledge entries. This is just not an exhaustive list.

[0112] In an embodiment of the present invention, the medical entities and associations in the incremental knowledge entries can first be added to the existing knowledge graph to expand the nodes and edges of the graph. In particular, the existing nodes and edges can also be dynamically modified, such as updating the attribute values ​​of the nodes or adding new edges. Subsequently, based on the updated knowledge graph, a large language model can be used to generate medical text related to the incremental knowledge entries, and the newly added medical text can be enhanced by corpus, so that the newly added medical text can be included in the artificial intelligence corpus. The above is only an exemplary explanation and is not intended to limit all possible situations for updating the artificial intelligence corpus. It is just that this is not an exhaustive list.

[0113] In this way, by acquiring incremental health data, we can ensure that the data stored in the corpus always reflects the patient's latest health status. By generating incremental multimodal features, we can capture the changing trends of the patient's health status. By generating incremental knowledge entries, we can continuously enrich the medical knowledge in the corpus, keep the knowledge dynamically updated, and reflect the changing trends of the patient's health status. By updating the AI ​​corpus, the corpus can always be synchronized with the patient's health data and medical knowledge.

[0114] In some embodiments, blood sugar, heart rate, number of steps, sleep, etc. can be collected through devices such as smart bracelets, blood glucose meters, and insulin pumps, and the data can be uploaded in real time. At the same time, the changes in the data collected in real time can be synchronously pushed to the doctor's side, and the doctor can revise the automatically annotated corpus. The revised results are returned to the corpus and used as training samples to enhance the model capabilities. Furthermore, through the interface docking of the hospital's electronic medical record system, real-time synchronization of patient medical records and test data can be achieved, and the medical record summary, diagnostic process, medication strategy description, etc. in the patient corpus can be automatically updated. In particular, the application of the health follow-up applet can allow patients to actively fill in diet, medication, physical discomfort, etc. every day, and achieve real-time synchronization with the corpus system.

[0115] In some embodiments, the method for constructing an artificial intelligence corpus based on multimodal data further includes: performing quality assessment on the artificial intelligence corpus in a preset manner to generate a corpus quality assessment result; identifying high-uncertainty data in the artificial intelligence corpus based on the corpus quality assessment result, and optimizing the corpus using an active learning method to obtain an optimized artificial intelligence corpus; performing expert annotation correction on the optimized artificial intelligence corpus, and updating the artificial intelligence corpus based on the correction results.

[0116] The corpus quality assessment results are outputs of quantitative or semantic analysis of the data quality in the artificial intelligence corpus. In embodiments of the present invention, the corpus quality assessment results may include data accuracy, data consistency, data coverage, data interpretability, generalization ability, etc.

[0117] In an embodiment of the present invention, specific quantitative indicators can be first set for each evaluation criterion. Subsequently, an algorithm can be used to automatically perform quality assessment. For example, a natural language processing quality detection model can be used to detect the language fluency and standardization of medical terminology in the corpus text. Alternatively, a knowledge graph comparison method can be used to compare the medical entities and relationships in the corpus with a standard medical knowledge base. The above is merely an example and does not limit all possible scenarios for generating corpus quality assessment results. This is not intended to be exhaustive.

[0118] In an embodiment of the present invention, data with low credibility or high ambiguity can first be screened based on the quality assessment results, and the screened data can be treated as high-uncertainty data. Subsequently, the high-uncertainty data can be selected as the key optimization target, and the model combined with manual active learning can be used to achieve a cyclic optimization of the artificial intelligence corpus. The above is only an example and does not limit all possible scenarios for obtaining an optimized artificial intelligence corpus. This is just a non-exhaustive list.

[0119] In an embodiment of the present invention, medical experts can review and annotate the optimized corpus, and the revised data annotated by the experts can be recorded as the correction results. Subsequently, the AI ​​corpus can be updated based on the correction results. For example, this can replace existing low-quality data, add new data, or delete redundant or erroneous data.

[0120] By generating corpus quality assessment results, we can systematically identify problems within the corpus and provide clear targets for subsequent corpus optimization, expansion, and revision. Active learning enables focused optimization of high-uncertainty data, significantly improving corpus quality. Expert annotation and revision can eliminate errors or ambiguities in the corpus, ensuring the accuracy and authority of the corpus.

[0121] In some implementations, corpus quality assessment involves multiple dimensions, which may include indicators such as accuracy, consistency, coverage, interpretability, and generalization ability.

[0122] Accuracy can measure the correctness of the annotated data in the corpus. For example, the calculation process of accuracy can be expressed by the following formula:

[0123] Where, represents the accuracy score, which measures the overall accuracy of the model; Indicates the precision rate, which indicates the proportion of the corpus predicted correctly by the model to all predictions; Recall is the ratio of the predicted correct corpus to all correct corpus.

[0124] Furthermore, the accuracy It can be expressed by the following formula:

[0125] Where, (True Positives) indicates the number of correctly classified positive samples; (FalsePositives) indicates the number of incorrectly classified positive samples.

[0126] Furthermore, the recall rate It can be expressed by the following formula:

[0127] Where, (False Negatives) indicates the number of positive samples that are misclassified as negative samples.

[0128] Among them, consistency can measure whether different experts maintain consistency in the annotation of the same data. For example, the calculation process of consistency can be expressed by the following formula:

[0129] Where, Represents the consistency coefficient, ranging from [-1,1], and the higher the value, the better the consistency; represents the matching probability between actual annotators; represents the probability of label matching in random cases.

[0130] In particular, when , it indicates that the annotations of the corpus have high consistency.

[0131] Among them, coverage can measure whether the corpus contains enough diabetes medical terms. For example, coverage The calculation process can be expressed by the following formula:

[0132] Where, represents the set of medical terms in the corpus; Represents a set of standard medical terms; Indicates the number of correctly matched medical terms in the corpus.

[0133] Among them, interpretability can measure whether the corpus supports effective interpretation by medical experts and artificial intelligence systems. For example, the calculation process of interpretability can be expressed by the following formula:

[0134] Where, Represents the interpretability score of the corpus; Represents cosine similarity.

[0135] Furthermore, the cosine similarity can be expressed by the following formula:

[0136] Where, Indicates medical terminology; Indicates standard medical terminology.

[0137] Among them, generalization ability can measure whether the corpus is applicable to different medical scenarios. For example, the calculation process of generalization ability can be expressed by the following formula:

[0138] Where, Represents the generalization ability score, and the higher the value, the more scenarios the corpus is applicable to; Indicates the number of test sets; Represents the test set The terms in the corpus are correctly covered; Represents the test set The total number of terms.

[0139] In some embodiments, based on the corpus quality assessment results, high-uncertainty data in the artificial intelligence corpus is identified, and the corpus is optimized using active learning methods. In the process of obtaining the optimized artificial intelligence corpus, the methods for identifying high-uncertainty data include: uncertainty scoring based on prediction probability, scoring based on information entropy, and scoring based on model inconsistency.

[0140] Specifically, by calculating the uncertainty score of the predicted probability, we can identify high-uncertainty data and give priority to labeling, thereby improving the quality of the corpus. For example, the calculation process of the uncertainty score can be expressed by the following formula:

[0141] Where, Represents a sample uncertainty score; Represents the model prediction sample The most likely label confidence level.

[0142] In particular, data with high uncertainty will be labeled first to reduce erroneous samples.

[0143] Specifically, the identification can be performed by information entropy. For example, the greater the information entropy, the more ambiguous the model's classification of the sample is.

[0144] Specifically, identification can be based on model inconsistency scores. For example, using multiple models or multiple versions of the same model to predict the same sample, if the prediction results differ widely, it indicates high uncertainty. Common scoring methods are voting entropy or relative entropy.

[0145] In some embodiments, expert annotation corrections are performed based on the optimized artificial intelligence corpus, and the process of updating the artificial intelligence corpus based on the correction results can introduce expert annotation corrections to improve credibility. Exemplarily, the evaluation can be performed using an expert consensus score, that is, when the expert consensus score falls below a threshold, a manual review mechanism is triggered. In particular, the expert consensus score can be expressed by the following formula:

[0146] Where, represents the expert agreement score; represents the number of experts; Indicates the Expert ratings.

[0147] The embodiment of the present invention provides an artificial intelligence corpus construction device based on multimodal data, such as Figure 2 As shown, the device may include: a data acquisition module 201, used to acquire health data corresponding to the disease type; a feature extraction module 202, used to generate multimodal features corresponding to the patient code based on the health data; an entry generation module 203, used to extract medical entities and association relationships based on the multimodal features to generate knowledge entries; a corpus construction module 204, used to construct an artificial intelligence corpus based on the knowledge entries.

[0148] In some embodiments, the data acquisition module 201 includes: a source screening submodule for screening available data sources according to disease types; and a data pulling submodule for acquiring health data according to available data sources.

[0149] In some embodiments, the feature extraction module 202 includes: a data determination submodule, which is used to determine structured data, unstructured data and image data corresponding to the patient code from the health data; a structured determination submodule, which is used to determine structured values ​​based on the structured data; an unstructured determination submodule, which is used to determine unstructured features based on the unstructured data; an image determination submodule, which is used to determine image features based on the image data using a pre-trained model; and a feature fusion submodule, which is used to fuse structured values, unstructured features and image features to obtain multimodal features.

[0150] In some embodiments, the entry generation module 203 includes: an initial identification submodule, which is used to identify medical entities and the association relationships corresponding to the medical entities based on multimodal features, and generate initial knowledge entries; a knowledge fusion submodule, which is used to perform knowledge fusion on medical entities and association relationships based on the initial knowledge entries to obtain knowledge entries.

[0151] In some embodiments, the corpus construction module 204 includes: a graph establishment submodule, which is used to establish a knowledge graph based on knowledge items, using medical entities as nodes and association relationships as relationships; a text generation submodule, which is used to generate medical text based on the knowledge graph using a pre-trained large language model; and a corpus enhancement submodule, which is used to perform corpus enhancement on medical text to obtain an artificial intelligence corpus.

[0152] In some embodiments, the artificial intelligence corpus construction device based on multimodal data further includes: an incremental health data module 205 ( Figure 2), for obtaining incremental health data corresponding to the patient code; incremental feature module 206 ( Figure 2 ), for generating incremental multimodal features corresponding to patient codes according to incremental health data; incremental entry module 207 ( Figure 2 ), used to generate incremental knowledge entries based on incremental multimodal features; corpus updating module 208 ( Figure 2 ), which is used to update the artificial intelligence corpus based on incremental knowledge entries.

[0153] In some embodiments, the apparatus for constructing an artificial intelligence corpus based on multimodal data further includes: a quality assessment module 209 ( Figure 2 ), used to perform quality assessment according to a preset method based on the artificial intelligence corpus and generate a corpus quality assessment result; a corpus optimization module 210 ( Figure 2 ), which is used to identify high-uncertainty data in the artificial intelligence corpus based on the corpus quality assessment results, optimize the corpus using active learning, and obtain the optimized artificial intelligence corpus; the expert annotation module 211 ( Figure 2 ), which is used to perform expert annotation correction based on the optimized artificial intelligence corpus and update the artificial intelligence corpus according to the correction results.

[0154] For the description of specific functions and examples of each module and submodule of the device according to the embodiment of the present invention, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0155] The artificial intelligence corpus construction device based on multimodal data in the embodiment of the present invention can effectively screen out relevant health data by acquiring health data corresponding to disease types and utilizing the correspondence between data and disease types. By generating multimodal features corresponding to patient codes, information from different data modalities can be integrated, avoiding the limitations of single-modal data. By generating knowledge items, complex multimodal data can be converted into structured knowledge. By constructing an artificial intelligence corpus, the storage and retrieval of knowledge items can be facilitated, improving the utilization efficiency of medical knowledge. At the same time, it can be used as a basic data source for the development of applications such as question-answering systems and recommendation systems.

[0156] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0157] According to an embodiment of the present invention, the present invention further provides an electronic device and a readable storage medium.

[0158] The embodiment of the present invention provides a scenario diagram of a method for constructing an artificial intelligence corpus based on multimodal data, such as Figure 3 shown.

[0159] As previously mentioned, the method for constructing an artificial intelligence corpus based on multimodal data provided by an embodiment of the present invention is applied to electronic devices. The electronic devices are intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0160] Specifically, the electronic device can perform the following operations: Obtain health data corresponding to the disease type; based on the health data, generate multimodal features corresponding to the patient code; based on the multimodal features, extract medical entities and association relationships to generate knowledge items; based on the knowledge items, build an artificial intelligence corpus.

[0161] It should be understood that Figure 3 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 3 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present invention.

[0162] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0163] like Figure 4 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. RAM 403 may also store various programs and data required for the operation of device 400. Computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.

[0164] Various components in device 400 are connected to I / O interface 405, including an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0165] The computing unit 401 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the medical model training method based on multimodal data. For example, in some embodiments, the medical model training method based on multimodal data can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the medical model training method based on multimodal data described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the medical model training method based on multimodal data in any other appropriate manner (eg, by means of firmware).

[0166] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0167] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0168] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0170] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0171] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0172] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0173] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for constructing an artificial intelligence corpus based on multimodal data, characterized in that: include: Obtain health data corresponding to disease types; generating, based on the health data, a multimodal feature corresponding to the patient code; Extracting medical entities and association relationships based on the multimodal features to generate knowledge items; An artificial intelligence corpus is constructed based on the knowledge items.

2. The method according to claim 1, characterized in that The obtaining of health data corresponding to the disease type includes: Based on the disease type, available data sources were screened; The health data is obtained according to the available data source.

3. The method according to claim 1, characterized in that Generating a multimodal feature corresponding to the patient code based on the health data includes: Determining, from the health data, structured data, unstructured data, and image data corresponding to the patient code; Determining a structured value based on the structured data; Determining unstructured features based on the unstructured data; Determining image features using a pre-trained model based on the image data; The structured values, the unstructured features and the image features are fused to obtain the multimodal features.

4. The method according to claim 1, wherein The extracting of medical entities and association relationships based on the multimodal features to generate knowledge items includes: identifying the medical entity and the association relationship corresponding to the medical entity based on the multimodal features, and generating initial knowledge entries; Based on the initial knowledge items, the medical entities and the association relationships are fused to obtain the knowledge items.

5. The method according to claim 1, wherein The step of constructing the artificial intelligence corpus based on the knowledge items includes: According to the knowledge items, the medical entities are used as nodes and the association relationships are used as relationships to establish a knowledge graph; Generate medical text based on the knowledge graph using a pre-trained large language model; Corpus enhancement is performed on the medical text to obtain the artificial intelligence corpus.

6. The method according to claim 1, characterized in that Also includes: Acquiring incremental health data corresponding to the patient code; generating, based on the incremental health data, an incremental multimodal feature corresponding to the patient code; generating incremental knowledge entries according to the incremental multimodal features; The artificial intelligence corpus is updated according to the incremental knowledge entries.

7. The method according to claim 1, characterized in that Also includes: Performing quality assessment according to a preset method based on the artificial intelligence corpus to generate a corpus quality assessment result; According to the corpus quality assessment results, high uncertainty data in the artificial intelligence corpus is identified, and the corpus is optimized using an active learning method to obtain an optimized artificial intelligence corpus; According to the optimized artificial intelligence corpus, expert annotation correction is performed, and the artificial intelligence corpus is updated according to the correction results.

8. An artificial intelligence corpus construction device based on multimodal data, characterized in that: include: A data acquisition module is used to acquire health data corresponding to the disease type; a feature extraction module, configured to generate multimodal features corresponding to the patient code based on the health data; An entry generation module, configured to extract medical entities and association relationships based on the multimodal features and generate knowledge entries; The corpus construction module is used to construct the artificial intelligence corpus based on the knowledge items.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are for causing a computer to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Multilingual health information intelligent customization method and system based on user portrait

    CN121366688A

  • Multi-lingual health information intelligent customization method and system based on user portrait

    CN121366688B