Medical ultrasound field knowledge base construction method, related equipment and program product
Through the multimodal unsupervised pre-training model, the feature extraction and alignment of medical text data and other modal data is solved, and the problem of limited knowledge coverage in traditional methods is achieved, achieving more comprehensive knowledge expression and efficient knowledge base construction.
Patent Information
- Application Number
- CN202510831990.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-07
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The traditional medical ultrasound knowledge base construction method relies on manual setting rules and logic, and cannot integrate multimodal information, resulting in limited knowledge coverage.
A multimodal unsupervised pre-training model is used to extract and align medical text data and other modal data, integrate text features with other modal features, and establish a knowledge base in the field of multimodal medical ultrasound.
It improves the knowledge coverage and construction efficiency of the knowledge base, achieves more comprehensive knowledge expression, and reduces the dependence on manual rules formulation.
Smart Images

Figure CN120338078A_ABST
Abstract
Description
[0001] This application claims the priority of a domestic application titled "Method for Constructing a Knowledge Base in the Field of Medical Ultrasound, Related Devices and Program Products", with the application number 202510582273.X, which was filed with the China National Patent Office on May 7, 2025. The entire content of this application is incorporated herein by reference. Technical Field
[0002] This application relates to the technical field of knowledge base construction, and more specifically, to a method for constructing a knowledge base in the field of medical ultrasound, related devices and program products. Background Art
[0003] In the field of medical and health, contrast-enhanced ultrasound (CEUS) is a diagnostic imaging technique used for the liver and other organs. In the field of medical ultrasound, there are a large number of medical guidelines, medical records, medical images, physiological data, and examination and test results. These information describe different processes or bases of diagnosis and treatment. Constructing a professional knowledge base for the field of medical ultrasound can not only provide support for clinical decision-making, but also serve as an important resource for education and research.
[0004] Traditional knowledge base construction methods generally adopt rule-based construction methods. By experts defining rules and logical relationships, the text information in the field is converted into structured knowledge to construct the knowledge base. However, this method relies on manually setting clear rules and logics, and can only organize text-modal information, unable to integrate knowledge information in other modalities, and the knowledge coverage of the constructed domain knowledge base is limited. Summary of the Invention
[0005] In view of the above problems, this application is proposed to provide a method for constructing a knowledge base in the field of medical ultrasound, related devices and program products, so as to improve the efficiency and knowledge coverage of knowledge base construction and improve the quality of the knowledge base. The specific solutions are as follows:
[0006] In a first aspect, a method for constructing a knowledge base in the field of medical ultrasound is provided, including:
[0007] Obtain multi-modal medical ultrasound data, where the multi-modal medical ultrasound data includes medical text data and its associated other modal data;
[0008] Perform entity extraction on the medical text data to obtain target entities, and extract the relationships between target entities in the medical text data;
[0009] Call a multi-modal unsupervised pre-training model to extract features from the medical text data and its associated other modal data, obtaining text features and other modal features. The multi-modal unsupervised pre-training model is configured to extract features of multi-modal data and perform modal alignment;
[0010] Fuse the text features and other modal features to obtain fused features, and establish an association relationship between the fused features and the target entities included in the medical text data;
[0011] Add the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities to the knowledge base in the field of medical ultrasound.
[0012] In a second aspect, there is provided a device for constructing a knowledge base in the field of medical ultrasound, including:
[0013] A data acquisition unit for acquiring multi-modal medical ultrasound data, where the multi-modal medical ultrasound data includes medical text data and its associated other modal data;
[0014] An entity and relationship extraction unit for performing entity extraction on the medical text data to obtain target entities, and extracting the relationships between target entity pairs in the medical text data;
[0015] A multi-modal feature extraction unit for calling a multi-modal unsupervised pre-training model to extract features from the medical text data and its associated other modal data, obtaining text features and other modal features. The multi-modal unsupervised pre-training model is configured to extract features of multi-modal data and perform modal alignment;
[0016] A multi-modal feature fusion unit for fusing the text features and other modal features to obtain fused features, and establishing an association relationship between the fused features and the target entities included in the medical text data;
[0017] A knowledge base update unit for adding the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities to the knowledge base in the field of medical ultrasound.
[0018] In a third aspect, there is provided an electronic device, including: a memory and a processor;
[0019] The memory is used to store a program;
[0020] The processor is used to execute the program to implement each step of the method for constructing a knowledge base in the field of medical ultrasound provided in the first aspect of the embodiments of the present application.
[0021] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step of the method for constructing a knowledge base in the field of medical ultrasound provided in the first aspect of the embodiments of the present application is implemented.
[0022] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, each step of the method for constructing a knowledge base in the field of medical ultrasound provided in the first aspect of the embodiments of the present application is implemented.
[0023] With the above technical solutions, the present application obtains multi-modal medical ultrasound data, including medical text data and data of other modalities associated with the medical text data. Examples include image data associated with medical text (such as ultrasound examination reports), audio-video data, etc. On this basis, entity extraction is performed on the medical text data therein to obtain target entities, and the relationships between the target entities in the medical text data are extracted, so as to obtain the extracted target entities and the relationship information between the target entities. Further, the present application is also configured with a multi-modal unsupervised pre-training model, which is configured to extract the features of multi-modal data and perform modality alignment. With the help of this model, the features of the medical text data and its associated other modality data are extracted to obtain aligned text features and other modality features. The text features and other modality features are fused to obtain fused features, and the fused features are associated with the target entities included in the medical text data. Furthermore, the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities can be added to the knowledge base in the field of medical ultrasound. The knowledge base contains not only text modality information but also other modality information, improving the knowledge coverage of the knowledge base, thereby providing a more comprehensive knowledge expression. And, the method of the present application does not rely on manually formulated rules, improving the construction efficiency of the knowledge base. Description of the Drawings
[0024] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0025] Figure 1 It is a schematic diagram of an implementation system architecture for the method of constructing a knowledge base in the field of medical ultrasound provided by the embodiments of the present application;
[0026] Figure 2 It is a schematic diagram of the process flow of a method for constructing a knowledge base in the field of medical ultrasound provided by the embodiments of the present application;
[0027] Figure 3 It exemplifies a schematic diagram of the training process of a multi-view semi-supervised training model based on an auxiliary objective function;
[0028] Figure 4 Illustrates a schematic structural diagram of a relation extraction model based on distant supervision;
[0029] Figure 5 Illustrates a process of aligning text and image modality features;
[0030] Figure 6 Illustrates a schematic flow diagram of entity linking;
[0031] Figure 7 Illustrates a schematic processing process diagram of an entity linking prediction model;
[0032] Figure 8 Schematic structural diagram of a knowledge base construction device in the field of medical ultrasound provided by an embodiment of the present application;
[0033] Figure 9 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0035] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0036] The data involved in the technical solutions of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related regulations.
[0037] The present application provides a method for constructing a knowledge base in the field of medical ultrasound, which can be applied to a system architecture as shown in Figure 1 The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 illustrated by including one server as an example).
[0038] Either the terminal 100 or the server 200 can be used alone to execute the method for constructing a knowledge base in the field of medical ultrasound provided by the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the method for constructing a knowledge base in the field of medical ultrasound provided in the embodiments of the present application.
[0039] The terminal 100 in the embodiments of the present application may be a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not make any restrictions on this.
[0040] The method for constructing a knowledge base in the field of medical ultrasound of the present application can be applied to the scenario of constructing knowledge bases in multiple sub-fields in the field of medical ultrasound, including but not limited to: the construction of a professional knowledge base for liver ultrasound, etc. The construction process of the professional knowledge base for liver ultrasound will be used as an example for illustration in the following text.
[0041] The embodiments of the present application provide a method for constructing a knowledge base in the field of medical ultrasound. Taking the application of this method to a computer device as an example, the computer device may specifically be Figure 1 the terminal 100 in Figure 2 or a system composed of the terminal 100 and the server 200. Referring to
[0042] Step S100: Obtain multi-modal medical ultrasound data, where the multi-modal medical ultrasound data includes medical text data and other associated modal data.
[0043] Specifically, the multi-modal medical ultrasound data includes data in text modality and other modalities, such as image modality, audio-video modality, etc. The multi-modal medical ultrasound data may be multi-source heterogeneous data, including but not limited to the following five types of data sources:
[0044] Anatomy and physiology knowledge: mainly from medical textbooks. Taking the knowledge of liver anatomy and physiology as an example, it is mainly the normal anatomical structure of the liver and the physiological mechanisms of the normal functions of the liver.
[0045] Ultrasound examination techniques and equipment parameters: including the basic principles of ultrasonic imaging, Doppler ultrasound, and the parameter settings and optimization methods of various corresponding imaging modes. Mainly obtained from device-related technical manuals.
[0046] Disease diagnosis information: including common diseases and their ultrasound manifestations, typical ultrasound features of various cases, and grading diagnosis criteria. Mainly relies on large-scale clinical case records and diagnostic standard guidelines.
[0047] Image data and annotation information: including a large number of labeled ultrasound image data, as well as image data sets of various cases for training AI models or reference analysis, etc. These can be obtained from the imaging examination and test reports of medical institutions.
[0048] Clinical case data: mainly ultrasound examination reports and diagnostic results, as well as case follow-up data required for establishing prediction models for disease progression or treatment effectiveness.
[0049] Based on the multi-modal medical ultrasound data of the above examples, a knowledge base in the field of medical ultrasound is constructed. The knowledge base includes the fusion feature representation of entities, relationships between entities, and multi-modal data associated with entities. Among them, the entity types include but are not limited to patient information, medical terms, measurement indicators, imaging materials, text descriptions, etc., and the relationship patterns involve causal relationships, subordination relationships, spatio-temporal relationships, operation relationships, etc. In addition, for each entity, necessary attributes can be defined. For example, for disease entities, related attributes include incidence rate, risk factors, treatment methods, etc.
[0050] By defining the knowledge structure, a clear goal and framework are provided for subsequent information extraction, ensuring that the extracted information is organized, utilizable, and in line with the professional field.
[0051] Step S110: Perform entity extraction on the medical text data to obtain target entities, and extract the relationships between target entity pairs in the medical text data.
[0052] Specifically, for the text-modal data (i.e., medical text data) in the multi-modal medical ultrasound data obtained in the previous step, an entity extraction model is used for entity extraction to obtain the target entities included in the medical text data. Further, extract the relationships between target entity pairs in the medical text data.
[0053] Step S120: Invoke a multi-modal unsupervised pre-training model to perform feature extraction on the medical text data and its associated other-modal data, to obtain text features and other-modal features. The multi-modal unsupervised pre-training model is configured to extract features of multi-modal data and perform modal alignment.
[0054] In order to make full use of the complementarity between different modalities and achieve more comprehensive and accurate knowledge expression, this application also pre-configures a multi-modal unsupervised pre-training model. This model can perform feature extraction on the input multi-modal data and achieve feature alignment across modalities, that is, associate data of different modalities (such as text, images, structured data, etc.) to the same semantic space.
[0055] In this step, by invoking the multi-modal unsupervised pre-training model, perform feature extraction on the medical text data and its associated other-modal data, to obtain the text features corresponding to the medical text data, and other-modal features corresponding to the other-modal data.
[0056] Taking an ultrasound examination report as an example, it contains both ultrasound images and corresponding examination report texts. Through a multi-modal unsupervised pre-training model, the image features of the ultrasound images and the text features of the examination report texts can be extracted respectively. Moreover, since the multi-modal unsupervised pre-training model has the ability of cross-modal feature alignment, it can ensure that the extracted image features and text features are aligned, thus facilitating feature fusion in subsequent steps.
[0057] Step S130: Fuse the text features and other modal features to obtain fused features, and establish an association relationship between the fused features and the target entities included in the medical text data.
[0058] Specifically, for the text features and other modal features extracted in the previous step, fusion processing can be performed to obtain fused features. Through the above-mentioned feature extraction, alignment and fusion processing, the complementarity between different modalities can be fully utilized to achieve more comprehensive and accurate knowledge representation. Exemplarily, semantic associations can be established between the lesion features in the image, the text description, and the numerical parameters, which can improve the accuracy and interpretability of the diagnosis results.
[0059] The fused features obtained in this step correspond to the medical text data and other associated modal data. Since knowledge information is stored in the form of entity objects in the knowledge base, for the convenience of building the knowledge base, an association relationship can be established between the fused features and the target entities included in the medical text data in this step.
[0060] Step S140: Add the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities to the medical ultrasound domain knowledge base.
[0061] Through the processing in the foregoing steps, the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities (the fused features are used as the feature representations of the medical text data where the target entities are located and other associated modal data) have been obtained from the multi-modal medical ultrasound data, and the above types of information can be added to the medical ultrasound domain knowledge base.
[0062] The method provided by the embodiments of the present application obtains multimodal medical ultrasound data, including medical text data and data of other modalities associated with the medical text data. Examples include image data associated with medical text (such as ultrasound examination reports), audio-visual data, etc. On this basis, entity extraction is performed on the medical text data therein to obtain target entities, and the relationships between target entity pairs in the medical text data are extracted, so as to obtain the extracted target entities and the relationship information between target entity pairs. Further, the present application is also configured with a multimodal unsupervised pre-training model, which is configured to extract the features of multimodal data and perform modality alignment. With the help of this model, feature extraction is performed on the medical text data and its associated data of other modalities to obtain aligned text features and features of other modalities. The text features and features of other modalities are fused to obtain fused features, and an association relationship is established between the fused features and the target entities included in the medical text data. Furthermore, the target entities, the relationships between target entity pairs, and the fused features associated with the target entities can be added to the knowledge base in the field of medical ultrasound. The knowledge base contains not only text modality information but also information of other modalities, improving the knowledge coverage of the knowledge base, thereby providing a more comprehensive knowledge representation. Moreover, the method of the present application does not rely on manually formulated rules, improving the construction efficiency of the knowledge base.
[0063] In some embodiments of the present application, the process of performing entity extraction on the medical text data in the foregoing step S110 to obtain target entities is described.
[0064] Specifically, an entity extraction model can be used to perform entity extraction on the medical text data to obtain the target entities in the medical text data.
[0065] The medical text data can be pre-divided into basic units. For example, it can be divided according to sentences, and then entities are extracted from each sentence respectively to obtain the target entities corresponding to each sentence.
[0066] Among them, the entity extraction model can be trained using sample data marked with entity extraction results.
[0067] The training process of a conventional entity extraction model requires a large amount of manual annotation, with slow efficiency and high cost. In this embodiment, a weak supervision learning paradigm is introduced, which does not rely on a large amount of precisely annotated data for training. The entity extraction model is trained using a multi-view semi-supervised learning method.
[0068] Multi-view Semi-supervised Learning is a machine learning method that combines multi-view learning and semi-supervised learning, using data from different views and limited labeled data to improve the learning effect of the model. Among them, multi-view learning refers to learning from different views (i.e., different feature sets or data representations) of the same object or instance. Each view can be regarded as a different description of the same thing. For example, in medical diagnosis, different views may include the description of the patient's symptoms, laboratory test results, etc. By integrating these different views, the model can obtain more comprehensive information and thus improve its performance. Semi-supervised learning is a method of training using a small amount of labeled data and a large amount of unlabeled data. Its core idea is that even without labels, there is still valuable information in the unlabeled data, such as the structural features of the data distribution.
[0069] The role of multi-view semi-supervised training is to improve the "representation" learning of the model. Refer to Figure 3 , which exemplifies a multi-view semi-supervised training model based on an auxiliary objective function, including a Primary Prediction Module and several Auxiliary Prediction Modules. The auxiliary prediction modules can learn from the predictions of the primary prediction module because the primary prediction module has better and unrestricted-view inputs. Although the inputs of the auxiliary prediction modules are restricted input samples, after training, they can still learn to make the predictions of the primary prediction module, so they can improve the quality of the "representation". This in turn improves the entire model because they share an Encoder module. Figure 3 exemplifies using BERT as the encoding module. The pre-trained language model of BERT generates word vectors representing context semantic information and automatically extracts a large number of word-level and semantic-level features in the text. Entity label sequences are decoded in the primary prediction module and the auxiliary prediction modules, and then entity recognition and extraction are performed. Figure 3 The auxiliary prediction module shown can adopt a small neural network. For example, a Conditional Random Field (CRF) model is used to explicitly consider context associations at the output end.
[0070] Figure 3 exemplifies a Cross-View Training (CVT) strategy. Cross-View Training belongs to a semi-supervised algorithm that uses unlabeled and labeled data to distributively train word representations. When CVT learns data without labels, through several auxiliary prediction modules, each auxiliary prediction module receives an intermediate representation as input and outputs a label distribution . Each is selected to use only a part of the input , the specific selection may depend on the task and the model architecture. The auxiliary prediction module is only used during training; the predictions in the model inference phase come from the .
[0071] Figure 3 In the model structure shown, the main prediction module and the auxiliary prediction module share the same encoding module, and the encoding module adopts the BERT structure. The main prediction module makes predictions based on the training data carrying labels and is trained with the goal of outputting fitted semantic labels; the auxiliary prediction module generates training corpus by adding noise from multiple perspectives randomly (that is, adding Gaussian noise to the image data, or randomly occluding some areas for random noise addition; for text data, it is randomly replacing words or deleting certain words). For the training data without labels, they are processed separately by the main prediction module and the auxiliary prediction module. The input of the auxiliary prediction module is a partially restricted perspective (for example, only seeing a part of the sentence), to train and fit the prediction results obtained by the main prediction module for the input of the complete perspective. Since the main prediction module and the auxiliary prediction module share the encoding module, the above cross-perspective training method can improve the "representation" learning of the model.
[0072] After the entity extraction model is trained, the entity extraction task can be performed through the main prediction module.
[0073] Figure 3 The right side in shows a specific example of performing the entity extraction task using the trained entity extraction model.
[0074] In this embodiment, by introducing the multi-perspective semi-supervised learning method to train the entity extraction model, the purpose of reducing the manual annotation cost can be achieved, and the ability to process large-scale data can be improved.
[0075] In some embodiments of the present application, the process of extracting the relationship between target entity pairs in the medical text data in the foregoing step S110 is described.
[0076] This embodiment provides an end-to-end relation extraction technology based on distant supervision. Specifically, the trained distant supervision relation extraction model can be used to extract the relationship between target entity pairs in the medical text data.
[0077] Traditional supervised schemes require a large amount of manual annotation of data and cannot achieve fast and effective knowledge construction. In order to quickly obtain rich entity relationships, an end-to-end relation extraction technology based on distant supervision is proposed.
[0078] Distant supervision-based relation extraction is a technique that automatically annotates training data using an external knowledge base and learns the relationships between entities from it. This method generates labels for a large amount of unannotated text by assuming that the entity relationships known in the knowledge base also exist in the text, and then uses these labels to train machine learning models.
[0079] Traditional distant supervision methods assume that all sentences containing the same pair of entities express the predefined relationships in an external knowledge base (such as the FreeBase knowledge base, etc.), but in reality, a large number of sentences are noisy (do not express the predefined relationships), resulting in low-quality annotated data.
[0080] Therefore, in this embodiment, a distant supervision relation extraction model based on a sentence-level attention mechanism (Sentence-Level Attention) is provided. In the model, dynamic weight allocation is performed through the sentence-level attention mechanism, that is, the attention weights are calculated for multiple sentences (i.e., the sentence set) corresponding to each entity pair, highlighting the key sentences expressing the predefined relationships and suppressing the noisy sentences.
[0081] This embodiment introduces a training method for a distant supervision relation extraction model, which specifically may include the following steps:
[0082] S1. For sample entities A and B, extract the sentences that contain both sample entities A and B from the medical text data to form a sentence set, and identify the relationship r existing between sample entities A and B from the external knowledge base, and use the relationship r as the relationship type label of the sentence set.
[0083] S2. Send the sentence set and the syntactic tree representation of each sentence in it into the distant supervision relation extraction model, calculate the relation vector representation of each sentence and the attention weight of each sentence, and perform weighted summation on the relation vector representations of each sentence according to the attention weight of the sentence to obtain a weighted relation vector representation, and predict the relation type corresponding to the sentence set based on the weighted relation vector representation.
[0084] S3. Calculate the value of the loss function according to the predicted relation type of the sentence set and the relation type label of the sentence set, and update the model parameters according to the value of the loss function until the training end condition is reached.
[0085] The model training method introduced in this embodiment not only includes a sentence set at the input level, but further includes the syntactic tree representation of each sentence. The syntactic tree representation can cover the information at the semantic space level of the sample entity pair, which is beneficial to calculating the relational vector representation of each sentence more effectively. Inside the model, by calculating the sentence-level attention weights, the model will learn to assign higher weights to the sentences that truly express the predefined relationship r, so as to reduce the influence of noisy sentences during weighted aggregation and reduce the impact of noisy sentences in the sentence set on model training.
[0086] After obtaining the trained distant supervision relation extraction model, the medical text data containing the target entity pair and the syntactic tree representation of the medical text data can be sent into the distant supervision relation extraction model to obtain the relationship between the target entity pairs output by the model, thus completing the relation extraction task.
[0087] Refer to Figure 4 , which illustrates a schematic structural diagram of a distant supervision relation extraction model.
[0088] As Figure 4 shown, the distant supervision relation extraction model may include the following network structures:
[0089] An input layer for inputting a sentence set and the syntactic tree representation of each sentence therein.
[0090] Figure 4 In the example sentence 1, "[DIS]" indicates that entity A, "Krabbe disease", belongs to the disease type, and "SYM" indicates that entity B, "lower limb spasm", belongs to the symptom type.
[0091] By collecting each sentence that simultaneously contains sample entities A and B, a sentence set is formed. Based on each sentence, the corresponding syntactic tree representation is generated. The sentence set and the syntactic tree representation are sent into the model through the input layer.
[0092] An encoding layer includes a text encoding module and a graph neural network module. The text encoding module is used to encode each sentence to obtain the sentence-level vector representations of sample entities A and B, and the graph neural network module is used to encode the syntactic tree representation of each sentence to obtain the semantic space-level vector representations of sample entities A and B.
[0093] Figure 4 For the described encoding layer, the BERT network is used as the text encoding module, and other text encoding structures can also be adopted in addition. The graph neural network module can encode the syntactic tree representation of each sentence through graph convolution to obtain the semantic space-level vector representations of sample entities A and B.
[0094] The relationship representation layer is used to calculate the relationship vector representation of each sentence based on the sentence-level vector representations and semantic space-level vector representations of Entities A and B.
[0095] Specifically, based on the sentence-level vector representations of sample Entities A and B, the relationship vector representation at the sentence level can be calculated. For example, v1 relation =v A -v B , where v A and v B respectively represent the sentence-level vector representations of sample Entities A and B.
[0096] Based on the semantic space-level vector representations of sample Entities A and B, the relationship vector representation at the semantic space level can be calculated. For example, through an aggregation method such as max-pool, the semantic space-level vector representations of sample Entities A and B are aggregated to obtain the relationship vector representation v2 at the semantic space level. relation .
[0097] For each sentence, the relationship vector representation v1 at the sentence level relation and the relationship vector representation v2 at the semantic space level relation can be combined to obtain the relationship vector representation of the sentence.
[0098] The vector fusion layer is used to calculate the attention weight of each sentence and perform weighted summation on the relationship vector representations of each sentence according to the attention weight of the sentence to obtain the weighted relationship vector representation.
[0099] Figure 4 In [ ], it is defined that the sentence set contains m sentences, then the attention weight a = [a1, a2,..., a m , that is, there is a corresponding attention weight for each sentence. By performing weighted summation on the relationship vector representation of the sentence according to the attention weight, the weighted relationship vector representation can be obtained.
[0100] The output layer is used to predict the relationship type corresponding to the sentence set based on the weighted relationship vector representation.
[0101] Specifically, the output layer can use a softmax classifier to judge the relationship type.
[0102] The relation extraction model based on remote supervision provided by the embodiments of the present application can reduce the cost of manually labeled data, and through the sentence-level attention mechanism, highlight the importance of sentences expressing the predefined relation r, and reduce the impact of noisy sentences in the sentence set on model training. In addition, by simultaneously inputting sentences and syntactic tree representations of sentences at the input layer, the relation vector representation of sentences can be calculated from the sentence level and the semantic space level, improving the correctness of the relation vector representation, and further contributing to improving the accuracy of the final relation classification result.
[0103] In some embodiments of the present application, the training process of the multi-modal unsupervised pre-training model used in the foregoing step S120 is described.
[0104] Medical ultrasound data has multiple modalities, such as text modality, images, videos and other multi-modal data, which are also important media for knowledge storage. In order to make good use of data of different modalities, a multi-modal information unified representation technology based on mask prediction is proposed in this embodiment. By modeling the mutual correlation between modalities, a shared embedding space that can represent multi-modal features simultaneously is learned, realizing the association of semantic spaces such as text and images, and further improving the performance of the model in downstream tasks.
[0105] In this embodiment, the initial model is trained unsupervised using a mask prediction task to obtain a trained multi-modal model, where the mask prediction task includes at least two of the following:
[0106] Predicting the masked text, predicting whether the image region corresponding to the text is masked, predicting whether the text corresponding to the image region is masked, predicting whether the text and the image match.
[0107] Among them, mask generation can be achieved by randomly selecting partial data segments of modalities and replacing them with special tokens (such as [MASK]). For images, some regions in the image can be occluded, and for text, some words or sentences can be occluded.
[0108] Among them, the initial model can be a model with multi-modal data processing capabilities.
[0109] The multi-modal model trained by adopting the above mask prediction task has the ability to uniformly represent multi-modal information.
[0110] Furthermore, in order to achieve alignment between cross-modal features, on the basis of the above mask prediction task, a contrastive learning method can be further adopted to train the multi-modal model to train the cross-modal semantic alignment ability of the multi-modal model, and obtain a trained multi-modal unsupervised pre-training model.
[0111] Combined Figure 5 , which exemplifies a process of aligning text and image modality features.
[0112] Select positive example training samples and negative example training samples. Among them, the text and the image in the positive example training samples are associated. For example, the text and the image in the same ultrasound examination report are used as positive example training samples. The text and the image in the negative example training samples are not related.
[0113] The multimodal model includes an input layer and an encoder layer. The encoder layer includes encoding modules for different modalities. Figure 5 For example, for text, the BERT network can be used for encoding. For images, feature projection can be performed through a linear projection layer (LinearProjection of Flattened Patches), and then encoding can be performed through a Transformer Encoder.
[0114] In the training stage, with the goal of narrowing the distance between the encoded features (vector representations) of different modalities in the positive example training samples and widening the distance between the features of different modalities in the negative example training samples, a multimodal unsupervised pre-training model with cross-modal semantic alignment ability is obtained through training.
[0115] Further optionally, the multimodal unsupervised pre-training model may further include a position encoding module for modeling the spatial position of the input image data. By adding position encoding to the image data, position features and layout features can be extracted, so as to better represent the anatomical structure and its associated diagnostic information. Among them, the position features can clarify the spatial position (such as two-dimensional coordinates) of each pixel or partition in the image, helping the model distinguish different anatomical regions. The layout features capture the overall anatomical layout by modeling the spatial distribution relationship (such as adjacency relationship, shape features) between anatomical structures.
[0116] In some embodiments of the present application, the process of fusing the text features and other modal features to obtain the fused features in the foregoing step S130 is described.
[0117] An optional feature fusion strategy, for example, can directly splice the text features and other modal features, and use the spliced features as the fused features.
[0118] The above feature fusion strategy can quickly fuse multimodal features, but there is a lack of information exchange and complementarity between different modal features. For this reason, the present application provides another feature fusion strategy:
[0119] First, splice the text features and other modal features to obtain spliced features.
[0120] Further, the concatenated features are fed into the Transformer model to achieve cross-modal feature fusion by leveraging the attention mechanism of the Transformer, and the hidden layer features extracted by the Transformer decoder are obtained as the fused features.
[0121] In this embodiment, through the attention mechanism of the Transformer model, cross-modal feature fusion processing can be achieved, enabling information exchange and complementarity between features of different modalities in the concatenated features and enhancing the expression ability of the fused features.
[0122] In a possible implementation, the decoder of the above Transformer model can adopt an autoregressive tree decoder (Tree Decoder, TD).
[0123] The existing autoregressive tree decoder is applied to scenarios of parsing mathematical expressions or characters. This application innovatively applies this decoder to the cross-modal feature fusion process to optimize information complementarity between different modalities.
[0124] The autoregressive tree decoder includes four modules: Parent Module, Child Module, Memory Module, and Rleation Module. It synchronously realizes the self-supervised pre-training objective of the mask prediction task during information fusion and transmission. Among them, Parent State represents the state of the parent node, which contains the information passed from the encoder and the state of the parent node itself. After passing through the attention mechanism, part of it is passed to the ChildState of the Child Module according to different task objectives, and part of it forms the Memory Module for storing and managing the intermediate states and information generated during the decoding process. The Child State passes through the attention mechanism to perform mask prediction for the downstream task and ensures relevance to the information stored in the Memory Module at the current level through the Relation Module. The detailed structure of the autoregressive tree decoder can refer to the relevant technical description and will not be elaborated here.
[0125] In the multi-modal feature fusion technology provided in this embodiment, features of different modalities are combined using the splicing method, and the Transformer architecture is further used for cross-modal feature fusion. Essentially, the spliced features are not a whole and do not have the semantic continuity of context. Therefore, simply using the attention mechanism of the Transformer architecture cannot well obtain the hidden layer representations of different modality features from a higher dimension. Therefore, combining the nature of the attention mechanism and the structural characteristics of the current input, an autoregressive tree-like decoder is introduced. By constructing four modules with hierarchical structures and logical orders, different dimensional information of the data can be captured, and at the same time, the transmission and fusion between different modality information can be satisfied, improving the expression ability of the final fused features.
[0126] In some embodiments of the present application, the process of adding the target entity, the relationship between target entity pairs, and the fused features associated with the target entity to the knowledge base in the field of medical ultrasound is described for the foregoing step S140.
[0127] The multi-modal medical ultrasound data obtained in the present application can be multi-source, which may lead to problems such as knowledge duplication and unclear associations between knowledge. The same entity may have different statements in different data sources. Therefore, entity linking is a very important step in knowledge fusion. Entity linking is the process of determining whether the entities in multi-source heterogeneous data point to the same object in the real world.
[0128] The target entity in the text can be identified through entity extraction technology. This target entity may be ambiguous or even unknown. In the face of ambiguity, entity linking technology is needed to link the entity to the unique entity in the objective world. Through entity linking technology, on the one hand, the ambiguity of knowledge elements in concept can be eliminated, redundant and incorrect knowledge elements can be removed to ensure the quality of knowledge base construction; on the other hand, new entities can also be actively discovered to maintain the timeliness and completeness of the knowledge base.
[0129] Combined with Figure 6 , it exemplifies a process of entity linking.
[0130] For the target entity extracted from medical text data, the confidence of entity linking between the target entity and each entity object in the medical ultrasound field knowledge base is calculated. If there is a candidate entity object with a confidence greater than the threshold, the target entity is linked to the candidate entity object as a synonym of the candidate entity object.
[0131] Among them, the construction process of the knowledge base in the field of medical ultrasound can also be understood as an update process, that is, the process of adding the currently newly extracted target entities, the relationships between target entities, and the fusion features associated with target entities to the knowledge base. When involving entity linking during the process of storing target entities in the database, it is necessary to calculate the confidence between the target entity and each existing entity object in the knowledge base, and select the storage method of the target entity according to the confidence level.
[0132] Combined with Figure 6 As shown, if there is a candidate entity object in the knowledge base whose confidence with the target entity is greater than the threshold, the target entity can be used as a synonym of the candidate entity object and linked to the candidate entity object.
[0133] If there is no candidate entity in the knowledge base whose confidence with the target entity is greater than the threshold, experts can be arranged to verify whether the target entity is legal. If it is not legal, it will be directly discarded. If it is legal, the target entity can be newly added to the database, that is, added to the knowledge base as a new entity object.
[0134] After storing the target entity in the database, the relationships between target entities can be further updated to the knowledge base, and the fusion features associated with target entities can be added to the knowledge base.
[0135] This embodiment introduces an optional confidence calculation process for entity linking of target entities and each entity object in the knowledge base in the field of medical ultrasound.
[0136] Combined with Figure 7 As shown, Figure 7 It exemplifies the processing process of an entity linking prediction model.
[0137] The inputs of the input layer include: the target entity, the syntactic tree representation of the medical text data where the target entity is located, and the entity object to be linked in the knowledge base.
[0138] The encoding layer encodes the target entity, the entity object to be linked, and the syntactic tree representation of the medical text data where the target entity is located respectively, to obtain the target entity feature representation, the entity object feature representation to be linked, and the syntactic tree feature representation.
[0139] Among them, for the encoding process of the target entity and the entity object to be linked, a text encoder can be used, such as the BERT model. For the encoding process of the syntactic tree representation of the medical text data where the target entity is located, a graph convolutional network GCN can be used.
[0140] The attention layer performs cross-attention calculation on the entity object feature representation to be linked and the target entity feature representation to obtain two sub-attention context vectors ( Figure 7(defined as sub-attention context vector 1 and sub-attention context vector 2 in [reference]), and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the first differential attention context vector.
[0141] Similarly, through the attention layer, cross-attention calculation is performed on the feature representation of the entity object to be linked and the syntactic tree feature representation, and two sub-attention context vectors are obtained ( Figure 7 defined as sub-attention context vector 3 and sub-attention context vector 4 in [reference]), and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the second differential attention context vector.
[0142] In this embodiment, a combination strategy of cross-attention mechanism and difference operator is proposed, which can also be called differential attention mechanism. Through the cross-attention mechanism, information interaction and fusion are realized, and the difference operator is introduced to eliminate attention noise, further enhancing the ability of context modeling between entities and improving the accuracy of the calculation result of entity link confidence.
[0143] Based on the first differential attention context vector and the second differential attention context vector, the output layer determines the confidence of the entity link between the target entity and the entity object to be linked (which can also be called the similarity between the target entity and the entity object to be linked).
[0144] Since the first differential attention context vector and the second differential attention context vector measure the similarity between the target entity and the entity object to be linked from two dimensions of the target entity and the syntactic tree representation of the medical text data where the target entity is located respectively, the first differential attention context vector and the second differential attention context vector can be weighted and averaged, and based on the weighted average vector, the similarity between the target entity and the entity object to be linked is determined as the confidence of their entity link.
[0145] To sum up:
[0146] In the overall solution process of this application, a multi-perspective semi-supervised entity extraction method and a remote-supervised relation extraction technology are adopted to achieve the purpose of reducing the manual annotation cost, improve the ability of large-scale data processing, a self-supervised multi-modal information representation technology based on mask prediction is proposed, and a multi-task learning strategy is used to realize the unified representation of the entity semantic space. Moreover, an attention mechanism with a difference operator is introduced into the entity link prediction model to achieve highly accurate knowledge links, ensuring the alignment of multi-modal information as a whole.
[0147] In the entity extraction process, a multi-view semi-supervised training model is adopted. Through cross-view training, accurate deep information extraction of multi-source data is achieved. And through the architecture of the shared encoder, the entity extraction task is completed during the entity recognition process, which also demonstrates the idea of transfer learning from the side.
[0148] In the relation extraction process, a remotely supervised relation extraction model is adopted. It is assumed that if there is a certain relationship between two entities in the knowledge base, then the text instances containing these two entities also express the same relationship, so as to learn complex relationship patterns from the original text without explicit feature engineering, thus greatly reducing the need for manual annotation.
[0149] A multi-modal unsupervised pre-training model based on masked prediction is proposed. By designing specific masking tasks, the internal connections between different modal data (such as text, images, etc.) are learned, so as to construct a general feature space that can capture cross-modal information. With the multi-task joint learning method and the autoregressive tree-shaped decoder design, the problem of integrating and aligning data from different modalities is hierarchically solved to ensure that they have similar semantic meanings in the common space.
[0150] Based on the results of entity extraction and relation extraction, by constructing an entity link prediction model that includes cross-attention and differential operators, while eliminating knowledge ambiguity, the dynamic update of the knowledge base is ensured, meeting the requirements of incremental knowledge base construction.
[0151] Next, the medical ultrasound domain knowledge base construction device provided by the embodiments of the present application will be described. The medical ultrasound domain knowledge base construction device described below can be mutually referred to the medical ultrasound domain knowledge base construction method described above.
[0152] See Figure 8 , Figure 8 which is a schematic structural diagram of a medical ultrasound domain knowledge base construction device disclosed in the embodiments of the present application.
[0153] As Figure 8 shown, the device may include:
[0154] A data acquisition unit 11, configured to acquire multi-modal medical ultrasound data, where the multi-modal medical ultrasound data includes medical text data and its associated other modal data;
[0155] An entity and relation extraction unit 12, configured to perform entity extraction on the medical text data to obtain target entities, and extract the relations between target entity pairs in the medical text data;
[0156] The multimodal feature extraction unit 13 is configured to call a multimodal unsupervised pre-training model to extract features from the medical text data and its associated other modal data, obtaining text features and other modal features. The multimodal unsupervised pre-training model is configured to extract features of multimodal data and perform modal alignment;
[0157] The multimodal feature fusion unit 14 is configured to fuse the text features and other modal features to obtain fused features, and establish an association relationship between the fused features and the target entities included in the medical text data;
[0158] The knowledge base update unit 15 is configured to add the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities to the medical ultrasound domain knowledge base.
[0159] In a possible implementation, the device of the present application may further include:
[0160] The multimodal model training unit is configured to train the multimodal unsupervised pre-training model. This training process includes:
[0161] Performing unsupervised training on the initial model using a masked prediction task to obtain a trained multimodal model, where the masked prediction task includes at least two of the following:
[0162] Predicting masked text, predicting whether the image region corresponding to the text is masked, predicting whether the text corresponding to the image region is masked, predicting whether the text and the image match;
[0163] Training the multimodal model using a contrastive learning method to train the cross-modal semantic alignment ability of the multimodal model, obtaining a trained multimodal unsupervised pre-training model.
[0164] In a possible implementation, the process by which the multimodal feature fusion unit fuses the text features and other modal features to obtain fused features includes:
[0165] Concatenating the text features and other modal features to obtain concatenated features;
[0166] Feeding the concatenated features into a Transformer model to achieve cross-modal feature fusion using the attention mechanism of the Transformer, obtaining the hidden layer features extracted by the Transformer decoder as the fused features.
[0167] In a possible implementation, the decoder of the Transformer model uses an autoregressive tree-shaped decoder.
[0168] In a possible implementation, the process of the entity and relationship extraction unit extracting target entities from the medical text data includes:
[0169] Using an entity extraction model to extract target entities from the medical text data, where the entity extraction model is trained using a multi-perspective semi-supervised learning method.
[0170] In a possible implementation, the process of the entity and relationship extraction unit extracting the relationships between the target entity pairs in the medical text data includes:
[0171] Feeding the medical text data containing the target entity pairs and the syntactic tree representation of the medical text data into a remote supervision relationship extraction model to obtain the relationships between the target entity pairs output by the model;
[0172] Among them, the training process of the remote supervision relationship extraction model includes:
[0173] For sample entities A and B, extract the sentences that contain both sample entities A and B from the medical text data to form a sentence set, and identify the relationship r existing between sample entities A and B from an external knowledge base, and use the relationship r as the relationship type label of the sentence set;
[0174] Feeding the sentence set and the syntactic tree representation of each sentence therein into a remote supervision relationship extraction model, calculating the relationship vector representation of each sentence and the attention weight of each sentence, and performing weighted summation on the relationship vector representations of each sentence according to the attention weight of the sentence to obtain a weighted relationship vector representation, and predicting the relationship category corresponding to the sentence set based on the weighted relationship vector representation;
[0175] Calculating the value of the loss function according to the predicted relationship type of the sentence set and the relationship type label of the sentence set, and updating the model parameters according to the value of the loss function until the training end condition is reached.
[0176] In a possible implementation, the remote supervision relationship extraction model includes:
[0177] An input layer for inputting the sentence set and the syntactic tree representation of each sentence therein;
[0178] An encoding layer, including a text encoding module and a graph neural network module, where the text encoding module is used to encode each sentence to obtain the sentence-level vector representations of sample entities A and B, and the graph neural network module is used to encode the syntactic tree representation of each sentence to obtain the semantic space-level vector representations of sample entities A and B;
[0179] A relationship representation layer, which is used to calculate the relationship vector representation of each sentence based on the sentence-level vector representation and the semantic space-level vector representation of sample entities A and B;
[0180] A vector fusion layer, which is used to calculate the attention weight of each sentence, and weighted-sum the relationship vector representations of each sentence according to the attention weight of the sentence to obtain a weighted relationship vector representation;
[0181] An output layer, which is used to predict the relationship type corresponding to the sentence set based on the weighted relationship vector representation.
[0182] In a possible implementation, the process of the knowledge base update unit adding the target entity to the medical ultrasound domain knowledge base includes:
[0183] Calculate the confidence of entity linking for the target entity and each entity object in the medical ultrasound domain knowledge base. If there is a candidate entity object with a confidence greater than the threshold, link the target entity as a synonym of the candidate entity object to the candidate entity object.
[0184] In a possible implementation, the process of the knowledge base update unit calculating the confidence of entity linking for the target entity and each entity object in the medical ultrasound domain knowledge base includes:
[0185] Encode the target entity, the entity object to be linked, and the syntactic tree representation of the medical text data where the target entity is located respectively to obtain a target entity feature representation, an entity object feature representation to be linked, and a syntactic tree feature representation;
[0186] Perform cross-attention calculation on the entity object feature representation to be linked and the target entity feature representation to obtain two sub-attention context vectors, and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the first differential attention context vector;
[0187] Perform cross-attention calculation on the entity object feature representation to be linked and the syntactic tree feature representation to obtain two sub-attention context vectors, and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the second differential attention context vector;
[0188] Determine the confidence of entity linking between the target entity and the entity object to be linked based on the first differential attention context vector and the second differential attention context vector.
[0189] In a possible implementation, the process of the knowledge base update unit determining the confidence of entity linking between the target entity and the entity object to be linked based on the first differential attention context vector and the second differential attention context vector includes:
[0190] Perform a weighted average on the first differential attention context vector and the second differential attention context vector, and based on the vector after the weighted average, determine the confidence level of entity linking between the target entity and the entity object to be linked.
[0191] An embodiment of the present application also provides an electronic device. Refer to Figure 9 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, and the like. Figure 9 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.
[0192] As Figure 9 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603, so as to implement the method for constructing a knowledge base in the field of medical ultrasound in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0193] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0194] An embodiment of the present application also provides a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the methods for constructing a knowledge base in the field of medical ultrasound provided by the embodiment of the present application.
[0195] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the medical ultrasound field knowledge base construction methods provided in the embodiments of the present application.
[0196] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0197] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0199] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are all or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0200] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
Claims
1. A method for constructing a knowledge base in the field of medical ultrasound, characterized in that, Including: Obtain multi-modal medical ultrasound data, where the multi-modal medical ultrasound data includes medical text data and its associated other modal data; Perform entity extraction on the medical text data to obtain target entities, and extract the relationships between target entity pairs in the medical text data; Call a multi-modal unsupervised pre-training model to extract features from the medical text data and its associated other modal data to obtain text features and other modal features. The multi-modal unsupervised pre-training model is configured to extract features of multi-modal data and perform modal alignment; Fuse the text features and other modal features to obtain fused features, and establish an association relationship between the fused features and the target entities included in the medical text data; Add the target entities, the relationships between the target entity pairs, and the fused features associated with the target entities to the knowledge base in the field of medical ultrasound.
2. The method according to claim 1, wherein The training process of the multi-modal unsupervised pre-training model includes: Perform unsupervised training on the initial model using a masked prediction task to obtain a trained multi-modal model. Among them, the masked prediction task includes at least two of the following: Predict the masked text, predict whether the image region corresponding to the text is masked, predict whether the text corresponding to the image region is masked, and predict whether the text and the image match; Use the contrastive learning method to train the multi-modal model to train the cross-modal semantic alignment ability of the multi-modal model, and obtain a trained multi-modal unsupervised pre-training model.
3. The method according to claim 1, wherein The process of fusing the text features and other modal features to obtain fused features includes: Concatenate the text features and other modal features to obtain a concatenated feature; Send the concatenated feature into a Transformer model to use the attention mechanism of the Transformer to achieve cross-modal feature fusion, and obtain the hidden layer features extracted by the Transformer decoder as the fused features.
4. The method according to claim 3, wherein The decoder of the Transformer model uses an autoregressive tree-shaped decoder.
5. The method according to claim 1, wherein The process of performing entity extraction on the medical text data to obtain target entities includes: Use an entity extraction model to perform entity extraction on the medical text data to obtain target entities. The entity extraction model is trained using a multi-perspective semi-supervised learning method.
6. The method according to claim 1, wherein The process of extracting the relationships between target entity pairs in the medical text data includes: Send the medical text data containing target entity pairs and the syntactic tree representation of the medical text data into a distant supervision relation extraction model to obtain the relationships between target entity pairs output by the model; Among them, the training process of the distant supervision relation extraction model includes: For sample entities A and B, extract the sentences that contain both sample entities A and B from the medical text data to form a sentence set, and identify the relationship r existing between sample entities A and B from an external knowledge base. Use the relationship r as the relationship type label of the sentence set; Send the set of sentences and the syntactic tree representations of each sentence in the set into a distant supervision relation extraction model, calculate the relation vector representation of each sentence and the attention weight of each sentence, perform weighted summation on the relation vector representations of each sentence according to the attention weight of the sentence to obtain a weighted relation vector representation, and predict the relation category corresponding to the set of sentences based on the weighted relation vector representation; Calculate the value of the loss function according to the predicted relation type of the set of sentences and the relation type label of the set of sentences, and update the model parameters according to the value of the loss function until the training end condition is reached.
7. The method according to claim 6, wherein The distant supervision relation extraction model includes: An input layer for inputting the set of sentences and the syntactic tree representations of each sentence in the set; An encoding layer including a text encoding module and a graph neural network module. The text encoding module is used to encode each sentence to obtain the sentence-level vector representations of sample entities A and B, and the graph neural network module is used to encode the syntactic tree representations of each sentence to obtain the semantic space-level vector representations of sample entities A and B; A relation representation layer for calculating the relation vector representation of each sentence based on the sentence-level vector representations and semantic space-level vector representations of sample entities A and B; A vector fusion layer for calculating the attention weight of each sentence and performing weighted summation on the relation vector representations of each sentence according to the attention weight of the sentence to obtain a weighted relation vector representation; An output layer for predicting the relation type corresponding to the set of sentences based on the weighted relation vector representation.
8. The method according to claim 1, wherein The process of adding the target entity to the medical ultrasound domain knowledge base includes: Calculate the confidence of entity linking for the target entity and each entity object in the medical ultrasound domain knowledge base. If there is a candidate entity object with a confidence greater than the threshold, link the target entity as a synonym of the candidate entity object to the candidate entity object.
9. The method according to claim 8, characterized in that The process of calculating the confidence of entity linking for the target entity and each entity object in the medical ultrasound domain knowledge base includes: Encode the target entity, the entity object to be linked, and the syntactic tree representation of the medical text data where the target entity is located respectively to obtain the target entity feature representation, the entity object to be linked feature representation, and the syntactic tree feature representation; Perform cross-attention calculation on the entity object to be linked feature representation and the target entity feature representation to obtain two sub-attention context vectors, and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the first difference attention context vector; Perform cross-attention calculation on the entity object to be linked feature representation and the syntactic tree feature representation to obtain two sub-attention context vectors, and calculate the difference between the two sub-attention context vectors through a difference operator, and the result is used as the second difference attention context vector; Determine the confidence of entity linking between the target entity and the entity object to be linked based on the first difference attention context vector and the second difference attention context vector.
10. The method according to claim 9, characterized in that, The process of determining the confidence of entity linking between the target entity and the entity object to be linked based on the first differential attention context vector and the second differential attention context vector includes: Performing weighted averaging on the first differential attention context vector and the second differential attention context vector, and determining the confidence of entity linking between the target entity and the entity object to be linked based on the vector after weighted averaging.
11. A knowledge base construction device in the field of medical ultrasound, characterized in that Including: A data acquisition unit, configured to acquire multimodal medical ultrasound data, where the multimodal medical ultrasound data includes medical text data and its associated other modal data; An entity and relationship extraction unit, configured to perform entity extraction on the medical text data to obtain a target entity, and extract the relationships between target entity pairs in the medical text data; A multimodal feature extraction unit, configured to call a multimodal unsupervised pre-training model to extract features from the medical text data and its associated other modal data, to obtain text features and other modal features, where the multimodal unsupervised pre-training model is configured to extract features of multimodal data and perform modal alignment; A multimodal feature fusion unit, configured to fuse the text features and other modal features to obtain fused features, and establish an association relationship between the fused features and the target entity included in the medical text data; A knowledge base update unit, configured to add the target entity, the relationships between the target entity pairs, and the fused features associated with the target entity to a medical ultrasound domain knowledge base.
12. An electronic device, characterized in that, Including: A memory and a processor; The memory is configured to store a program; The processor is configured to execute the program to implement each step of the method for constructing a medical ultrasound domain knowledge base according to any one of claims 1 to 10.
13. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements each step of the method for constructing a medical ultrasound domain knowledge base according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that, When this computer program is executed by a processor, it implements each step of the method for constructing a medical ultrasound domain knowledge base according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video understanding method and device
CN116994171A
Industry knowledge base system and method based on entity link and relation extraction
CN117151220A
Knowledge graph question and answer method based on knowledge enhancement
CN117407541A
Power equipment operation inspection cognitive large model training method and system
CN117612189A
Ultrasonic multi-mode pre-training method and device, computer equipment and storage medium
CN119886266A
Cited By
PDA (Personal Digital Assistant) key area enhanced display method based on deep learning
CN120807511A