Medical ancient book knowledge graph construction method and system
By processing text and image of ancient medical books, and using multimodal large language model to fusion features, extracting entity relationship triplets, the problem of difficulty in extracting entity relationships in ancient medical books is solved, and the efficient construction and accuracy of the knowledge graph are achieved.
Patent Information
- Application Number
- CN202510586685.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The coexistence of ancient medical books on pictures and texts, obscure language, irregular grammar, and mutation of words and other problems have led to difficulties in extracting entity relationships and the inability to effectively build a knowledge graph.
By performing text and image processing on ancient medical books, text features and image features are extracted, and these features are fused using a pre-trained multimodal large language model to extract entity relationship triplets, and finally a knowledge graph is constructed.
It improves the accuracy and comprehensiveness of the extraction of entities and their relationships, enhances the accuracy and construction efficiency of the knowledge graph, can better understand the content of ancient books and realizes the accurate identification and extraction of entity relationships.
Smart Images

Figure CN120106201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for constructing a knowledge graph of ancient medical books. Background Art
[0002] Ancient medical books are an important part of cultural heritage. With the development of science and technology, knowledge graphs are gradually applied to the field of ancient medical books, such as semantic search, intelligent question and answer, and decision support. However, in the process of constructing knowledge graphs, the coexistence of pictures and texts in ancient medical books, as well as the common problems of obscure language, irregular grammar, word variation, etc., bring great challenges to the extraction of entity relationships from ancient medical books, which in turn leads to the inability to establish attributes, entities, relationships and other characteristics in the knowledge graph. Therefore, there is an urgent need for a method to efficiently and accurately construct knowledge graphs for ancient medical books. Summary of the invention
[0003] In view of this, the purpose of the present invention is to provide a method and system for constructing a knowledge graph of ancient medical books, which can effectively improve the accuracy and comprehensiveness of the extraction of entities and their relationships, better construct a knowledge graph, and improve the accuracy and construction efficiency of the knowledge graph.
[0004] In a first aspect, an embodiment of the present invention provides a method for constructing a knowledge graph of ancient medical books, the method comprising: Perform text processing on ancient medical books to obtain the text features of ancient books; Performing image processing on the ancient medical books to obtain image features of the ancient books; Extracting entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained by performing entity recognition and relationship extraction on the ancient book text features and the ancient book image features through the multimodal large language model; A knowledge graph is constructed based on the entity relationship triples.
[0005] In some embodiments, the text processing of ancient medical books to obtain ancient book text features includes: Collect ancient texts from ancient medical texts; The ancient book text is processed to obtain ancient book text features; wherein the text processing includes: denoising, word segmentation and part-of-speech tagging.
[0006] In some embodiments, the performing image processing on the ancient medical book to obtain the ancient book image features includes: Obtaining an ancient book image corresponding to the ancient medical book; Performing image recognition on the ancient book image to obtain image elements; Extracting element features of the image elements in the ancient book image; Classifying the ancient book image according to the element features to obtain an image category; The element features and the image category are determined as ancient book image features.
[0007] In some implementations, extracting entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model includes: Based on a pre-trained embedding matrix of traditional Chinese medicine terms, text embedding is performed on the sequence of ancient book text features to obtain a text feature vector; Performing image embedding on the sequence of ancient book image features to obtain an image feature vector; fusing the text feature vector and the image feature vector to obtain a multimodal feature vector; Entities and relations are extracted based on the multimodal feature vector to obtain entity-relationship triples.
[0008] In some implementations, extracting entities and relationships based on the multimodal feature vector to obtain entity-relationship triples includes: Extracting context information from the multimodal feature vector based on named entity recognition technology and a preset entity set; Determine entities in the text and / or diagram of the ancient medical book and the locations and types of the entities according to the context information; The relationship between every two entities in the text and / or diagram of the ancient medical book is extracted according to a preset relationship set to obtain an entity relationship triple.
[0009] In some embodiments, the method further comprises: Using a graph database, the entities in the knowledge graph are stored as nodes, and the relationships in the knowledge graph are stored as edges; The node corresponds to a node label, and the node label includes: entity type, entity attribute and node connection relationship.
[0010] In some embodiments, the method further comprises: Obtaining a target path consisting of a plurality of entities having connection relationships in the knowledge graph; Determine the occurrence probability of a first relationship under a first entity according to the entity-relationship triple in the knowledge graph; wherein the first entity and the first relationship correspond to the same entity-relationship triple, and the first entity is any entity among the multiple entities constituting the target path; The credibility of the target path is determined according to the corresponding occurrence probabilities of multiple entities on the target path.
[0011] In some embodiments, the method further comprises: Get the entity to be queried input by the user; Searching the knowledge graph for multiple candidate entities and multiple candidate relationships associated with the entity to be queried; Determining an importance metric value of each of the candidate entities and a weight of each of the candidate relationships; Determining a recommendation score for each candidate entity according to the importance measurement value of each candidate entity and the weight of each candidate relationship; A target entity corresponding to the entity to be queried is recommended from among the multiple candidate entities according to the recommendation score.
[0012] In some implementations, the training process of the multimodal large language model includes: Using the sample text sequence and a preset first loss function to train the multimodal large language model to obtain a first loss function value; Using sample image features and a preset second loss function to train the multimodal large language model to obtain a second loss function value; Training the multimodal large language model using the sample text sequence, the sample image features and a preset third loss function to obtain a third loss function value; Determine a target loss function value according to preset hyperparameters and the first loss function value, the second loss function value, and the third loss function value; The multimodal large language model is trained according to the target loss function value until the target loss function value converges to a preset value, and the training ends.
[0013] In a second aspect, an embodiment of the present invention provides a system for constructing a knowledge graph of ancient medical books, the system comprising the following modules: Text processing module: used to process the text of ancient medical books and obtain the text features of ancient books; Image processing module: used for performing image processing on the ancient medical books to obtain image features of the ancient books; Triple extraction module: used to extract entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained after the multimodal large language model performs entity recognition and relationship extraction on the ancient book text features and the ancient book image features; Knowledge graph construction module: used to construct a knowledge graph based on the entity relationship triples.
[0014] In a third aspect, an embodiment of the invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein when the processor executes the computer program, the steps of the method for constructing a knowledge graph of ancient medical books mentioned in the first aspect are implemented.
[0015] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for constructing a knowledge graph of ancient medical books mentioned in the first aspect are implemented.
[0016] The embodiments of the present invention bring at least the following beneficial effects: The present invention provides a method and system for constructing a knowledge graph of ancient medical books. The scheme includes: performing text processing on ancient medical books to obtain ancient book text features; performing image processing on ancient medical books to obtain ancient book image features; extracting entity relationship triples based on ancient book text features and ancient book image features through a pre-trained multimodal large language model; and constructing a knowledge graph based on the entity relationship triples.
[0017] The above technical solution provides a data basis for realizing multimodal information fusion by extracting ancient book text features and ancient book image features; then, entity relationship triples are extracted based on ancient book text features and ancient book image features through a multimodal large language model. In this process, the features of text and image are comprehensively considered. Text and image often contain complementary information, which can effectively improve the accuracy and comprehensiveness of entity and relationship extraction. Through multimodal feature fusion, the content of ancient books can be understood more comprehensively, and the entities and their relationships in ancient book texts can be accurately identified and extracted, and the information interaction ability of entity relationship triples in the dimensions of text and image can be improved, which effectively handles the complex problems of text in ancient books; at the same time, the multimodal large language model can ensure the accuracy of relationship extraction and improve the extraction efficiency. On this basis, knowledge graphs can be better constructed based on entity relationship triples, improving the accuracy and construction efficiency of knowledge graphs.
[0018] Other features and advantages of the present invention will be described in the following description, or some features and advantages can be inferred or determined without doubt from the description, or can be learned by implementing the above-mentioned technology of the present invention.
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are specifically listed below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A flowchart of a method for constructing a knowledge graph of ancient medical books provided in an embodiment of the present invention; Figure 2 A flowchart of the intelligent interpretation and knowledge extraction steps in a method for constructing a knowledge graph of ancient medical books provided in an embodiment of the present invention; Figure 3 A flowchart of the reasoning steps of a knowledge graph in a method for constructing a knowledge graph of ancient medical books provided in an embodiment of the present invention; Figure 4 A flowchart of the intelligent recommendation step of a knowledge graph in a method for constructing a knowledge graph of ancient medical books provided in an embodiment of the present invention; Figure 5 A schematic diagram of the structure of a medical ancient book knowledge graph construction system provided by an embodiment of the present invention; Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0022] icon: 510-text processing module; 520-image processing module; 530-triplet extraction module; 540-knowledge graph construction module; 101 - processor; 102 - memory; 103 - bus; 104 - communication interface. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] At present, in the management of coal production, transportation, sales and storage, there is a possibility of human tampering in data collection and reporting, which affects the authenticity of the data and leads to poor accuracy and efficiency of coal supervision.
[0025] In order to efficiently and accurately construct a knowledge graph for ancient medical books, an embodiment of the present invention provides a method and system for constructing a knowledge graph for ancient medical books. The method extracts entity relationship triples based on ancient book text features and ancient book image features through a multimodal large language model, and constructs a knowledge graph based on this. This method can effectively improve the accuracy and comprehensiveness of the extraction of entities and their relationships, better construct a knowledge graph, and improve the accuracy and construction efficiency of the knowledge graph.
[0026] To facilitate understanding of this embodiment, a method for constructing a knowledge graph of ancient medical books disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the following steps may be included.
[0027] Text data collection and processing step S101: Perform text processing on ancient medical books to obtain ancient book text features.
[0028] This embodiment includes: collecting ancient book texts from ancient medical book texts; performing text processing on the ancient book texts to obtain ancient book text features; wherein the text processing includes: denoising, word segmentation, and part-of-speech tagging.
[0029] In a specific embodiment, a technology such as OCR (Optical Character Recognition) can be used to convert the text in the scanned copy or electronic version of the ancient medical book into a text format that can be recognized by a computer to obtain the ancient book text; it is assumed that the ancient book text is represented by T={t 1 ,t 2 ,……,t n}, t n Indicates the nth character in the ancient text.
[0030] The extracted ancient book text T is first denoised, and the denoising function D(t i ) can adopt the statistical noise filtering method (where t i represents any word among n words), for example, calculating the frequency of character occurrence. If the frequency of a character is much lower than the frequency of normal text characters and does not conform to the characteristics of traditional Chinese medicine terms, it is determined to be a noise character and removed. The denoised text can be expressed as T'={D(t 1 ),D(t 2 ),……,D(t n )}.
[0031] Perform word segmentation on the denoised text T', the word segmentation function For example, we can combine dictionary matching with statistical machine learning. Let the dictionary set be Dic, and for the text , if there is a substring s in Dic, s is segmented as a word; at the same time, the statistical model such as Hidden Markov Model (HMM) is used to predict the word segmentation of some words that are not matched in the dictionary set Dic. The text after word segmentation can be expressed as T''={S(D(t 1 )),S(D(t 2 )),……,S(D(t n ))}.
[0032] The text T'' after word segmentation is tagged with part of speech, and the part of speech tagging function P(t'' k ) can use the conditional random field (CRF) model, as shown in the following formula:
[0033] Specifically, let the characteristic function be , where i represents the position of the word in the text, l Represents the part-of-speech tag. The CRF model is trained by a large amount of labeled traditional Chinese medicine text data to obtain the model parameters , then the annotated text can be expressed as T'''={P(S(D(t 1 ))),P(S(D(t 2 ))),……,P(S(D(t n )))}.
[0034] Next, we can use the word vector to represent the annotated text T''' and get the ancient text features. Assume that the word vector of the jth word in the ancient text is , then word t i It can be expressed as ,in For this word t i The weight in ancient texts can be determined by algorithms such as TF-IDF (term frequency-inverse document frequency), that is, ,in is word j in text t i The word frequency in is the number of texts containing word j, and N is the total number of texts.
[0035] Image data acquisition and processing S102: Perform image processing on ancient medical books to obtain image features of the ancient books.
[0036] In this embodiment, the method may include: first, obtaining an ancient book image corresponding to an ancient medical book; and performing image recognition on the ancient book image to obtain image elements.
[0037] In this embodiment, the ancient book image can be represented as I={i 1 ,i 2 ,……,i m}, for image recognition, a deep learning image recognition algorithm (expressed as R( )) can be used, such as a convolutional neural network. For example, the convolutional layer of the convolutional neural network is Conv, the pooling layer is Pool, and the fully connected layer is FC. The process of image recognition of ancient book images can be expressed as R(I)=FC(Pool(Conv(I))); through image recognition, image elements such as illustrations, medicinal material maps and tables in ancient book images can be identified.
[0038] Then, the element features of the image elements in the ancient book image are extracted. Specifically, the image elements in the ancient book image can be extracted by using a preset image feature extraction function (expressed as F( )) through a convolutional neural network. The extracted element features can be expressed as X f =F(R(I)), where X f For example, the feature map of a certain intermediate layer in a convolutional neural network can be used as a representation of the element features of the image.
[0039] Next, the ancient book images are classified according to the element features to obtain image categories; further, the element features and image categories are determined as ancient book image features.
[0040] Specifically, the convolutional neural network can be used to use a preset image classification function (expressed as C( )) to classify the ancient book images according to the element features, so as to classify the ancient book images into different image categories. The image category can be expressed as Xc=C(R(I)).
[0041] The above classification process can be achieved by constructing a multi-classification neural network model, such as using the Softmax function as the output layer activation function, calculating the probability that the image belongs to each category p(y|I)=Softmax(W·R(I)+b), where W is the weight matrix, b is the bias vector, and y is the category label, and taking the category with the highest probability as the image classification result. For example, if the ancient book image is a medicinal material atlas, the elemental features such as the shape and color of the medicinal material can be obtained through feature extraction; the classification function is used to classify the ancient book image according to the elemental features, determine the medicinal material category to which it belongs, and use the medicinal material category as the image category.
[0042] Element features and image categories can be used together as ancient book image features, and ancient book image features and the aforementioned ancient book text features are the basis for jointly constructing multimodal data, so that subsequent multimodal large language models can be deeply integrated and analyzed, thereby providing comprehensive data support for the intelligent interpretation and knowledge extraction of ancient Chinese medicine books.
[0043] Intelligent interpretation and knowledge extraction step S103: extract entity relationship triples based on ancient book text features and ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained after entity recognition and relationship extraction of ancient book text features and ancient book image features through a multimodal large language model.
[0044] Knowledge graph construction step S104: construct a knowledge graph based on entity relationship triples.
[0045] In the above step S103, in order to enable the multimodal large language model to be directly applied to intelligent interpretation and knowledge extraction, the multimodal large language model needs to be trained in advance, and the parameters of the multimodal large language model need to be obtained through training. The purpose of training the multimodal large language model is to finally determine the parameters that can meet the requirements. Based on this, before describing step S103 in this embodiment, a training method for a multimodal large language model is first given, as shown in the following steps (1)-(5).
[0046] (1) Use the sample text sequence and the preset first loss function to train the multimodal large language model to obtain the first loss function value.
[0047] The goal of this embodiment is to predict the next word in a text sequence, that is, to train the word prediction capability of a multimodal large language model using a sample text sequence and a preset first loss function.
[0048] Specifically, the sample text sequence can be represented as X t , the sample text sequence is predicted by the multimodal large language model to be trained, and the predicted probability distribution of the next (i.e., the k+1th) word element output is: P(x t(k+1) |x t1 ,x t2 ,...,x tk )=softmax(head q ·W qm ·head k +head v ·W vm ) Among them, head q ,head k ,head v is the query, key, and value vector in the multi-head attention mechanism, W qm ,W vm is the corresponding weight matrix.
[0049] Exemplarily, the cross entropy loss shown in the following formula is used as the first loss function:
[0050] The first loss function value L is calculated by minimizing the cross entropy loss function LM .
[0051] (2) Use the sample image features and the preset second loss function to train the multimodal large language model to obtain the second loss function value.
[0052] The goal of this embodiment is to generate corresponding text description Y for sample image features through a multimodal large language model t ; That is, using sample image features and a preset second loss function, the ability of the multimodal large language model to generate image descriptions is trained.
[0053] Specifically, the sample image features can be expressed as X i (Note: i here stands for image, image), the multimodal large language model to be trained generates a text description Y based on the sample image features t The probability distribution of is: P(y tj |y t1 ,y t2 ,...,y t(j-1) ,X i )=softmax(head q' ·W qdm ·head k '+head v '·W vdm ) Among them, head q ', head k ', head v ' is the vector in the multi-head attention mechanism for the image description generation task, W qdm ,W vdm is the corresponding weight matrix.
[0054] Exemplarily, the negative log-likelihood loss function shown in the following formula is used as the second loss function, and the second loss function value L is calculated: IC :
[0055] (3) Use the sample text sequence, sample image features and the preset third loss function to train the multimodal large language model to obtain the third loss function value.
[0056] The goal of this embodiment is to achieve cross-modal matching, that is, to use sample text sequences, sample image features and a preset third loss function to train the cross-modal matching judgment ability of the multimodal large language model.
[0057] Specifically, the sample text sequence X is judged by the multimodal large language model to be trained t And sample image features X i Whether it matches. In the judgment process, a binary classifier is used to output the matching probability p. m : p m =sigmoid(MLP([F(X t ,X i )]) Among them, MLP is a multi-layer perceptron.
[0058] Exemplarily, the binary cross entropy shown in the following formula is used as the third loss function, and the third loss function value L is calculated: CM : L CM =-[ymlogp m +(1-y m )log(1-p m )] Among them, y m is the actual matching tag.
[0059] (4) Determine the target loss function value based on the preset hyperparameters and the first loss function value, the second loss function value, and the third loss function value.
[0060] Specifically, during the model training process, the total loss function is determined according to the preset hyperparameters and the first loss function value, the second loss function value, and the third loss function value: L = αL LM +βL IC +γL CM Among them, α, β, γ are hyperparameters for balancing different pre-training tasks, and L is the target loss function value.
[0061] (5) The multimodal large language model is trained according to the target loss function value until the target loss function value converges to a preset value, and the training ends.
[0062] At this point, the training process of the multimodal large language model is completed.
[0063] Based on the above embodiments, this embodiment further describes the above intelligent interpretation and knowledge extraction step S103. Figure 2 As shown, the following steps may be included.
[0064] Step S201, based on the pre-trained TCM term embedding matrix, text embedding is performed on the sequence of ancient book text features to obtain a text feature vector.
[0065] Step S202, performing image embedding on the sequence of ancient book image features to obtain an image feature vector.
[0066] In a specific embodiment, the multimodal large language model can be, for example, an encoder-decoder structure based on a Transformer. In the encoder part, the sequence Yt=[y t1 ,y t2 ,...,y tn ] (where y ti represents the i-th word in the text) and the sequence Y of ancient book image features i =[y i1 ,y i2 ,...,y im ] (where y ij represents the j-th eigenvector of the image, where i represents the image, image).
[0067] In the encoder, the sequence of ancient book text features is firstly embedded in text through a specific input embedding layer, and the sequence of ancient book image features is embedded in image.
[0068] Among them, text embedding can include: Et(y ti )=W t ·y ti +bt. Among them, W t is the text embedding matrix, b t is the bias term, Et(y ti ) is the text feature vector.
[0069] Image embedding can include: Ei(y{ij})=W i ·y ij +bi. Among them, W i is the image embedding matrix, b i is the bias term, and Ei(y{ij}) is the image feature vector.
[0070] In some embodiments, knowledge in the field of traditional Chinese medicine can also be introduced into the embedding layer or middle layer of the multimodal large language model. 1 ,t 2 ,...,t p ], a TCM term embedding matrix W can be pre-trained ct In this case, the sequence of ancient book text features can be embedded in text according to the following formula: E t (y ti )=W t ·y ti +b t +W ct·e(t i ) Among them, e(t i ) is a term in traditional Chinese medicine i A specific embedding representation of (which can be a one-hot encoding or other pre-trained vector representation), E t (y ti ) is a text feature vector. This text embedding method enables the multimodal large language model to better understand the professional terms and concepts in the field of traditional Chinese medicine during the learning process, and improves the accuracy and professionalism of the multimodal large language model in processing ancient Chinese medicine books.
[0071] Step S203: fuse the text feature vector and the image feature vector to obtain a multimodal feature vector.
[0072] In this embodiment, the embedded text feature vector and the image feature vector are fused through a fusion layer. An example fusion method may be: concatenating the text feature vector and the image feature vector, and then performing a linear transformation on the concatenated vector, for example: F(Y t ,Y i )=W f ·[E t (Y t );E i (Y i )]+b f Among them, F(Y t ,Y i ) is the multimodal feature vector, E t (Y t ) is the text feature vector, E i (Y i ) is the image feature vector, W f is the weight matrix of the fusion layer, b f is the bias term, and [;] represents the concatenation operation.
[0073] Another example fusion method can be: using the attention mechanism for cross-modal fusion. In this process, it is assumed that the attention weight of the text feature vector to the image feature vector is , the attention weight of the image feature vector to the text feature vector is , then the fused multimodal feature vector It can be expressed as:
[0074] in, Represents the word t i The text feature vector of represents the jth image feature vector of the ancient book image, , ( a is the attention calculation function, for example, the dot product attention function can be used .
[0075] Step S204: extract entities and relationships based on the multimodal feature vector to obtain entity-relationship triples. This embodiment may include: According to named entity recognition technology and a preset entity set, context information is extracted from a multimodal feature vector; entities and the positions and types of entities in the text and / or diagrams of ancient medical books are determined according to the context information; the relationship between every two entities in the text and / or diagrams of ancient medical books is extracted according to a preset relationship set to obtain entity relationship triples.
[0076] Specifically, for knowledge extraction, assume that the entity set of TCM terms, symptoms, and medicinal materials is represented by E={e 1 ,e 2 ,……,e q}, the relation set is R={r 1 ,r 2 ,……,r s}. The named entity recognition technology is used to determine the location and type of entities in the text and / or diagrams of ancient medical books, and an entity recognition function is constructed. For text, the entity recognition function can be expressed as , for a graph, the entity recognition function can be expressed as .by For example, using the entity recognition function Perform entity recognition. If there are entities in the text and / or graph, k ,but , otherwise 0.
[0077] For relation extraction, this embodiment can be based on rules or deep learning models. Assume that the relation extraction function is , if in the text t i Medium Entity k1 and e k2 There is a relationship l ,but .
[0078] According to the above embodiment, entities in the text and / or diagram of ancient medical books and the relationship between every two entities are extracted, thereby obtaining entity relationship triples.
[0079] Furthermore, for step S104, when constructing a knowledge graph based on entity-relationship triples, the knowledge graph can be constructed with entities as nodes and relationships as edges. The knowledge graph can be expressed as G=(E,R), where the attributes of the nodes can include relevant information of the entity (such as the origin, nature, taste and meridians of the medicinal materials, etc.).
[0080] In the knowledge graph, let the entity set be E and the relationship set be R. For a specific entity (For example, a Chinese herbal medicine entity "ginseng") can be represented by the vector e i To represent its semantic features. Assuming that the dimension of the entity vector is d, then e i =[e i1 ,e i2 ,……,e id ], where e ij Indicates that the entity is j The feature values in the dimension.
[0081] For relationships (such as "having efficacy" relationship), the vector r k Indicates that r k =[r k1 ,r k2 ,……,r kd ].
[0082] When there is a triple (e i ,r k ,e j ) represents entity e i Through the relationship k With entity e j When the two terms are connected (for example, "ginseng has the effect of replenishing vital energy"), in the representation learning of knowledge graphs, models based on energy functions are often used, such as the TransE model, whose basic formula is: , that is, by minimizing the energy function f(e i ,r k ,e j )=||e i +e k -e j || is used to learn vector representations of entities and relations, where ||·|| represents the norm of the vector (such as the L2 norm).
[0083] This embodiment uses a multimodal knowledge graph to expand a single text-based node into a form that includes multimodal data such as digital images and text to present knowledge. The multimodal data is integrated with the corresponding entities, relationships, and attributes of the text to complete the construction of the multimodal knowledge graph.
[0084] This embodiment can construct a knowledge graph to achieve effective storage and management of ancient Chinese medicine knowledge in a structured manner, so as to facilitate efficient retrieval and intelligent recommendation services in the future, and make better development in smart medical fields such as semantic search, knowledge question and answer, and clinical decision support. For example, a query based on a knowledge graph can be expressed as Q(E, R, condition), where condition is a query condition, such as querying for medicinal treatment methods related to a certain disease.
[0085] After constructing the knowledge graph, the method provided in this embodiment may further include storage and query of the knowledge graph, reasoning and intelligent recommendation based on the knowledge graph, etc., refer to the following embodiments.
[0086] In this embodiment, the storage method of the knowledge graph may include: using a graph database to store entities in the knowledge graph as nodes and storing relationships in the knowledge graph as edges; wherein the nodes correspond to node labels, and the node labels include: entity type, entity attributes and node connection relationships.
[0087] Taking the use of graph databases (such as Neo4j) to store knowledge graphs as an example, entities can be stored as nodes and relationships as edges. Specifically, the node set in the graph database is N, and the edge set is L. For a node Corresponding to an entity e i , its storage structure may contain node labels, which include: entity type (such as "medicinal materials", "diseases", etc.), entity attributes (such as entity name, origin, etc.) and node connection relationships with other nodes.
[0088] In this embodiment, in terms of querying the knowledge graph, for example, querying medicinal materials with a certain efficacy, it can be expressed as: MATCH(e 1 : Medicinal Materials)-[r: Has efficacy]->(e 2 :Effect name: "Specific effect")RETURNe 1 The MATCH clause is used to specify the graph pattern of the query, that is, to find nodes e from the type of "herbal medicine" 1 Connected to node e of type "Efficacy" with name "SpecificEfficacy" through relationship r of "hasEfficacy" 2 , the RETURN clause specifies to return the medicinal material node e1 that meets the conditions.
[0089] Reference Figure 3 In this embodiment, the reasoning method of the knowledge graph may include: Step S301: Obtain a target path consisting of multiple entities with connection relationships in the knowledge graph.
[0090] In terms of reasoning of the knowledge graph, this embodiment can use path reasoning in the knowledge graph. Assume that P(e i ,e j ) indicates that from entity e i To entity e j A target path consists of a series of entities connected by relationships, such as P(e i ,e j )=(e i ,r 1 ,e i+1 ,r 2 ,……,e j ).
[0091] Step S302: Determine the occurrence probability of the first relationship under the first entity according to the entity relationship triple in the knowledge graph; wherein the first entity and the first relationship correspond to the same entity relationship triple, and the first entity is any entity among the multiple entities constituting the target path.
[0092] For example, by counting the number of related entity relationship triples in the knowledge graph, the first entity e m The first relation r k The probability of occurrence can be expressed as p(r k |e m ).
[0093] Step S303: Determine the credibility of the target path according to the corresponding occurrence probabilities of multiple entities on the target path.
[0094] The credibility of the inference can be calculated by the combination of relations on the target path. Specifically, a probability-based method can be used to determine the credibility of the target path P according to the following formula: p(P)=p(r 1 |e i )×p(r 2 |e i+1 )×…… Among them, p(P) represents the credibility of the target path P.
[0095] Reference Figure 4 In this embodiment, the intelligent recommendation method of the knowledge graph may include: Step S401: Obtain the entity to be queried input by the user.
[0096] Step S402: Search the knowledge graph for multiple candidate entities and multiple candidate relationships associated with the entity to be queried.
[0097] For intelligent recommendation, you can recommend medicinal treatment plans related to a certain disease to users. Get the query entity entered by the user. For example, the query entity entered by the user is the disease entity e s The system first finds the entity e in the knowledge graph that matches the query entity. s All candidate relationships and candidate entities are related.
[0098] Step S403: Determine the importance measurement value of each candidate entity and the weight of each candidate relationship.
[0099] Step S404: Determine the recommendation score of each candidate entity according to the importance measurement value of each candidate entity and the weight of each candidate relationship.
[0100] Among them, the candidate entity e i The recommendation score S(e i ) can be expressed as:
[0101] Among them, w rk It is a relationship k The weight of f(e s ,r k ,e i ) is the correlation calculated based on the entity and relationship representation mentioned above, such as the inverse of the energy function value in the TransE model or other similarity metrics, g(e i ) represents the candidate entity e i Importance metrics (such as normalized values of connectivity, citation frequency, etc.).
[0102] Step S405: recommending a target entity corresponding to the entity to be queried from among multiple candidate entities according to the recommendation scores.
[0103] According to the recommendation score S(e i ) Sort the candidate entities from high to low, and recommend the candidate entity ranked in the first place or the top N places as the target entity to the user.
[0104] To sum up, from the method for constructing the medical ancient book knowledge graph mentioned in the above embodiment, it can be seen that the method includes: performing text processing on the medical ancient books to obtain the ancient book text features; performing image processing on the medical ancient books to obtain the ancient book image features; extracting entity relationship triples based on the ancient book text features and ancient book image features through a pre-trained multimodal large language model; and constructing a knowledge graph based on the entity relationship triples.
[0105] The above technical solution provides a data basis for realizing multimodal information fusion by extracting ancient book text features and ancient book image features; then, entity relationship triples are extracted based on ancient book text features and ancient book image features through a multimodal large language model. In this process, the features of text and image are comprehensively considered. Text and image often contain complementary information, which can effectively improve the accuracy and comprehensiveness of entity and relationship extraction. Through multimodal feature fusion, the content of ancient books can be understood more comprehensively, and the entities and their relationships in ancient book texts can be accurately identified and extracted, and the information interaction ability of entity relationship triples in the dimensions of text and image can be improved, which effectively handles the complex problems of text in ancient books; at the same time, the multimodal large language model can ensure the accuracy of relationship extraction and improve the extraction efficiency. On this basis, knowledge graphs can be better constructed based on entity relationship triples, improving the accuracy and construction efficiency of knowledge graphs.
[0106] In addition, in practical applications, the present invention helps to promote the construction of smart medicine and the inheritance of traditional Chinese medicine culture to a certain extent.
[0107] Corresponding to the above method embodiment, the present invention provides a system for constructing a knowledge graph of ancient medical books, such as Figure 5 As shown, the system includes the following modules: Text processing module 510: used to process the text of ancient medical books to obtain the text features of the ancient books; Image processing module 520: used to perform image processing on the ancient medical book to obtain image features of the ancient book; Triple extraction module 530: used to extract entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained after the multimodal large language model performs entity recognition and relationship extraction on the ancient book text features and the ancient book image features; Knowledge graph construction module 540: used to construct a knowledge graph based on the entity relationship triples.
[0108] In some implementations, the text processing module 510 is used to: Collect ancient texts from ancient medical texts; The ancient book text is processed to obtain ancient book text features; wherein the text processing includes: denoising, word segmentation and part-of-speech tagging.
[0109] In some embodiments, the image processing module 520 is configured to: Obtaining an ancient book image corresponding to the ancient medical book; Performing image recognition on the ancient book image to obtain image elements; Extracting element features of the image elements in the ancient book image; Classifying the ancient book image according to the element features to obtain an image category; The element features and the image category are determined as ancient book image features.
[0110] In some implementations, the triple extraction module 530 is configured to: Based on a pre-trained embedding matrix of traditional Chinese medicine terms, text embedding is performed on the sequence of ancient book text features to obtain a text feature vector; Performing image embedding on the sequence of ancient book image features to obtain an image feature vector; fusing the text feature vector and the image feature vector to obtain a multimodal feature vector; Entities and relations are extracted based on the multimodal feature vector to obtain entity-relationship triples.
[0111] In some implementations, the triple extraction module 530 is configured to: Extracting context information from the multimodal feature vector based on named entity recognition technology and a preset entity set; Determine entities in the text and / or diagram of the ancient medical book and the locations and types of the entities according to the context information; The relationship between every two entities in the text and / or diagram of the ancient medical book is extracted according to a preset relationship set to obtain an entity relationship triple.
[0112] In some embodiments, the ancient medical book knowledge graph construction system further includes a knowledge graph storage module, which is used to: Using a graph database, the entities in the knowledge graph are stored as nodes, and the relationships in the knowledge graph are stored as edges; The node corresponds to a node label, and the node label includes: entity type, entity attribute and node connection relationship.
[0113] In some embodiments, the ancient medical book knowledge graph construction system further includes a path reasoning module, which is used to: Obtaining a target path consisting of a plurality of entities having connection relationships in the knowledge graph; Determine the occurrence probability of a first relationship under a first entity according to the entity-relationship triple in the knowledge graph; wherein the first entity and the first relationship correspond to the same entity-relationship triple, and the first entity is any entity among the multiple entities constituting the target path; The credibility of the target path is determined according to the corresponding occurrence probabilities of multiple entities on the target path.
[0114] In some embodiments, the ancient medical book knowledge graph construction system further includes an intelligent recommendation module, which is used to: Get the entity to be queried input by the user; Searching the knowledge graph for multiple candidate entities and multiple candidate relationships associated with the entity to be queried; Determining an importance metric value of each of the candidate entities and a weight of each of the candidate relationships; Determining a recommendation score for each candidate entity according to the importance measurement value of each candidate entity and the weight of each candidate relationship; A target entity corresponding to the entity to be queried is recommended from among the multiple candidate entities according to the recommendation score.
[0115] In some embodiments, the ancient medical book knowledge graph construction system further includes a model training module, which is used to: Using the sample text sequence and a preset first loss function to train the multimodal large language model to obtain a first loss function value; Using sample image features and a preset second loss function to train the multimodal large language model to obtain a second loss function value; Training the multimodal large language model using the sample text sequence, the sample image features and a preset third loss function to obtain a third loss function value; Determine a target loss function value according to preset hyperparameters and the first loss function value, the second loss function value, and the third loss function value; The multimodal large language model is trained according to the target loss function value until the target loss function value converges to a preset value, and the training ends.
[0116] It can be seen from the medical ancient book knowledge graph construction system mentioned in the above embodiment that the system extracts entity relationship triples based on ancient book text features and ancient book image features through a multimodal large language model. In this process, the features of text and image are comprehensively considered. Text and image often contain complementary information, which can effectively improve the accuracy and comprehensiveness of entity and relationship extraction. Through multimodal feature fusion, the content of ancient books can be understood more comprehensively, and accurate identification and extraction of entities and their relationships in ancient book texts can be achieved. The information interaction ability of entity relationship triples in the dimensions of text and image is improved, and the complex problems of text in ancient books are effectively handled; at the same time, the multimodal large language model can ensure the accuracy of relationship extraction and improve the extraction efficiency. On this basis, knowledge graphs can be better constructed based on entity relationship triples, and the accuracy and construction efficiency of knowledge graphs can be improved.
[0117] The ancient medical book knowledge graph construction system provided in this embodiment has the same technical features as the ancient medical book knowledge graph construction method provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects. For the sake of brief description, for matters not mentioned in the embodiment, reference can be made to the corresponding contents in the above-mentioned ancient medical book knowledge graph construction method embodiment.
[0118] This embodiment also provides an electronic device. The structural diagram of the electronic device is as follows: Figure 6 As shown, the device includes a processor 101 and a memory 102; wherein the memory 102 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the above-mentioned method for constructing the ancient medical book knowledge graph.
[0119] Figure 6 The electronic device shown further includes a bus 103 and a communication interface 104 , and the processor 101 , the communication interface 104 and the memory 102 are connected via the bus 103 .
[0120] The memory 102 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The bus 103 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0121] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and send the encapsulated IPv4 message or IPv4 message to the user terminal through the network interface.
[0122] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 101. The above processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and completes the steps of the method of the above embodiment in combination with its hardware.
[0123] An embodiment of the present invention further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for constructing a knowledge graph of ancient medical books in the aforementioned embodiment are executed.
[0124] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0125] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0127] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that can be executed by a processor. Based on this understanding, the technical solution of the present invention can essentially or in other words, the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0128] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for constructing a knowledge graph of ancient medical books, characterized in that: The method comprises: Perform text processing on ancient medical books to obtain the text features of ancient books; Performing image processing on the ancient medical books to obtain image features of the ancient books; Extracting entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained by performing entity recognition and relationship extraction on the ancient book text features and the ancient book image features through the multimodal large language model; A knowledge graph is constructed based on the entity relationship triples.
2. The method according to claim 1, characterized in that The text processing of ancient medical books to obtain ancient book text features includes: Collect ancient texts from ancient medical texts; The ancient book text is processed to obtain ancient book text features; wherein the text processing includes: denoising, word segmentation and part-of-speech tagging.
3. The method according to claim 1, characterized in that The performing image processing on the ancient medical book to obtain the image features of the ancient book includes: Obtaining an ancient book image corresponding to the ancient medical book; Performing image recognition on the ancient book image to obtain image elements; Extracting element features of the image elements in the ancient book image; Classifying the ancient book image according to the element features to obtain an image category; The element features and the image category are determined as ancient book image features.
4. The method according to claim 1, characterized in that: The extracting entity relationship triples based on the ancient book text features and the ancient book image features by using a pre-trained multimodal large language model includes: Based on a pre-trained embedding matrix of traditional Chinese medicine terms, text embedding is performed on the sequence of ancient book text features to obtain a text feature vector; Performing image embedding on the sequence of ancient book image features to obtain an image feature vector; fusing the text feature vector and the image feature vector to obtain a multimodal feature vector; Entities and relations are extracted based on the multimodal feature vector to obtain entity-relationship triples.
5. The method according to claim 4, characterized in that The extracting of entities and relations based on the multimodal feature vector to obtain entity-relationship triples includes: Extracting context information from the multimodal feature vector according to named entity recognition technology and a preset entity set; Determine entities in the text and / or diagram of the ancient medical book and the locations and types of the entities according to the context information; The relationship between every two entities in the text and / or diagram of the ancient medical book is extracted according to a preset relationship set to obtain an entity relationship triple.
6. The method according to claim 1, characterized in that The method further comprises: Using a graph database, the entities in the knowledge graph are stored as nodes, and the relationships in the knowledge graph are stored as edges; The node corresponds to a node label, and the node label includes: entity type, entity attribute and node connection relationship.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining a target path consisting of a plurality of entities having connection relationships in the knowledge graph; Determine the occurrence probability of a first relationship under a first entity according to the entity-relationship triple in the knowledge graph; wherein the first entity and the first relationship correspond to the same entity-relationship triple, and the first entity is any entity among the multiple entities constituting the target path; The credibility of the target path is determined according to the corresponding occurrence probabilities of multiple entities on the target path.
8. The method according to claim 1, characterized in that The method further comprises: Get the entity to be queried input by the user; Searching the knowledge graph for multiple candidate entities and multiple candidate relationships associated with the entity to be queried; Determining an importance metric value of each of the candidate entities and a weight of each of the candidate relationships; Determining a recommendation score for each candidate entity according to the importance measurement value of each candidate entity and the weight of each candidate relationship; A target entity corresponding to the entity to be queried is recommended from among the multiple candidate entities according to the recommendation score.
9. The method according to claim 1, characterized in that: The training process of the multimodal large language model includes: Using the sample text sequence and a preset first loss function to train the multimodal large language model to obtain a first loss function value; Using sample image features and a preset second loss function to train the multimodal large language model to obtain a second loss function value; Training the multimodal large language model using the sample text sequence, the sample image features and a preset third loss function to obtain a third loss function value; Determine a target loss function value according to preset hyperparameters and the first loss function value, the second loss function value, and the third loss function value; The multimodal large language model is trained according to the target loss function value until the target loss function value converges to a preset value, and the training ends.
10. A system for constructing a knowledge graph of ancient medical books, characterized in that: The system includes the following modules: Text processing module: used to process the text of ancient medical books and obtain the text features of ancient books; Image processing module: used for performing image processing on the ancient medical books to obtain image features of the ancient books; Triple extraction module: used to extract entity relationship triples based on the ancient book text features and the ancient book image features through a pre-trained multimodal large language model; wherein the entity relationship triples are obtained after the multimodal large language model performs entity recognition and relationship extraction on the ancient book text features and the ancient book image features; Knowledge graph construction module: used to construct a knowledge graph based on the entity relationship triples.