Medical knowledge graph management system

By designing a medical knowledge graph management system, the problem of low efficiency in traditional medical knowledge management is solved, efficient knowledge management and personalized retrieval are achieved, and the efficiency of utilization of medical knowledge is improved.

CN120258108APending Publication Date: 2025-07-04GENERAL HOSPITAL OF PLA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510286771.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional medical knowledge management methods are difficult to meet the needs of rapid retrieval and efficient use of medical knowledge, especially when medical knowledge is growing explosively.

Method used

A medical knowledge graph management system is designed, including a data reception module, a graph processing module, a graph database, an information retrieval module, a user behavior database and a behavior acquisition module. Through the collaborative work of these modules, the extraction of medical knowledge triplets, the update and complementation of graphs, and personalized knowledge retrieval functions are provided.

Benefits of technology

It improves the management efficiency and search efficiency of medical knowledge, provides personalized knowledge retrieval services, and can better manage and utilize medical knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258108A_ABST
    Figure CN120258108A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a management system of a medical knowledge graph. The system comprises a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database and a behavior acquisition module. According to the method and the system, specialized medical knowledge in a certain disease field or general medical knowledge in the whole field can be managed based on a knowledge graph technology, so that the management efficiency and retrieval efficiency of the knowledge can be improved, and a personalized knowledge retrieval function can be provided for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a management system for a medical knowledge graph. Background Art

[0002] The knowledge system in the medical field is extremely large and complex, covering multiple aspects such as drugs, diseases, symptoms, treatment plans, clinical records, clinical trials, biochemical genes, medical examinations, and medical tests. With the continuous progress of medical technology and the in-depth development of clinical practice, medical knowledge shows an explosive growth trend. Traditional knowledge management methods (such as paper literature management methods) are difficult to meet the needs of quickly retrieving and efficiently using medical knowledge. As a new type of knowledge representation and management technology, the knowledge graph has received extensive attention and application in the field of artificial intelligence in recent years. The knowledge graph technology realizes the systematic representation and efficient management of complex knowledge systems by converting knowledge triples (entity-relationship-entity) into directed graphs. If the management of specialized medical knowledge in a certain disease field (such as cognitive impairment diseases, etc.) or general medical knowledge in the whole field can be based on the knowledge graph technology, it will surely improve the management efficiency and retrieval efficiency of knowledge. Summary of the Invention

[0003] The purpose of the present invention is to provide a management system for a medical knowledge graph in view of the defects of the prior art. The system includes: a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database, and a behavior collection module; among them, the graph processing module is used for extracting medical knowledge triples, updating the medical knowledge graph, and supplementing the medical knowledge graph; the graph database is used for storing a series of data tables of the medical knowledge graph; the information retrieval module is used for performing personalized graph retrieval in combination with user behavior characteristics; the user behavior database is used for storing user information; the behavior collection module is used for collecting user behavior data and analyzing user behavior characteristics. Through the present invention, the management of specialized medical knowledge in a certain disease field or general medical knowledge in the whole field can be based on the knowledge graph technology, which can not only improve the management efficiency and retrieval efficiency of knowledge, but also provide users with personalized knowledge retrieval functions.

[0004] To achieve the above object, an embodiment of the present invention provides a management system for a medical knowledge graph, the system including: a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database, and a behavior collection module;

[0005] The data receiving module is connected to the graph processing module; the graph database is respectively connected to the graph processing module and the information retrieval module; the user behavior database is respectively connected to the information retrieval module and the behavior collection module;

[0006] The data receiving module is used to receive the first original data packet sent by the client; preprocess the first original data packet to obtain the corresponding first preprocessed data packet; and send the first preprocessed data packet and the first original data packet to the graph processing module;

[0007] The graph processing module is used to perform duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet and the graph database; extract medical knowledge triples from the first preprocessed data packet to obtain the corresponding first triple set; and update the medical knowledge graph according to the first triple set and the graph database;

[0008] The graph processing module is also used to periodically predict whether there are potential new connection relationships on the latest medical knowledge graph and perform new edge complement processing on the current medical knowledge graph based on the prediction result to obtain the latest medical knowledge graph;

[0009] The graph database is used to store a series of data tables of the medical knowledge graph; the series of data tables at least includes a first resource record table, a first edge record table, a plurality of first node record tables, a first attribute record table, a first edge feature vector record table and a first node feature vector record table;

[0010] The information retrieval module is used to receive the first retrieval request sent by the client; extract the corresponding first user identifier and first retrieval text from the first retrieval request; perform retrieval keyword extraction processing on the first retrieval text to obtain the corresponding first keyword set; perform user behavior data addition processing according to the first user identifier, the first keyword set and the user behavior database; and perform knowledge graph retrieval processing according to the first user identifier, the first keyword set, the user behavior database and the graph database to obtain the corresponding first retrieval graph and send it back to the client; the first retrieval request includes the first user identifier and the first retrieval text; the first keyword set consists of one or more first keywords;

[0011] The user behavior database is used to store a first user record table and a plurality of first behavior record tables;

[0012] The behavior collection module is used to collect user behavior data through the client and perform user behavior data addition processing based on the collected data and the user behavior database;

[0013] The behavior collection module is also used to perform behavior analysis regularly according to the user behavior database.

[0014] Preferably, the first original data packet is composed of one or more first file packets; each of the first file packets is composed of a first medical file and a corresponding first file parameter set; the first file parameter set includes file type, file name, and content summary; the file type at least includes PDF files, table files, image files, audio files, and video files; when the file type is a PDF file, the corresponding first medical file is a medical literature material in a certain medical knowledge field, and the corresponding content summary includes the file source information, release time information, and content abstract information corresponding to the current first medical file; when the file type is a table file, the corresponding first medical file is a formatted medical data table; when the file type is an image file, audio file, or video file, the corresponding first medical file is a medical examination image, medical examination audio, or medical examination video generated by a certain medical examination corresponding to a certain disease, and the corresponding content summary includes the disease information, examination category information, and examination description information corresponding to the current first medical file; the medical examination images at least include X-ray images, CT images, ultrasound images, magnetic resonance images, and nuclear medicine images; the medical examination audio at least includes heart sound auscultation audio, lung auscultation audio, and abdominal auscultation audio; the medical examination videos at least include endoscopic examination videos and dynamic medical imaging videos;

[0015] The first preprocessed data packet is composed of one or more first preprocessed data; the first preprocessed data corresponds to the first file packet one by one; when the file type of the first file packet is a PDF file, the corresponding first preprocessed data includes a first data type and a first sentence sequence, and the first data type is specifically the first type, and the first sentence sequence is composed of multiple first sentences sorted in sequence; when the file type of the first file packet is a table file, the corresponding first preprocessed data includes the first data type and a first data table, and the first data type is specifically the second type; when the file type of the first file packet is an image file, the corresponding first preprocessed data includes the first data type, a first description, and a first image, and the first data type is the third type; when the file type of the first file packet is an audio file, the corresponding first preprocessed data includes the first data type, the first description, and a first audio, and the first data type is the fourth type; when the file type of the first file packet is a video file, the corresponding first preprocessed data includes the first data type, the first description, and a first video, and the first data type is the fifth type;

[0016] The first triple set includes multiple first triples; the first triples are composed of corresponding entity A, entity relationship R, and entity B according to the knowledge triple structure of entity-relationship-entity; the entity parameters of entity A and B are both composed of entity name, entity type, and entity attribute set; the entity type includes multiple medical entity types, at least including multiple drug entity types, multiple disease entity types, multiple disease symptom entity types, multiple disease treatment plan entity types, multiple disease clinical record entity types, multiple disease clinical experiment entity types, multiple biochemical gene entity types, multiple medical examination type entity types, multiple medical test type entity types, and the entity type is consistent with the type range of the knowledge graph node type; the entity attribute set corresponds one-to-one with the entity type, the entity attribute set includes multiple entity attributes, the entity attribute includes attribute name and attribute value, and the quantity and type of the entity attributes corresponding to each type of medical entity type are fixed; the entity relationship R includes multiple medical entity association types, and the entity relationship R is consistent with the relationship range of the knowledge graph edge association relationship;

[0017] The medical knowledge graph is composed of a first node set and a first edge set; the first node set includes multiple first nodes; the first edge set includes multiple first edges; the node parameters of the first node include a first node identifier, a first node name, a first node type, and a first node attribute set, the first node attribute set includes multiple first node attributes, and the first node attribute is composed of an attribute name and an attribute value; the edge parameters of the first edge include a first edge identifier, a first association type, and a first association node group; the first association node group includes a head node identifier and a tail node identifier;

[0018] When the first resource record table is not empty, it consists of one or more first resource records; the first resource record includes a resource identifier field, a resource type field, a resource name field, a content summary field, and a storage address field; the resource identifier field is set as the primary key;

[0019] When the first edge feature vector record table is not empty, it consists of one or more first edge feature vector records; the first edge feature vector record corresponds one-to-one with the first edge of the medical knowledge graph; the first edge feature vector record includes an edge vector identifier field and an edge feature vector field; the edge vector identifier field is set as the primary key;

[0020] When the first node feature vector record table is not empty, it consists of one or more first node feature vector records; the first node feature vector record corresponds one-to-one with the first node of the medical knowledge graph; the first node feature vector record includes a node vector identifier field and a node feature vector field; the node vector identifier field is set as the primary key;

[0021] When the first attribute record table is not empty, it consists of one or more first attribute records; each first attribute record includes an attribute identifier field, an attribute name field, and an attribute value sequence field; the attribute value sequence field is used to store an attribute value sequence; when the attribute value sequence is not empty, it is composed of one or more sequence elements sorted in chronological order, and each sequence element includes an addition time and an attribute value; the attribute identifier field is set as the primary key;

[0022] The first node record table corresponds one-to-one with the first node type of the medical knowledge graph, that is, the entity type; when the first node record table is not empty, it consists of one or more first node records; each first node record corresponds one-to-one with the first node of the medical knowledge graph; each first node record includes a node identifier field, a node name field, multiple attribute fields, and a node vector field; the total number and types of the attribute fields are consistent with the total number and types of the entity attributes corresponding to the entity type corresponding to the current first node record table; the node identifier field is set as the primary key field; each of the attribute fields and the node vector field is set as a foreign key field; each of the attribute fields forms a one-to-one foreign key-primary key mapping relationship with the attribute identifier field of a first attribute record; the node vector field forms a one-to-one foreign key-primary key mapping relationship with the node vector identifier field of a first node feature vector record;

[0023] When the first edge record table is not empty, it consists of one or more first edge records; each first edge record corresponds one-to-one with the first edge of the medical knowledge graph; each first edge record includes an edge identifier field, an association type field, a head node identifier field, a tail node identifier field, and an edge vector field; the total number and range of types of the association type field are consistent with the total number and range of relationships of the entity relationship R; the edge identifier field is set as the primary key field; the head node identifier field, the tail node identifier field, and the edge vector field are all set as foreign key fields; the head node identifier field forms a one-to-one foreign key-primary key mapping relationship with the node identifier field of a first node record; the tail node identifier field forms a one-to-one foreign key-primary key mapping relationship with the node identifier field of another first node record; the edge vector field forms a one-to-one foreign key-primary key mapping relationship with the edge vector identifier field of a first edge feature vector record;

[0024] The first line is in one-to-one correspondence with the user; each first line corresponds to a unique table index; when the first line is not empty, it consists of one or more first-line records; the first-line record includes a behavior record identification field, a behavior time field, a behavior type field, and a knowledge type field; the behavior type field includes at least retrieval, browsing, liking, forwarding, and collection; the knowledge type field includes at least all the entity types; the record identification field is set as the primary key;

[0025] When the first user record table is not empty, it consists of one or more first user records; the first user record is in one-to-one correspondence with the user; the first user record includes a user identification field, a name field, an age field, a gender field, a user type field, a personalized feature field, and a behavior table index field; the user type field includes at least multiple types of medical practitioner types and multiple types of medical student types; the personalized feature field is used to store a personalized feature set, which consists of one or more personalized features when not empty, and the personalized feature is one of the entity types; the user identification field is set as the primary key; the behavior table index field is the table index of the first behavior record table corresponding to the user corresponding to the current first user record.

[0026] Preferably, when the map processing module is specifically used for performing duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet, and the map database:

[0027] Take each first file packet of the first original data packet as the corresponding current file packet one by one; and extract the corresponding first medical file and the first file parameter set from the current file packet as the corresponding current medical file and current file parameter set; and extract the corresponding file type, file name, and content brief description from the current file parameter set as the corresponding current type, current name, and current description;

[0028] Perform word embedding encoding on the current name based on the bag-of-words encoding algorithm to obtain the corresponding first name encoding vector; perform word segmentation on the current description to obtain the corresponding current word segmentation sequence, and then perform word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain the corresponding first description encoding vector; and perform vector splicing on the obtained first name encoding vector and the first description encoding vector to obtain the corresponding first splicing vector;

[0029] And initialize the duplicate file check status as not duplicate; mark all the first resource records in the first resource record table of the atlas database whose resource type fields match the current type as corresponding records to be traversed; perform a round of traversal on all the records to be traversed; during this round of traversal, regard the currently traversed record to be traversed as the corresponding current record; perform word embedding encoding on the resource name field of the current record based on the bag-of-words encoding algorithm to obtain the corresponding second name encoding vector; perform word segmentation processing on the content summary field of the current record to obtain the corresponding current word segmentation sequence, and then perform word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain the corresponding second description encoding vector; splice the obtained second name encoding vector and the second description encoding vector to obtain the corresponding second splicing vector; calculate the vector similarity of the first and second splicing vectors based on the cosine vector similarity algorithm to obtain the corresponding current similarity; identify whether the current similarity exceeds the preset first similarity threshold; if so, set the duplicate file check status as duplicate and stop this round of traversal; if not, go to the next record to be traversed and continue traversing until the traversal of the last record to be traversed ends;

[0030] And identify the obtained duplicate file check status;

[0031] If the duplicate file check status is not duplicate, add a new first resource record in the first resource record table as the corresponding current new record; perform file storage on the current medical file and use the storage address as the corresponding current storage address; set a unique identifier for the current new record as the corresponding current record identifier; set the corresponding resource identifier field, resource type field, resource name field, content summary field and storage address field in the current new record based on the current record identifier, the current type, the current name, the current description and the current storage address;

[0032] If the duplicate file check status is duplicate, delete the current file package from the first original data packet, and synchronously delete the corresponding first preprocessed data in the first preprocessed data packet for the current file package.

[0033] Preferably, when the atlas processing module is specifically used for extracting medical knowledge triples from the first preprocessed data packet to obtain the corresponding first triple set:

[0034] Take each of the first preprocessed data packets as the corresponding current preprocessed data one by one; and take the first data type of the current preprocessed data as the corresponding current data type;

[0035] And identify the current data type;

[0036] If the current data type is the first type, identify the preset text triple extraction model set; if the text triple extraction model set is the first model set, perform medical knowledge triple recognition on the first sentence sequence of the current preprocessed data based on the preset first entity naming model, first relation extraction model, and first attribute extraction model corresponding to each entity type to obtain the corresponding first triple subset; if the text triple extraction model set is the second model set, perform medical knowledge triple recognition on the first sentence sequence of the current preprocessed data based on the preset second entity naming model, second relation extraction model, and second attribute extraction model corresponding to each entity type to obtain the corresponding first triple subset; the text triple extraction model set includes the first model set and the second model set; the first entity naming model, the first relation extraction model, and each of the first attribute extraction models are each implemented based on a stacking model framework with a multi-class logistic regression model as the base meta-model; the second entity naming model, the second relation extraction model, and each of the second attribute extraction models are each implemented based on an NLP model framework with a BERT model as the core encoder; the first triple subset includes one or more of the first triples;

[0037] If the current data type is the second type, perform knowledge triple recognition on the first data table of the current preprocessed data based on the preset formatted data table-medical knowledge triple conversion template to obtain the corresponding first triple subset;

[0038] If the current data type is the third type, perform medical knowledge triple recognition on the first description and the first image of the current preprocessed data and the preset image attribute prediction model corresponding to each medical examination image to obtain the corresponding first triple subset; each of the image attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network;

[0039] If the current data type is the fourth type, medical knowledge triple recognition is performed based on the first description and the first audio of the current preprocessed data and the preset audio attribute prediction models corresponding to various medical examination audios to obtain the corresponding first triple subset; each of the audio attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network;

[0040] If the current data type is the fifth type, medical knowledge triple recognition is performed based on the first description and the first video of the current preprocessed data and the preset video attribute prediction models corresponding to various medical examination videos to obtain the corresponding first triple subset; each of the video attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network;

[0041] All the first triple subsets obtained based on the first preprocessing data packet are merged to obtain the corresponding first set; and the duplicate first triples in the first set are removed, and the first set after deduplication is used as the corresponding first triple set.

[0042] Further, the first entity naming model is used to perform entity classification processing on each word segmentation in the first text input to the model according to a preset entity classification set and output the corresponding first entity classification vector; the entity classification set contains all the entity types; the first text is a single-sentence text; the first entity classification vector is composed of multiple first word segmentation classification vectors, and the first word segmentation classification vector corresponds one-to-one with the text word segmentation in the first text; the first word segmentation classification vector is composed of multiple first entity classification probabilities, and the first entity classification probability corresponds one-to-one with the entity type;

[0043] The first entity naming model consists of a first preprocessing unit, Na parallel base models BM-1 a and a meta-model MM-1; all the base models BM-1 a are implemented based on the principle of a multi-class logistic regression model, and all the base models BM-1 a have the same model structure but different hyperparameters. Na is the number of all hyperparameter combinations, and 1 ≤ base model index a ≤ Na; the meta-model MM-1 is also implemented based on the principle of a multi-class logistic regression model, and its hyperparameters are one of all the hyperparameter combinations of the multi-class logistic regression model;

[0044] The first relation extraction model is used to perform relation classification processing on the first entity pair features input to the model according to a preset set of association relations and output the corresponding first relation classification vector; the set of association relations includes all the medical entity association types corresponding to all the entity relations R; the first entity pair features include a first context word segmentation sequence, a head entity feature, and a tail entity feature; the first context word segmentation sequence is formed by sorting a plurality of first word segmentation texts; the head entity feature includes a head entity index and a head entity type; the tail entity feature includes a tail entity index and a tail entity type; the head and tail entity indexes are the word segmentation text indexes of the corresponding head and tail entities in the first context word segmentation sequence; the head and tail entity types are one type of the entity types; the first relation classification vector is composed of a plurality of first relation classification probabilities, and each of the first relation classification probabilities corresponds to one type of the medical entity association types;

[0045] The first relation extraction model consists of a second preprocessing unit, Nb parallel base models BM-2 b and a meta-model MM-2; all the base models BM-2 b are implemented based on the principle of a multi-class logistic regression model, and all the base models BM-2 b have the same model structure but different hyperparameters. Nb is the number of all hyperparameter combinations, and 1 ≤ base model index b ≤ Nb; the meta-model MM-2 is also implemented based on the principle of a multi-class logistic regression model, and its hyperparameters are one of all the hyperparameter combinations of the multi-class logistic regression model;

[0046] The number of models of the first attribute extraction model is equal to the number of entity types, and the first attribute extraction models correspond to the entity types one by one; each of the first attribute extraction models is used to perform attribute classification processing on the first entity features input to the model according to all the attribute type categories of the corresponding single entity attribute type set and output the corresponding first attribute classification vector; the single entity attribute type set corresponds to one of the entity types and is composed of a plurality of the attribute types, and each of the attribute types matches the attribute name of one type of the entity attributes corresponding to the current entity type; the first entity features include a second context word segmentation sequence and a first entity index; the second context word segmentation sequence is formed by sorting a plurality of second word segmentation texts; the first entity index is a word segmentation text index in the second context word segmentation sequence; the first attribute classification vector is composed of a plurality of first word segmentation attribute vectors, and the first word segmentation attribute vectors correspond one by one to the second word segmentation texts in the second context word segmentation sequence; the vector length of the first word segmentation attribute vector matches the total number of attribute types of the corresponding single entity attribute type set and is composed of a plurality of first attribute classification probabilities, and the first attribute classification probabilities correspond to the attribute types of the corresponding single entity attribute type set one by one;

[0047] The first attribute extraction model consists of a third preprocessing unit, Nc parallel base models BM-3 c and a meta-model MM-3; all the base models BM-3 c are implemented based on the principle of the multi-class logistic regression model, and all the base models BM-3 c have the same model structure but different hyperparameters. Nc is the number of all hyperparameter combinations, and 1 ≤ base model index c ≤ Nc; the meta-model MM-3 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameters are one of all hyperparameter combinations of the multi-class logistic regression model.

[0048] Furthermore, the second entity naming model is used to perform entity classification processing on each token in the second text input to the model according to the preset entity classification set and output the corresponding second entity classification vector; the second text is a single-segment text composed of one or more single-sentence texts; the second entity classification vector is composed of multiple second token classification vectors, and the second token classification vectors correspond one-to-one to the tokens in the second text; the second token classification vector is composed of multiple second entity classification probabilities, and the second entity classification probabilities correspond one-to-one to the entity types;

[0049] The second entity naming model consists of a fourth preprocessing unit, a first BERT model, a first linear network, and a first Softmax layer; the first BERT model is one of the basic BERT model, BioBERT, and ClinicalBERT that have completed pre-training; the first linear network is implemented based on one or more fully connected layers;

[0050] The second relationship extraction model is used to perform entity pair relationship classification processing on the basis of the preset association relationship set, the second token sequence, and the second entity classification vector input to the model and output the corresponding first token pair classification matrix; the second token sequence is the token sequence generated by the fourth preprocessing unit of the second entity naming model, and the second entity classification vector is the entity classification vector output by the second entity naming model; each row or column of the first token pair classification matrix corresponds one-to-one to the second tokens in the second token sequence; each matrix unit not on the diagonal in the first token pair classification matrix is a second relationship classification vector of a token pair, and each matrix unit on the diagonal of the first token pair classification matrix is an invalid classification vector with a vector length consistent with that of the second relationship classification vector and all vector data being zero or all negative values; the second relationship classification vector is composed of multiple second relationship classification probabilities, and each of the second relationship classification probabilities corresponds to one type of the medical entity association type;

[0051] The second relation extraction model consists of a fifth preprocessing unit, a second BERT model, a second linear network, and a second Softmax layer; the model structures of the first and second BERT models are the same, and the model parameters are the same; the second linear network is implemented based on one or more fully connected layers;

[0052] The number of models of the second attribute extraction model is equal to the number of entity types, and the second attribute extraction models correspond to the entity types one by one; each of the second attribute extraction models is used to perform attribute classification processing according to all attribute type categories of the corresponding single-entity attribute type set, based on the second word segmentation sequence and the first entity position input to the model, and output the corresponding second attribute classification vector; the second word segmentation sequence is the word segmentation sequence generated by the fourth preprocessing unit of the second entity naming model, and the first entity position is the word segmentation index corresponding to one of the second word segmentations in the second word segmentation sequence; the second attribute classification vector is composed of multiple second word segmentation attribute vectors, and the second word segmentation attribute vectors correspond to the second word segmentations in the second word segmentation sequence one by one; the vector length of the second word segmentation attribute vector matches the total number of attribute types of the corresponding single-entity attribute type set and is composed of multiple second attribute classification probabilities, and the second attribute classification probabilities correspond to the attribute types of the corresponding single-entity attribute type set one by one;

[0053] The second attribute extraction model consists of a sixth preprocessing unit, a third BERT model, a third linear network, and a third Softmax layer; the model structures of the first and third BERT models are the same, and the model parameters are the same; the third linear network is implemented based on one or more fully connected layers.

[0054] Furthermore, each type of medical examination image corresponds to a preset first analysis attribute set, and each first analysis attribute set consists of one or more preset first analysis attributes; each first analysis attribute corresponds to a first attribute name and a first attribute value range;

[0055] Each type of medical examination audio corresponds to a preset second analysis attribute set, and each second analysis attribute set consists of one or more preset second analysis attributes; each second analysis attribute corresponds to a second attribute name and a second attribute value range;

[0056] Each type of medical examination video corresponds to a preset third analysis attribute set, and each third analysis attribute set consists of one or more preset third analysis attributes; each third analysis attribute corresponds to a third attribute name and a third attribute value range;

[0057] The number of the image attribute prediction models is equal to the number of the medical examination images, and the image attribute prediction models correspond to the medical examination images one by one; each of the image attribute prediction models is used to perform image attribute prediction on the input first medical image according to the attribute analysis requirements of the corresponding first analysis attribute set and output a corresponding first image attribute vector; the first medical image is a type of the medical examination images; the first image attribute vector includes a plurality of first attribute prediction data; the first attribute prediction data corresponds to the first attribute names of the first analysis attribute set one by one;

[0058] The image attribute prediction model is composed of a first feature extraction network and a first prediction network; the first feature extraction network is implemented based on a type of image encoder model; the first prediction network is implemented based on a type of non - linear regression prediction model; the image encoder model at least includes a CNN network and a residual network; the non - linear regression prediction model at least includes an MLP model, a decision tree model, an SVR model, and a random forest model;

[0059] The number of the audio attribute prediction models is equal to the number of the medical examination audios, and the audio attribute prediction models correspond to the medical examination audios one by one; each of the audio attribute prediction models is used to perform audio attribute prediction on the input first medical audio according to the attribute analysis requirements of the corresponding second analysis attribute set and output a corresponding first audio attribute vector; the first medical audio is a type of the medical examination audios; the first audio attribute vector includes a plurality of second attribute prediction data; the second attribute prediction data corresponds to the second attribute names of the second analysis attribute set one by one;

[0060] The audio attribute prediction model is composed of a second feature extraction network and a second prediction network; the second feature extraction network is implemented based on a type of audio encoder model; the second prediction network is implemented based on a type of the non - linear regression prediction model; the audio encoder model at least includes a CNN network, an RNN network, a residual network, an LSTM model, and a Transformer model;

[0061] The number of the video attribute prediction models is equal to the number of the medical examination videos, and the video attribute prediction models correspond to the medical examination videos one by one; each of the video attribute prediction models is used to perform video attribute prediction on the input first medical video according to the attribute analysis requirements of the corresponding third analysis attribute set and output a corresponding first video attribute vector; the first medical video is a type of the medical examination videos; the first video attribute vector includes a plurality of third attribute prediction data; the third attribute prediction data corresponds to the third attribute names of the third analysis attribute set one by one;

[0062] The video attribute prediction model is composed of a third feature extraction network and a third prediction network; the third feature extraction network is implemented based on a type of video encoder model; the third prediction network is implemented based on a type of the non - linear regression prediction model; the video encoder model includes at least a 3D CNN network, a TCN network, an RNN network, an LSTM model, and a Transformer model;

[0063] The input end of the third feature extraction network is connected to the input end of the video attribute prediction model, and the output end is connected to the input end of the third prediction network; the output end of the third prediction network is connected to the output end of the video attribute prediction model.

[0064] Preferably, when the atlas processing module is specifically used for updating the medical knowledge atlas according to the first triple set and the atlas database:

[0065] Perform a round of traversal on all the first triples in the first triple set; and during this round of traversal, regard the currently traversed first triple as the corresponding current triple; and regard the entity names, entity types, and entity attribute sets of entities A and B in the current triple as the corresponding names A, B, types A, B, and attribute sets A, B; and regard the entity relationship R in the current triple as a corresponding A - B association relationship; and regard the first node record tables corresponding to types A and B in the atlas database as the corresponding node record tables A and B; and regard the first node record whose node name field in the node record tables A and B matches the corresponding names A and B as the corresponding matching node records A and B; and regard the first edge record in the first edge record table of the atlas database whose association type field matches the A - B association relationship, and the first node identifier field has a mapping relationship with the node identifier field of the matching node record A, and the tail node identifier field has a mapping relationship with the node identifier field of the matching node record B as the corresponding matching edge record C; and identify the matching node records A and B;

[0066] If both the matching node records A and B are not empty, then regard the matching node records A and B as the corresponding current matching node records in sequence, and perform old node update processing based on the current matching node records and the current triple, and perform single - edge addition processing based on the matching node records A and B and the A - B association relationship when the matching edge record C is empty;

[0067] If one of the matching node records A and B is empty, the matching node record A or B that is not empty is used as the corresponding current matching node record, and the entity A or B corresponding to the matching node record A or B that is empty is used as the corresponding current newly added entity, and based on the current matching node record and the current triplet, the old node is updated, and based on the current newly added entity, a single node is added to obtain the corresponding newly added node record, and a pair of new matching node records A and B are formed by the current matching node record and the newly added node record, and a unilateral addition is performed based on the new matching node records A and B and the AB association relationship;

[0068] If the matching node records A and B are both empty, the entities A and B of the current triplet are taken as the corresponding current newly added entities in turn, and a single node addition process is performed based on the current newly added entities to obtain the corresponding newly added node records, and a pair of new matching node records A and B are formed by the two newly added node records corresponding to the entities A and B, and a unilateral addition process is performed based on the new matching node records A and B and the AB association relationship;

[0069] After this round of traversal of all the first triples of the first triple set is completed, the latest first node set is constructed based on all the first node record tables and the first attribute record tables on the graph database; and the latest first edge set is constructed based on the first edge record table on the graph database; and the obtained first node set and the first edge set constitute the latest medical knowledge graph.

[0070] Preferably, the graph processing module is specifically used for periodically predicting whether there are any potential new connection relationships on the latest medical knowledge graph and performing new edge filling processing on the current medical knowledge graph based on the prediction result to obtain the latest medical knowledge graph:

[0071] The first node set of the latest medical knowledge graph is regularly used as the corresponding node set V; the node set V is composed of multiple nodes v, each of the node v is assigned a unique integer value as the corresponding node index, the node v corresponds to the first node of the first node set one by one, and the node feature of the node v is composed of the first node name, the first node type and the first node attribute set of the corresponding first node;

[0072] Construct a virtual edge e with direction features and association relationship features between every two nodes v in the node set V, and form a corresponding virtual edge set E consisting of all the obtained virtual edges e; initialize the direction features and the association relationship features of all the virtual edges e in the virtual edge set E to the corresponding invalid directions and invalid relationships; the edge features of each virtual edge e at least include the direction features and the association relationship features; the direction features include forward, reverse, and invalid direction; if the node with the larger node index among the two nodes v corresponding to each virtual edge e is denoted as the large-index node and the node with the smaller node index is denoted as the small-index node, then when the direction feature is forward, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the large-index node to the small-index node, when it is reverse, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the small-index node to the large-index node, and when it is invalid direction, it indicates that the direction of the directed edge corresponding to the current virtual edge e is unknown; the association relationship features include all the medical entity association types and invalid relationships of the medical knowledge graph;

[0073] Perform a round of traversal on all the first edges in the first edge set of the latest medical knowledge graph; during this round of traversal, regard the currently traversed first edge as the corresponding current entity edge; regard the virtual edge e in the virtual edge set E corresponding to the current entity edge as the corresponding current virtual edge; reset the direction feature of the current virtual edge based on the node index of the first node corresponding to the head and tail node identifiers of the current entity edge, and reset the association relationship feature of the current virtual edge based on the first association type of the current entity edge; after this round of traversal, denote the latest virtual edge set E as the corresponding initial edge set E ini ;

[0074] And from the node set V and the initial edge set E ini Form a graph structure data denoted as the corresponding initial graph; input the initial graph into a preset knowledge graph edge prediction model, and the knowledge graph edge prediction model predicts the direction type and association relationship type of each virtual edge e in the input initial graph of the model to obtain the corresponding predicted edge set E * ; the knowledge graph edge prediction model is a prediction model with a graph neural network model as the core encoder; the predicted edge set E * Consists of multiple predicted edges e * ; the predicted edge set E * The predicted edge e * Of which corresponds one-to-one with the virtual edge e in the initial edge set E ini ; each predicted edge e *The edge features at least include the direction feature and the association relationship feature;

[0075] For all the virtual edges e in the initial edge set E ini where all the direction features and the association relationship features are the corresponding invalid directions and invalid relationships, mark the virtual edges e as the corresponding initial invalid edges; and for each prediction edge e in the prediction edge set E * where the direction feature is not an invalid direction, the association relationship feature is not an invalid relationship, and it corresponds to an initial invalid edge, * mark it as the corresponding newly added valid edge;

[0076] And when the total number of the obtained newly added valid edges is not zero, perform a round of traversal on all the obtained newly added valid edges; during this round of traversal, regard the currently traversed newly added valid edge as the corresponding currently newly added edge; and based on the direction feature of the currently newly added edge, perform corresponding head and tail node markings on the two first nodes corresponding to the currently newly added edge; and record the two first node records corresponding to the current head and tail nodes on the knowledge graph database as a pair of new matching node records A and B; and regard the association relationship feature of the currently newly added edge as a new A - B association relationship; and perform single - edge addition processing based on the matching node records A, B and the A - B association relationship;

[0077] And after the end of this round of traversal of all the newly added valid edges, add all the newly added valid edges to the current medical knowledge graph to obtain the latest medical knowledge graph.

[0078] Furthermore, the knowledge graph edge prediction model is used to perform prediction processing according to the direction types and association relationship types of each virtual edge e of the input initial graph to obtain the corresponding prediction edge set E * ;

[0079] The knowledge graph edge prediction model consists of a graph embedding encoding module, a graph feature encoder, a direction prediction head, a relationship prediction head, and an output module; the graph feature encoder is implemented based on a type of graph neural network model, and the graph neural network model at least includes a GCN model, a GNN model, and an ApeGNN model; the direction prediction head and the relationship prediction head are each implemented based on an MLP model.

[0080] Preferably, when the information retrieval module specifically sends the first retrieval graph obtained by performing knowledge graph retrieval processing according to the first user identifier, the first keyword set, the user behavior database, and the knowledge graph database to the client:

[0081] Take the first user record corresponding to the first user identifier in the first user record table of the user behavior database as the corresponding current user record; and extract the personalized feature field of the current user record as the corresponding current user feature set;

[0082] And the first keyword vector corresponding to the first keyword set; and use a preset node feature encoding algorithm to encode the first keyword vector to obtain the corresponding first feature vector; and take the node feature vector fields of each of the first node feature vector records in the first node feature vector record table in the graph database as the corresponding first comparison vectors; and calculate the vector similarity between the first feature vector and each of the first comparison vectors to obtain the corresponding second similarity; and record the first node feature vector record corresponding to the second similarity that exceeds the preset second similarity threshold as the candidate feature record; and take the first node record having a mapping relationship with each of the candidate feature records as the corresponding preliminary selection node record;

[0083] And identify the current user feature set; if the current user feature set is empty, record all the preliminary selection node records as the corresponding second selection node records; if the current user feature set is not empty, only record the preliminary selection node records that match each of the personalized features of the current user feature set as the corresponding second selection node records;

[0084] And perform a round of traversal on all the second selection node records; and during this round of traversal, take the currently traversed second selection node record as the corresponding current node record; and take the first node corresponding to the current node record in the medical knowledge graph as the corresponding current node; and select a sub-graph centered on the current node in the medical knowledge graph as the corresponding first sub-graph; the maximum node distance between each of the first nodes in the first sub-graph and the current node is K, where K is a preset positive integer, and the node distance is the total number of other nodes between two directly or indirectly connected first nodes;

[0085] And at the end of this round of traversal of all the second selection node records, send back the corresponding first retrieval graph composed of all the first sub-graphs to the client.

[0086] Preferably, the behavior acquisition module is specifically used for: when performing behavior analysis according to the user behavior database regularly:

[0087] Regularly take each of the first user records in the first user record table of the user behavior database as the corresponding current user record;

[0088] Extract the age field, gender field, and user type field of the current user record as the corresponding user age, user gender, and user type, and form a corresponding user basic parameter;

[0089] Use the first behavior record table corresponding to the current user record as the corresponding current behavior record table; extract the last X first behavior records in the current behavior record table and sort them in chronological order to form a corresponding first record sequence; extract the behavior time field, behavior type, and knowledge type field of each first behavior record in the first record sequence as the corresponding first time, first behavior type, and first entity type sequence to form a corresponding first behavior vector; and sort all the obtained first behavior vectors in chronological order to form a corresponding user behavior sequence; X is a preset positive integer;

[0090] Input the user basic parameter and the user behavior sequence into a preset user behavior prediction model, and the user behavior prediction model performs personalized feature prediction processing based on the user basic parameter and the user behavior sequence to obtain a corresponding prediction feature set; the user behavior prediction model is a classification prediction model implemented based on a deep learning model framework; when the prediction feature set is not empty, it consists of one or more prediction features, and the prediction feature is a type of entity type;

[0091] When the current prediction feature set is not empty, reset the personalized feature field of the current user record based on the current prediction feature set.

[0092] Furthermore, the user behavior prediction model is used to perform personalized feature prediction processing based on the user basic parameter and the user behavior sequence input into the model to obtain the corresponding prediction feature set;

[0093] The user behavior prediction model consists of a behavior feature extraction network, a basic feature encoding module, a feature fusion module, a personalized prediction head, and a prediction output module; the behavior feature extraction network includes at least an LSTM model and a bi-LSTM model; the personalized prediction head is implemented based on a type of multi-classification prediction model.

[0094] An embodiment of the present invention provides a management system for a medical knowledge graph. The system includes: a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database, and a behavior collection module; wherein, the graph processing module is used for extracting medical knowledge triples, updating the medical knowledge graph, and complementing the medical knowledge graph; the graph database is used for storing a series of data tables of the medical knowledge graph; the information retrieval module is used for performing personalized graph retrieval in combination with user behavior characteristics; the user behavior database is used for storing user information; the behavior collection module is used for collecting user behavior data and analyzing user behavior characteristics. The embodiment of the present invention can achieve the technical purpose of managing the specialized medical knowledge in a certain disease field or the general medical knowledge in the whole field based on the knowledge graph technology. Based on the embodiment of the present invention, not only the management efficiency and retrieval efficiency of knowledge are improved, but also a personalized knowledge retrieval function is provided for users. Description of the Drawings

[0095] Figure 1 It is a module structure diagram of a management system for a medical knowledge graph provided by an embodiment of the present invention;

[0096] Figure 2 It is a schematic diagram of a first node and a first edge provided by an embodiment of the present invention;

[0097] Figure 3 It is a model structure diagram of a first entity naming model, a first relationship extraction model, and a first attribute extraction model provided by an embodiment of the present invention;

[0098] Figure 4 It is a model structure diagram of a second entity naming model, a second relationship extraction model, and a second attribute extraction model provided by an embodiment of the present invention;

[0099] Figure 5 It is a model structure diagram of an image attribute prediction model, an audio attribute prediction model, and a video attribute prediction model provided by an embodiment of the present invention;

[0100] Figure 6 It is a model structure diagram of a knowledge graph edge prediction model provided by an embodiment of the present invention;

[0101] Figure 7 It is a data structure diagram of a first resource record table, a first edge record table, a plurality of first node record tables, a first attribute record table, a first edge feature vector record table, and a first node feature vector record table provided by an embodiment of the present invention;

[0102] Figure 8 It is a data structure diagram of a first user record table and a first behavior record table provided by an embodiment of the present invention;

[0103] Figure 9This is the model structure diagram of the user behavior prediction model provided by the embodiments of the present invention. Detailed implementation manners

[0104] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0105] The medical knowledge graph management system 1 provided by the embodiments of the present invention, as Figure 1 shown in the module structure diagram of a medical knowledge graph management system provided by the embodiments of the present invention, mainly includes: a data receiving module 11, a graph processing module 12, a graph database 13, an information retrieval module 14, a user behavior database 15, and a behavior collection module 16.

[0106] The connection relationships of the components of the system of the present invention are: the data receiving module 11 is connected to the graph processing module 12; the graph database 13 is respectively connected to the graph processing module 12 and the information retrieval module 14; the user behavior database 15 is respectively connected to the information retrieval module 14 and the behavior collection module 16.

[0107] (1) Data receiving module 11:

[0108] The data receiving module 11 is used to receive a first original data packet sent by the client 2; perform data preprocessing on the first original data packet to obtain a corresponding first preprocessed data packet; and send the first preprocessed data packet and the first original data packet to the graph processing module 12.

[0109] Here, the first original data packet of the embodiment of the present invention is composed of one or more first file packets; each first file packet is composed of a first medical file and a corresponding first file parameter set; the first file parameter set includes file type, file name, and content brief description; the file type at least includes PDF file, table file, image file, audio file, and video file; when the file type is a PDF file, the corresponding first medical file is a medical literature material in a certain medical knowledge field, and the corresponding content brief description includes file source information, release time information, and content summary information corresponding to the current first medical file; when the file type is a table file, the corresponding first medical file is a formatted medical data table; when the file type is an image file, audio file, or video file, the corresponding first medical file is a medical examination image, medical examination audio, or medical examination video generated by a certain type of medical examination corresponding to a certain type of disease, and the corresponding content brief description includes disease information, examination category information, and examination description information corresponding to the current first medical file; medical examination images at least include X-ray images, CT images, ultrasound images, magnetic resonance images, and nuclear medicine images; medical examination audio at least includes heart sound auscultation audio, lung auscultation audio, and abdominal auscultation audio; medical examination videos at least include endoscopic examination videos and dynamic medical imaging videos.

[0110] The first preprocessing data packet of the embodiment of the present invention is composed of one or more first preprocessing data; the first preprocessing data corresponds to the first file packet one by one; when the file type of the first file packet is a PDF file, the corresponding first preprocessing data includes a first data type and a first sentence sequence, and the first data type is specifically the first type, and the first sentence sequence is composed of multiple first sentences sorted in sequence; when the file type of the first file packet is a table file, the corresponding first preprocessing data includes a first data type and a first data table, and the first data type is specifically the second type; when the file type of the first file packet is an image file, the corresponding first preprocessing data includes a first data type, a first description, and a first image, and the first data type is the third type; when the file type of the first file packet is an audio file, the corresponding first preprocessing data includes a first data type, a first description, and a first audio, and the first data type is the fourth type; when the file type of the first file packet is a video file, the corresponding first preprocessing data includes a first data type, a first description, and a first video, and the first data type is the fifth type.

[0111] In a specific implementation manner of the embodiment of the present invention, the data receiving module 11 is specifically used for when preprocessing the data of the first original data packet to obtain the corresponding first preprocessing data packet:

[0112] Step A1: Take each first file package in the first original data packet as the corresponding current file package one by one; extract the corresponding first medical file and first file parameter set from the current file package as the corresponding current medical file and current file parameter set; and extract the corresponding file type, file name, and content summary from the current file parameter set as the corresponding current type, first name, and first description.

[0113] Step A2: Identify the current type.

[0114] Step A3: If the current type is a PDF file, perform text recognition on the current medical file based on OCR technology to obtain the corresponding first long text; perform data cleaning on the first long text to obtain the corresponding second long text; perform sentence recognition on the second long text to obtain the corresponding multiple first sentences to form the corresponding first sentence sequence; set the corresponding first data type to the first type; and form the corresponding first preprocessed data from the obtained first data type and first sentence sequence.

[0115] Step A4: If the current type is a table file, take the medical data table of the current medical file as the corresponding first data table; set the corresponding first data type to the second type; and form the corresponding first preprocessed data from the obtained first data type and first data table.

[0116] Step A5: If the current type is an image file, take the current medical file as the corresponding first image; set the corresponding first data type to the third type; and form the corresponding first preprocessed data from the obtained first data type, first description, and first image.

[0117] Step A6: If the current type is an audio file, take the current medical file as the corresponding first audio; set the corresponding first data type to the fourth type; and form the corresponding first preprocessed data from the obtained first data type, first description, and first audio.

[0118] Step A7: If the current type is a video file, take the current medical file as the corresponding first video; set the corresponding first data type to the fifth type; and form the corresponding first preprocessed data from the obtained first data type, first description, and first video.

[0119] Step A8: Form the corresponding first preprocessed data packet from all the first preprocessed data obtained based on the first original data packet.

[0120] (2) Atlas Processing Module 12:

[0121] The atlas processing module 12 is used to perform duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet, and the atlas database 13; extract medical knowledge triples from the first preprocessed data packet to obtain a corresponding first triple set; and update the medical knowledge atlas according to the first triple set and the atlas database 13.

[0122] The atlas processing module 12 is also used to periodically predict whether there are potential new connection relationships in the latest medical knowledge atlas and perform new edge complementation processing on the current medical knowledge atlas based on the prediction result to obtain the latest medical knowledge atlas.

[0123] Here, the first triple set of the embodiments of the present invention includes a plurality of first triples; the first triples are composed of corresponding entity A, entity relationship R, and entity B according to the knowledge triple structure of entity-relationship-entity; wherein, the entity parameters of entity A and B are both composed of entity name, entity type, and entity attribute set; the entity type includes multiple medical entity types, which at least include multiple drug entity types, multiple disease entity types, multiple disease symptom entity types, multiple disease treatment plan entity types, multiple disease clinical record entity types, multiple disease clinical experiment entity types, multiple biochemical gene entity types, multiple medical examination type entity types, multiple medical test type entity types, and the entity type is consistent with the type range of the knowledge graph node type; the entity attribute set corresponds to the entity type one by one, the entity attribute set includes multiple entity attributes, the entity attribute includes attribute name and attribute value, and the quantity and type of entity attributes corresponding to each type of medical entity type are fixed; the entity relationship R includes multiple medical entity association types, and the entity relationship R is consistent with the relationship range of the knowledge graph edge association relationship.

[0124] The medical knowledge atlas of the embodiments of the present invention is composed of a first node set and a first edge set; the first node set includes multiple first nodes; the first edge set includes multiple first edges; as Figure 2 shown, the node parameters of the first node of the embodiments of the present invention include a first node identifier, a first node name, a first node type, and a first node attribute set, the first node attribute set includes multiple first node attributes, and the first node attribute is composed of an attribute name and an attribute value; the edge parameters of the first edge include a first edge identifier, a first association type, and a first association node group; the first association node group includes a head node identifier and a tail node identifier.

[0125] In another specific implementation manner of the embodiments of the present invention, when the atlas processing module 12 is specifically used to perform duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet, and the atlas database 13:

[0126] Step B1: Take each first file packet of the first original data packet as the corresponding current file packet one by one; extract the corresponding first medical file and the first file parameter set from the current file packet as the corresponding current medical file and the current file parameter set; and extract the corresponding file type, file name, and content brief from the current file parameter set as the corresponding current type, current name, and current description;

[0127] Step B2: Perform word embedding encoding on the current name based on the bag-of-words encoding algorithm to obtain the corresponding first name encoding vector; perform word segmentation on the current description to obtain the corresponding current word segmentation sequence, and then perform word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain the corresponding first description encoding vector; perform vector splicing on the obtained first name encoding vector and the first description encoding vector to obtain the corresponding first splicing vector;

[0128] Step B3: Initialize the duplicate file check status as not duplicate; mark all the first resource records in the first resource record table of the atlas database 13 whose resource type fields match the current type as the corresponding records to be traversed; perform a round of traversal on all the records to be traversed; during this round of traversal, take the currently traversed record to be traversed as the corresponding current record; perform word embedding encoding on the resource name field of the current record based on the bag-of-words encoding algorithm to obtain the corresponding second name encoding vector; perform word segmentation on the content brief field of the current record to obtain the corresponding current word segmentation sequence, and then perform word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain the corresponding second description encoding vector; perform vector splicing on the obtained second name encoding vector and the second description encoding vector to obtain the corresponding second splicing vector; calculate the vector similarity of the first and second splicing vectors based on the cosine vector similarity algorithm to obtain the corresponding current similarity; identify whether the current similarity exceeds the preset first similarity threshold; if so, set the duplicate file check status as duplicate and stop this round of traversal; if not, move to the next record to be traversed and continue traversing until the last record to be traversed is traversed;

[0129] Here, the first similarity threshold is a preset threshold parameter;

[0130] Step B4: Identify the obtained duplicate file check status;

[0131] Step B5, if the duplicate file check status is "not duplicate", add a first resource record in the first resource record table as the corresponding current new record; store the current medical file and use the storage address as the corresponding current storage address; set a unique identifier for the current new record as the corresponding current record identifier; and set the corresponding resource identifier field, resource type field, resource name field, content brief description field, and storage address field in the current new record based on the current record identifier, current type, current name, current description, and current storage address.

[0132] Step B6, if the duplicate file check status is "duplicate", delete the current file package from the first original data package, and synchronously delete the corresponding first preprocessed data in the first preprocessed data package as well.

[0133] In another specific implementation manner of the embodiment of the present invention, the atlas processing module 12 is specifically configured to, when performing medical knowledge triple extraction processing on the first preprocessed data package to obtain the corresponding first triple set:

[0134] Step C1, take each first preprocessed data package in the first preprocessed data package as the corresponding current preprocessed data; and use the first data type of the current preprocessed data as the corresponding current data type;

[0135] Step C2, and identify the current data type;

[0136] Step C3, if the current data type is the first type, identify the preset text triple extraction model set; if the text triple extraction model set is the first model set, perform medical knowledge triple identification on the first sentence sequence of the current preprocessed data based on the preset first entity naming model, first relationship extraction model, and first attribute extraction models corresponding to various entity types to obtain the corresponding first triple subset; if the text triple extraction model set is the second model set, perform medical knowledge triple identification on the first sentence sequence of the current preprocessed data based on the preset second entity naming model, second relationship extraction model, and second attribute extraction models corresponding to various entity types to obtain the corresponding first triple subset;

[0137] Here, the text triple extraction model set of the embodiment of the present invention includes the first model set and the second model set; the first entity naming model, the first relationship extraction model, and each first attribute extraction model are each implemented based on a stacking model framework with a multi-class logistic regression model as the base and meta model, as Figure 3 shown; the second entity naming model, the second relationship extraction model, and each second attribute extraction model are each implemented based on an NLP model framework with a BERT model as the core encoder, as Figure 4as shown; the first triple subset includes one or more first triples;

[0138] Step C4, if the current data type is the second type, perform knowledge triple recognition on the first data table of the current preprocessed data based on a preset formatted data table - medical knowledge triple conversion template to obtain the corresponding first triple subset;

[0139] Here, the formatted data table - medical knowledge triple conversion template of the embodiments of the present invention is a preset data conversion template, and these templates can directly perform triple format conversion on the table records of various formatted data tables based on a series of corresponding rules between preset data table fields and triple elements;

[0140] Step C5, if the current data type is the third type, perform medical knowledge triple recognition on the first description and the first image of the current preprocessed data and a preset image attribute prediction model corresponding to various medical examination images to obtain the corresponding first triple subset;

[0141] Here, each type of medical examination image in the embodiments of the present invention corresponds to a preset first analysis attribute set, and each first analysis attribute set is composed of one or more preset first analysis attributes. Each first analysis attribute corresponds to a first attribute name and a first attribute value range; the image attribute prediction models in the embodiments of the present invention are each implemented based on a type of deep learning model framework composed of a feature extraction network and an attribute prediction network, as Figure 5 as shown;

[0142] Step C6, if the current data type is the fourth type, perform medical knowledge triple recognition on the first description and the first audio of the current preprocessed data and a preset audio attribute prediction model corresponding to various medical examination audios to obtain the corresponding first triple subset;

[0143] Here, each type of medical examination audio in the embodiments of the present invention corresponds to a preset second analysis attribute set, and each second analysis attribute set is composed of one or more preset second analysis attributes; each second analysis attribute corresponds to a second attribute name and a second attribute value range; the audio attribute prediction models in the embodiments of the present invention are each implemented based on a type of deep learning model framework composed of a feature extraction network and an attribute prediction network, as Figure 5 as shown;

[0144] Step C7, if the current data type is the fifth type, perform medical knowledge triple recognition on the first description and the first video of the current preprocessed data and a preset video attribute prediction model corresponding to various medical examination videos to obtain the corresponding first triple subset;

[0145] Here, each type of medical examination video in the embodiments of the present invention corresponds to a preset third analysis attribute set, and each third analysis attribute set is composed of one or more preset third analysis attributes. Each third analysis attribute corresponds to a third attribute name and a third attribute value range; the video attribute prediction models in the embodiments of the present invention are each implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network, such as Figure 5 shown;

[0146] Step C8, and merge all the first triple subsets obtained based on the first preprocessing data packet to obtain the corresponding first set; and perform deduplication processing on the repeated first triples in the first set, and use the deduplicated first set as the corresponding first triple set.

[0147] In another specific implementation manner of the embodiments of the present invention, the first entity naming model is specifically configured to perform entity classification processing on each word segmentation in the first text input to the model according to a preset entity classification set and output the corresponding first entity classification vector.

[0148] Here, the entity classification set in the embodiments of the present invention includes all entity types; the first text is a single-sentence text; the first entity classification vector is composed of multiple first word segmentation classification vectors, and the first word segmentation classification vectors correspond one by one to the word segmentations in the first text; the first word segmentation classification vector is composed of multiple first entity classification probabilities, and the first entity classification probabilities correspond one by one to the entity types.

[0149] Such as Figure 3 shown, the first entity naming model in the embodiments of the present invention is composed of a first preprocessing unit, Na parallel base models BM-1 a and a meta-model MM-1; it should be noted that: all the base models BM-1 a are implemented based on the principle of a multi-class logistic regression model, and all the base models BM-1 a have the same model structure but different hyperparameters. Na is the number of all hyperparameter combinations, and 1 ≤ base model index a ≤ Na; the meta-model MM-1 is also implemented based on the principle of a multi-class logistic regression model, and its hyperparameters are one of all the hyperparameter combinations of the multi-class logistic regression model.

[0150] The connection relationship of each component of the first entity naming model in the embodiments of the present invention is: the input end of the first preprocessing unit is connected to the model input end of the first entity naming model, and the output end is connected to the input ends of each base model BM-1 a ; the output ends of each base model BM-1 a are connected to one input end of the meta-model MM-1; the output end of the meta-model MM-1 is connected to the model output end of the first entity naming model.

[0151] The functions of each component of the first entity naming model in the embodiment of the present invention are as follows:

[0152] 1) The first preprocessing unit is used to perform word segmentation on the first text input to the model to obtain a corresponding first word segmentation sequence; and perform word embedding encoding on the first word segmentation sequence based on the TF-IDF encoding algorithm to obtain a corresponding first encoding vector and send it to each base model BM-1 a Send;

[0153] Among them, the first word segmentation sequence is composed of multiple first word segments sorted in order; the first encoding vector is composed of multiple first word segmentation encoding vectors concatenated in order, and the first word segmentation encoding vector corresponds to the first word segment one by one;

[0154] 2) Each base model BM-1 a is used to perform entity classification prediction according to the first encoding vector to obtain a corresponding first prediction vector and send it to the meta-model MM-1;

[0155] Among them, the first prediction vector is composed of multiple first sub-prediction vectors concatenated in order, the first sub-prediction vector is composed of multiple first prediction probabilities, the first sub-prediction vector corresponds to the first word segment one by one, and the first prediction probability corresponds to the entity type one by one;

[0156] 3) The meta-model MM-1 is used to splice the first sub-prediction vectors corresponding to each first word segment in the Na first prediction vectors received, and use the spliced vector as the corresponding second word segmentation encoding vector; and form a corresponding second encoding vector from all the obtained second word segmentation encoding vectors; and perform entity classification prediction according to the second encoding vector to obtain a corresponding first word segment classification vector.

[0157] In another specific implementation manner of the embodiment of the present invention, the first relationship extraction model is specifically used to perform relationship classification processing on the first entity pair features input to the model according to a preset set of association relationships and output a corresponding first relationship classification vector.

[0158] Here, the set of association relationships in the embodiment of the present invention includes all medical entity association types corresponding to all entity relationships R; the first entity pair features include a first context word segmentation sequence, a head entity feature, and a tail entity feature; the first context word segmentation sequence is composed of multiple first word segment texts sorted in order; the head entity feature includes a head entity index and a head entity type; the tail entity feature includes a tail entity index and a tail entity type; the head and tail entity indexes are the word segment text indexes of the corresponding head and tail entities in the first context word segmentation sequence; the head and tail entity types are a type of entity type; the first relationship classification vector is composed of multiple first relationship classification probabilities, and each first relationship classification probability corresponds to a type of medical entity association type.

[0159] Such as Figure 3As shown in the figure, the first relation extraction model of the embodiment of the present invention is composed of a second preprocessing unit, Nb parallel base models BM-2 b and a meta-model MM-2; it should be noted that: all base models BM-2 b are implemented based on the principle of the multi-class logistic regression model, and all base models BM-2 b have the same model structure but different hyperparameters. Nb is the number of all hyperparameter combinations, and 1 ≤ base model index b ≤ Nb; the meta-model MM-2 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameters are one of all hyperparameter combinations of the multi-class logistic regression model.

[0160] The connection relationships of the components of the first relation extraction model of the embodiment of the present invention are as follows: the input end of the second preprocessing unit is connected to the model input end of the first relation extraction model, and the output end is connected to the input ends of each base model BM-2 b ; the output ends of each base model BM-2 b are connected to one input end of the meta-model MM-2; the output end of the meta-model MM-2 is connected to the model output end of the first relation extraction model.

[0161] The functions of the components of the first relation extraction model of the embodiment of the present invention are as follows:

[0162] 1) The second preprocessing unit is used to extract the corresponding first context token sequence, head entity index, head entity type, tail entity index, and tail entity type from the first entity pair features input to the model; and perform one-hot encoding on the head and tail entity types to obtain the corresponding head and tail entity type encodings; and perform word embedding encoding on the first context token sequence based on the TF-IDF encoding algorithm to obtain the corresponding third encoding vector. The third encoding vector is sequentially spliced by multiple third token encoding vectors, and the third token encoding vector corresponds one-to-one with the first token text in the first context token sequence; and extract the third token encoding vectors corresponding to the head and tail entity indexes in the third encoding vector as the corresponding head and tail entity token encoding vectors; and splice the obtained head entity type encoding and the head entity token encoding vector to obtain the corresponding head entity encoding vector, and splice the obtained tail entity type encoding and the tail entity token encoding vector to obtain the corresponding tail entity encoding vector; and splice the obtained head and tail entity encoding vectors into the corresponding first entity pair encoding vector and send it to each base model BM-2 b ;

[0163] 2) Each base model BM-2 b is used to perform head and tail entity relationship classification prediction based on the first entity pair encoding vector to obtain the corresponding second prediction vector and send it to the meta-model MM-2;

[0164] Among them, the second prediction vector consists of multiple second prediction probabilities, and each second prediction probability corresponds to a type of medical entity association type.

[0165] 3) The meta-model MM-2 is used to splice the second prediction probabilities corresponding to various medical entity association types in the received Nb second prediction vectors and use the spliced vector as the corresponding fourth word segmentation encoding vector; and all the obtained fourth word segmentation encoding vectors form the corresponding fourth encoding vector; and relationship classification prediction is performed according to the fourth encoding vector to obtain the corresponding first relationship classification vector.

[0166] In the embodiment of the present invention, the number of models of the first attribute extraction model is equal to the number of entity types, that is, the first attribute extraction model corresponds one-to-one with the entity type. In another specific implementation manner of the embodiment of the present invention, each first attribute extraction model is specifically used to perform attribute classification processing on all attribute type categories of the corresponding single-entity attribute type set according to the first entity feature input by the model and output the corresponding first attribute classification vector.

[0167] Here, each single-entity attribute type set in the embodiment of the present invention corresponds to an entity type and consists of multiple attribute types, and each attribute type matches the attribute name of a type of entity attribute corresponding to the current entity type; the first entity feature includes a second context word segmentation sequence and a first entity index; the second context word segmentation sequence is composed of multiple second word segmentation texts sorted; the first entity index is a word segmentation text index in the second context word segmentation sequence; the first attribute classification vector is composed of multiple first word segmentation attribute vectors, and the first word segmentation attribute vectors correspond one-to-one with the second word segmentation texts in the second context word segmentation sequence; the vector length of the first word segmentation attribute vector matches the total number of attribute types of the corresponding single-entity attribute type set and consists of multiple first attribute classification probabilities, and the first attribute classification probabilities correspond one-to-one with the attribute types of the corresponding single-entity attribute type set.

[0168] As Figure 3 shown, the first attribute extraction model in the embodiment of the present invention consists of a third preprocessing unit, Nc parallel base models BM-3 c and a meta-model MM-3; it should be noted that: all the base models BM-3 c are implemented based on the principle of the multi-class logistic regression model, and the model structures of all the base models BM-3 c are the same, but the hyperparameters are different. Nc is the number of all hyperparameter combinations, and 1 ≤ base model index c ≤ Nc; the meta-model MM-3 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameter is one of all hyperparameter combinations of the multi-class logistic regression model.

[0169] In the embodiment of the present invention, the connection relationship of each component of the first attribute extraction model is as follows: the input end of the third preprocessing unit is connected to the model input end of the first attribute extraction model, and the output end is connected to the input ends of each base model BM-3 c ; the output ends of each base model BM-3 c are connected to one input end of the meta-model MM-3; the output end of the meta-model MM-3 is connected to the model output end of the first attribute extraction model.

[0170] In the embodiment of the present invention, the functions of each component of the first attribute extraction model are as follows:

[0171] 1) The third preprocessing unit is used to extract the corresponding second context token sequence and the first entity index from the first entity features input to the model; and perform word embedding encoding on the second context token sequence based on the TF-IDF encoding algorithm to obtain the corresponding fifth encoding vector. The fifth encoding vector is sequentially spliced by a plurality of fifth token encoding vectors, and the fifth token encoding vector corresponds to the second token text in the second context token sequence one by one; and initialize a first entity marker vector with a vector length consistent with the sequence length of the second context token sequence. The first entity marker vector is sequentially sorted by a plurality of first vector data; and set the first vector data whose vector data index in the first entity marker vector matches the first entity index to a preset entity marker, and set all other first vector data whose vector data index in the first entity marker vector does not match the first entity index to a preset non-entity marker; and perform vector splicing on the first vector data and the fifth token encoding vector corresponding to each second token text to obtain a corresponding token splicing vector, and perform vector splicing on all the obtained token splicing vectors to obtain the corresponding third splicing vector and send it to each base model BM-3 c ;

[0172] Here, the entity marker and the non-entity marker in the embodiment of the present invention are two preset marker parameter;

[0173] 2) Each base model BM-3 c is used to perform corresponding entity attribute classification prediction according to all attribute type categories of the corresponding single-entity attribute type set based on the third splicing vector to obtain the corresponding third prediction vector and send it to the meta-model MM-3;

[0174] Among them, the third prediction vector is sequentially spliced by a plurality of second sub-prediction vectors, and the second sub-prediction vector corresponds to the second token text one by one; the vector length of the second sub-prediction vector matches the total number of attribute types of the corresponding single-entity attribute type set and is composed of a plurality of first attribute prediction probabilities, and the first attribute prediction probability corresponds to the attribute type of the corresponding single-entity attribute type set one by one;

[0175] 3) The meta-model MM-3 is used to splice the second sub-prediction vectors corresponding to each second segmented text in the received Nc third prediction vectors, and use the spliced vector as the corresponding sixth segmented encoding vector; and form the corresponding sixth encoding vector from all the obtained sixth segmented encoding vectors; and perform single-entity attribute classification prediction according to all the attribute type categories of the corresponding single-entity attribute type set based on the sixth encoding vector to obtain the corresponding first attribute classification vector.

[0176] It should also be noted that the training methods of the first entity naming model, the first relationship extraction model, and the first attribute extraction model in the embodiments of the present invention are all implemented by using the supervised stacking model training method. Specifically: 1) Before training, a large amount of training data is collected based on the model input data format of the first entity naming model / first relationship extraction model / first attribute extraction model, and corresponding label data is set for each training data based on the model output data format of the first entity naming model / first relationship extraction model / first attribute extraction model. A corresponding training data set is composed of a large number of training-label data pairs, and a model training loss function is set based on a conventional classification prediction loss function, such as a conventional cross-entropy loss function or a cross-entropy loss function with a regularization penalty term, etc.; 2) The training method adopts a two-stage training method, that is, first, each base model BM-1 is independently trained based on the training data set and the model training loss function a / b / c and then the first entity naming model / first relationship extraction model / first attribute extraction model is trained as a whole based on the training data set and the model training loss function.

[0177] In another specific implementation manner of the embodiments of the present invention, the graph processing module 12 is specifically used for: when identifying the first triple subset of medical knowledge from the first sentence sequence of the current preprocessed data based on the preset first entity naming model, the first relationship extraction model, and the first attribute extraction model corresponding to various entity types

[0178] Step D1, taking each first sentence of the first sentence sequence as the corresponding current sentence;

[0179] And use the current sentence as the corresponding first text to input into the first entity naming model for entity classification processing to obtain the corresponding first word segmentation sequence and the first entity classification vector; extract the largest first entity classification probability among the respective first word segmentation classification vectors of the first entity classification vector as the corresponding first probability; identify whether each first probability exceeds a preset first probability threshold. If it exceeds, set a corresponding first word segmentation type as the entity type corresponding to the current first probability. If it does not exceed, set the corresponding first word segmentation type as a non-entity; and use the first word segmentation in the first word segmentation sequence corresponding to each first word segmentation type that is not a non-entity as the corresponding first entity name. When the number of first entity names obtained is greater than 1, combine all the obtained first entity names in pairs to obtain a corresponding set or multiple sets of first entity groups composed of two different first entity names; here, the first probability threshold is a preset threshold parameter;

[0180] And when the number of first entity groups corresponding to the current sentence is greater than 0, use each first entity group corresponding to the current sentence as the corresponding current entity group; use the first word segmentation sequence corresponding to the current entity group as the corresponding first context word segmentation sequence; use one of the two first entity names in the current entity group as the head entity name and the other as the tail entity name; use the word segmentation text indexes of the head and tail entity names in the first context word segmentation sequence as the corresponding head and tail entity indexes; use the first word segmentation types corresponding to the head and tail entity names as the corresponding head and tail entity types; form a corresponding head entity feature from the obtained head entity index and head entity type, and form a corresponding tail entity feature from the obtained tail entity index and tail entity type; form a corresponding first entity pair feature from the obtained first context word segmentation sequence, head entity feature, and tail entity feature; input the obtained first entity pair feature into the first relation extraction model for relation classification processing to obtain the corresponding first relation classification vector; use the largest first relation classification probability in the first relation classification vector as the corresponding second probability; identify whether the second probability exceeds a preset second probability threshold. If it exceeds, use the medical entity association type corresponding to the current second probability as the head-tail entity association relationship corresponding to the current entity group. If it does not exceed, set the head-tail entity association relationship corresponding to the current entity group as empty; and when the head-tail entity association relationship corresponding to the current entity group is not empty, form a corresponding head-tail entity information group from the head entity name, tail entity name, head entity type, tail entity type, and head-tail entity association relationship corresponding to the current entity group; here, the second probability threshold is a preset threshold parameter;

[0181] And use the first word segmentation sequence corresponding to the current sentence as a corresponding second context word segmentation sequence; when the number of the first entity names corresponding to the current sentence is greater than 1, perform a round of traversal on all the first entity names corresponding to the current sentence; during this round of traversal, use the currently traversed first entity name as the corresponding current entity name; use the first word segmentation type corresponding to the current entity name as the corresponding first entity type as the corresponding current entity type; use the first attribute extraction model corresponding to the current entity type as the corresponding current attribute extraction model; use the word segmentation text index of the current entity name on the current second context word segmentation sequence as a corresponding first entity index; and form a corresponding first entity feature from the second context word segmentation sequence corresponding to the current entity name and the first entity index, input it into the current attribute extraction model for corresponding attribute classification processing to obtain a corresponding first attribute classification vector; extract the maximum first attribute classification probability among the first word segmentation attribute vectors of the first attribute classification vector obtained this time as the corresponding third probability; record the third probabilities whose respective probability values exceed the preset third probability threshold as the corresponding fourth probabilities; when the total number of the fourth probabilities obtained this time is not zero, form a corresponding first attribute group from the attribute types and first word segmentations corresponding to the respective fourth probabilities as a group of corresponding attribute names and attribute values; after the end of this round of traversal, perform deduplication and merging processing on all the first attribute groups corresponding to the same first entity name to obtain a corresponding first attribute group set; here, the third probability threshold is a preset threshold parameter;

[0182] Step D2, and perform deduplication processing on all the head-tail entity information groups obtained based on the first sentence sequence; and perform attribute group deduplication and merging processing on one or more first attribute group sets corresponding to the same first entity name in all the first attribute group sets obtained based on the first sentence sequence to obtain a corresponding first merged attribute group set;

[0183] Step D3, and use each deduplicated head-tail entity information group as the corresponding current entity information group one by one; use the head entity name and head entity type of the current entity information group as the entity name and entity type of a corresponding entity A, and construct the entity attribute set corresponding to the current entity A from the first merged attribute group set corresponding to the current head entity name; use the tail entity name and tail entity type of the current entity information group as the entity name and entity type of a corresponding entity B, and construct the entity attribute set corresponding to the current entity B from the first merged attribute group set corresponding to the current tail entity name; use the head-tail entity association relationship of the current entity information group as the entity relationship R corresponding to the current entities A and B; and form a corresponding first triple from the entity A, entity relationship R, and entity B corresponding to the current entity information group;

[0184] Step D4, and form a corresponding first triple subset from all the first triples obtained this time.

[0185] In another specific implementation manner of the embodiment of the present invention, the second entity naming model is specifically used to perform entity classification processing on each word segment in the second text input to the model according to a preset entity classification set and output a corresponding second entity classification vector. Here, the second text in the embodiment of the present invention is a single-segment text, which is composed of one or more single-sentence texts; the second entity classification vector is composed of multiple second word segment classification vectors, and the second word segment classification vectors correspond one by one to the word segments in the second text; the second word segment classification vector is composed of multiple second entity classification probabilities, and the second entity classification probabilities correspond one by one to the entity types.

[0186] As Figure 4 shown, the second entity naming model of the embodiment of the present invention is composed of a fourth preprocessing unit, a first BERT model, a first linear network, and a first Softmax layer; it should be noted that: the first BERT model is one of the basic BERT model, BioBERT, and Clinical BERT that have completed pre-training; the first linear network is implemented based on one or more fully connected layers.

[0187] The connection relationship of the components of the second entity naming model in the embodiment of the present invention is: the input end of the fourth preprocessing unit is connected to the model input end of the second entity naming model, and the output end is connected to the input end of the first BERT model; the output end of the first BERT model is connected to the input end of the first linear network; the output end of the first linear network is connected to the input end of the first Softmax layer; the output end of the first Softmax layer is connected to the model output end of the second entity naming model.

[0188] The functions of the components of the second entity naming model in the embodiment of the present invention are:

[0189] 1) The fourth preprocessing unit is used to perform word segmentation processing on the second text to obtain a corresponding second word segment sequence; and perform embedding encoding processing on the second word segment sequence according to the embedding encoding rule of the model input vector of the BERT model to obtain a corresponding first input tensor and send it to the first BERT model;

[0190] Among them, the second word segment sequence is composed of multiple word segments sorted in order; the first input tensor is composed of multiple embedding encoding vectors, and the embedding encoding vectors of the first input tensor correspond one by one to the word segments of the second word segment sequence;

[0191] 2) The first BERT model is used to encode the first input tensor to obtain a corresponding first encoded tensor and send it to the first linear network;

[0192] Among them, the first encoded tensor is composed of multiple sub-encoded tensors, and the sub-encoded tensors of the first encoded tensor correspond one-to-one with the embedded encoded vectors of the first input tensor;

[0193] 3) The first linear network is used to perform feature extraction according to the first encoded tensor to obtain the corresponding first feature tensor and send it to the first Softmax layer;

[0194] Among them, the first feature tensor is composed of multiple feature vectors, and the feature vectors of the first feature tensor correspond one-to-one with the sub-encoded tensors of the first encoded tensor;

[0195] 4) The first Softmax layer is used to perform entity classification prediction processing according to the first feature tensor to obtain the corresponding second entity classification vector.

[0196] In another specific implementation manner of the embodiment of the present invention, the second relation extraction model is specifically used to perform entity pair relation classification processing according to a preset association relation set, the second word segmentation sequence input by the model, and the second entity classification vector, and output the corresponding first word pair classification matrix. Here, the second word segmentation sequence is the word segmentation sequence generated by the fourth preprocessing unit of the second entity naming model, and the second entity classification vector is the entity classification vector output by the second entity naming model; each row or each column of the first word pair classification matrix output by the model corresponds one-to-one with the second word segmentation of the second word segmentation sequence; each matrix unit not on the diagonal in the first word pair classification matrix is the second relation classification vector of a word pair, and each matrix unit on the diagonal of the first word pair classification matrix is an invalid classification vector with a vector length consistent with the vector length of the second relation classification vector and all vector data being zero or all negative values; the second relation classification vector is composed of multiple second relation classification probabilities, and each second relation classification probability corresponds to a type of medical entity association type.

[0197] As Figure 4 shown, the second relation extraction model of the embodiment of the present invention is composed of a fifth preprocessing unit, a second BERT model, a second linear network, and a second Softmax layer; it should be noted that: the model structures of the first and second BERT models are the same and the model parameters are the same; the second linear network is implemented based on one or more fully connected layers.

[0198] The connection relationships of the components of the second relation extraction model in the embodiment of the present invention are: the input end of the fifth preprocessing unit is connected to the model input end of the second relation extraction model, and the output end is connected to the input end of the second BERT model; the output end of the second BERT model is connected to the input end of the second linear network; the output end of the second linear network is connected to the input end of the second Softmax layer; the output end of the second Softmax layer is connected to the model output end of the second relation extraction model.

[0199] The functions of the components of the second relation extraction model in the embodiments of the present invention are as follows:

[0200] 1) The fifth preprocessing unit performs embedding encoding processing on the second tokenized sequence and the second entity classification vector of the model input according to the embedding encoding rule of the model input vector of the BERT model to obtain the corresponding second input tensor, and sends it to the second BERT model;

[0201] Among them, the second input tensor is composed of multiple embedding encoding vectors, and the embedding encoding vectors of the second input tensor correspond one by one to the second tokens of the second tokenized sequence;

[0202] 2) The second BERT model is used to encode the second input tensor to obtain the corresponding second encoded tensor and send it to the second linear network;

[0203] Among them, the second encoded tensor is composed of multiple sub-encoded tensors, and the sub-encoded tensors of the second encoded tensor correspond one by one to the embedding encoding vectors of the second input tensor;

[0204] 3) The second linear network is used to perform pairwise permutation and combination on the sub-encoded tensors in the second encoded tensor to obtain multiple sub-encoded tensor combinations; and perform tensor splicing according to the permutation order of the two sub-encoded tensors in each sub-encoded tensor combination to obtain the corresponding combined tensor; and perform feature extraction on each combined tensor to obtain the corresponding combined feature vector; and form a corresponding first token pair feature matrix based on all the obtained combined feature vectors and send it to the second Softmax layer;

[0205] Among them, each row or each column of the first token pair feature matrix corresponds one by one to the sub-encoded tensors of the second encoded tensor; each matrix unit not on the diagonal in the first token pair feature matrix is a corresponding combined feature vector, and each matrix unit on the diagonal of the first token pair feature matrix is an invalid feature vector with a vector length consistent with that of the combined feature vector and all vector data being zero or all negative values;

[0206] 4) The second Softmax layer is used to perform entity relationship classification prediction according to the combined feature vectors of the first token pair feature matrix to obtain the corresponding second relationship classification vector; and form the corresponding first token pair classification matrix from all the obtained second relationship classification vectors.

[0207] The number of models of the second attribute extraction model in the embodiments of the present invention is equal to the number of entity types, and the second attribute extraction model corresponds to the entity types one by one. In another specific implementation manner of the embodiments of the present invention, each second attribute extraction model is specifically configured to perform attribute classification processing according to all attribute type categories of the corresponding single-entity attribute type set, and based on the second word segmentation sequence and the first entity position input to the model, and output the corresponding second attribute classification vector. Here, the second word segmentation sequence is the word segmentation sequence generated by the fourth preprocessing unit of the second entity naming model, and the first entity position is the word segmentation index corresponding to a second word in the second word segmentation sequence; the second attribute classification vector output by the model is composed of multiple second word segmentation attribute vectors, and the second word segmentation attribute vectors correspond to the second words in the second word segmentation sequence one by one; the vector length of the second word segmentation attribute vector matches the total number of attribute types of the corresponding single-entity attribute type set and is composed of multiple second attribute classification probabilities, and the second attribute classification probabilities correspond to the attribute types of the corresponding single-entity attribute type set one by one.

[0208] As Figure 4 shown, the second attribute extraction model in the embodiments of the present invention is composed of a sixth preprocessing unit, a third BERT model, a third linear network, and a third Softmax layer; it should be noted that: First, the model structures of the third BERT models are the same and the model parameters are the same; the third linear network is implemented based on one or more fully connected layers.

[0209] The connection relationship of each component of the second attribute extraction model in the embodiments of the present invention is: the input end of the sixth preprocessing unit is connected to the model input end of the second attribute extraction model, and the output end is connected to the input end of the third BERT model; the output end of the third BERT model is connected to the input end of the third linear network; the output end of the third linear network is connected to the input end of the third Softmax layer; the output end of the third Softmax layer is connected to the model output end of the second attribute extraction model.

[0210] The functions of each component of the second attribute extraction model in the embodiments of the present invention are:

[0211] 1) The sixth preprocessing unit initializes a second entity marker vector with a vector length consistent with the sequence length of the second word segmentation sequence input to the model. The second entity marker vector is composed of multiple second vector data sorted in sequence; and sets the second vector data whose vector data index in the second entity marker vector matches the first entity position to a preset entity marker, and sets all other second vector data whose vector data index in the second entity marker vector does not match the first entity position to a preset non-entity marker; and performs embedding encoding processing according to the second word segmentation sequence and the second entity marker vector according to the embedding encoding rule of the model input vector of the BERT model to obtain the corresponding third input tensor and send it to the third BERT model;

[0212] Among them, the third input tensor is composed of multiple embedded encoding vectors, and the embedded encoding vectors of the third input tensor correspond one-to-one with the second word segments of the second word segmentation sequence;

[0213] 2) The third BERT model is used to encode the third input tensor to obtain the corresponding third encoded tensor and send it to the third linear network;

[0214] Among them, the third encoded tensor is composed of multiple sub-encoded tensors, and the sub-encoded tensors of the third encoded tensor correspond one-to-one with the embedded encoding vectors of the third input tensor;

[0215] 3) The third linear network is used to perform feature extraction based on the third encoded tensor to obtain the corresponding second feature tensor and send it to the third Softmax layer;

[0216] Among them, the second feature tensor is composed of multiple feature vectors, and the feature vectors of the second feature tensor correspond one-to-one with the sub-encoded tensors of the third encoded tensor;

[0217] 4) The third Softmax layer is used to perform corresponding single-entity attribute classification prediction processing on each feature vector of the second feature tensor according to all attribute type categories of the corresponding single-entity attribute type set to obtain the corresponding second word segment attribute vector; and all the obtained second word segment attribute vectors form the corresponding second attribute classification vector.

[0218] It should also be noted that the training methods of the second entity naming model, the second relation extraction model, and the second attribute extraction model in the embodiments of the present invention can be implemented with reference to the NLP task training method of the pre-trained BERT model, and are roughly similar to the traditional supervised training method: 1) Before training, a large amount of training data is collected based on the model input data format of the second entity naming model / second relation extraction model / second attribute extraction model, and corresponding label data is set for each training data based on the model output data format of the second entity naming model / second relation extraction model / second attribute extraction model. A large number of training-label data pairs constitute the corresponding training dataset, and the model training loss function is set based on the conventional classification prediction loss function, such as the conventional cross-entropy loss function or the cross-entropy loss function with a regularization penalty term, etc.; 2) Then, the second entity naming model / second relation extraction model / second attribute extraction model is trained as a whole based on the training dataset and the model training loss function.

[0219] Another specific implementation manner of the embodiment of the present invention is that the graph processing module 12 is specifically used for when identifying medical knowledge triples in the first sentence sequence of the current preprocessed data based on the preset second entity naming model, second relation extraction model, and second attribute extraction model corresponding to various entity types to obtain the corresponding first triple subset:

[0220] Step E1, perform a sharding and partitioning process on the first sentence sequence according to a preset sharding and partitioning rule to obtain one or more corresponding first fragment sentence sequences;

[0221] Among them, the first fragment sentence sequence consists of one or more first sentences;

[0222] Here, the sharding and partitioning rule is a preset sliding sharding rule for long text sharding;

[0223] Step E2, and take each first fragment sentence sequence as the corresponding current fragment sentence sequence one by one;

[0224] And perform sequential splicing on all the first sentences of the current fragment sentence sequence to obtain a corresponding current fragment splicing text; and use the current fragment splicing text as the corresponding second text to input into the second entity naming model for entity classification processing to obtain the corresponding second word segmentation sequence and second entity classification vector; and extract the largest second entity classification probability among the second word segmentation classification vectors of the second entity classification vector as the corresponding fifth probability; and identify whether each fifth probability exceeds a preset fourth probability threshold. If it exceeds, set a corresponding second word segmentation type as the entity type corresponding to the current fifth probability. If it does not exceed, set the corresponding second word segmentation type as non-entity; and use the second word segments in the second word segmentation sequence corresponding to each second word segmentation type that is not non-entity as the corresponding second entity names, and when the number of second entity names is greater than 1, perform pairwise combination on all the obtained second entity names to obtain a corresponding group or multiple groups of second entity groups composed of two different second entity names; here, the fourth probability threshold is a preset threshold parameter;

[0225] When the number of the second entity groups corresponding to the current segment sentence sequence is greater than 0, input the second word segmentation sequence and the second entity classification vector corresponding to the current segment sentence sequence into the second relation extraction model for entity pair relation classification processing to obtain the corresponding first word pair classification matrix; perform a round of traversal on all the second entity groups corresponding to the current segment sentence sequence; during this round of traversal, use the currently traversed second entity group as the corresponding current entity group; use one of the two word segmentation indexes of the two second entity names of the current entity group in the current second word segmentation sequence as the row index and the other as the column index, and use the second relation classification vector corresponding to the matrix cell whose row label and column label in the current first word pair classification matrix match the corresponding row and column indexes respectively as the corresponding current relation classification vector; use the largest second relation classification probability in the current relation classification vector as the corresponding sixth probability; identify whether the sixth probability exceeds the preset fifth probability threshold. If it exceeds, use the medical entity association type corresponding to the current sixth probability as the head-tail entity association relationship corresponding to the current entity group. If it does not exceed, set the head-tail entity association relationship corresponding to the current entity group to be empty; when the head-tail entity association relationship corresponding to the current entity group is not empty, use the row and column indexes as the corresponding head and tail entity positions, use the second entity name and the second word segmentation type corresponding to the head entity position as the corresponding head entity name and head entity type, use the second entity name and the second word segmentation type corresponding to the tail entity position as the corresponding tail entity name and tail entity type, and form a corresponding head-tail entity information group from the head entity position, tail entity position, head entity name, tail entity name, head entity type, tail entity type, and head-tail entity association relationship corresponding to the current entity group; at the end of this round of traversal, perform duplicate removal processing on all the head-tail entity information groups obtained from this round of traversal; here, the fifth probability threshold is a preset threshold parameter;

[0226] When the number of the head-tail entity information groups corresponding to the current segment sentence sequence is greater than 0, perform a round of traversal on all the head-tail entity information groups corresponding to the current segment sentence sequence; during this round of traversal, take the currently traversed head-tail entity information group as the corresponding current information group; take the head entity position of the current information group as a corresponding first entity position, take the second attribute extraction model corresponding to the head entity type of the current information group as the corresponding current attribute extraction model, and input the second word segmentation sequence corresponding to the current segment sentence sequence and the current first entity position into the current attribute extraction model to perform attribute classification processing to obtain a corresponding second attribute classification vector denoted as the corresponding head entity attribute classification vector; take the tail entity position of the current information group as a new first entity position, take the second attribute extraction model corresponding to the tail entity type of the current information group as the new current attribute extraction model, and input the second word segmentation sequence corresponding to the current segment sentence sequence and the new first entity position into the new current attribute extraction model to perform attribute classification processing to obtain a new second attribute classification vector denoted as the corresponding tail entity attribute classification vector; sequentially take the head and tail entity attribute classification vectors as the corresponding current vectors, extract the largest second attribute classification probability in each second word segmentation attribute vector of the current vector as the corresponding seventh probability, and denote the seventh probability whose probability value exceeds the preset sixth probability threshold as the corresponding eighth probability, and when the total number of the eighth probabilities is not zero, form a corresponding second attribute group with the attribute type and the second word segmentation corresponding to each eighth probability as a group of corresponding attribute names and attribute values; after the end of this round of traversal, perform deduplication and merging processing on all the second attribute groups corresponding to the same second entity name in all the second attribute groups corresponding to the current segment sentence sequence to obtain a corresponding second attribute group set; here, the sixth probability threshold is a preset threshold parameter;

[0227] Step E3, perform deduplication processing on all the head-tail entity information groups obtained based on the first sentence sequence; in all the second attribute group sets obtained based on the first sentence sequence, perform attribute group deduplication and merging processing on all the second attribute group sets corresponding to the same second entity name to obtain the corresponding second merged attribute group set;

[0228] Step E4, and take each head-tail entity information group after duplicate removal as the corresponding current information group one by one; use the head entity name and head entity type of the current information group as the entity name and entity type of a corresponding entity A, and construct the entity attribute set corresponding to the current entity A from the second merged attribute group set corresponding to the current head entity name; use the tail entity name and tail entity type of the current information group as the entity name and entity type of a corresponding entity B, and construct the entity attribute set corresponding to the current entity B from the second merged attribute group set corresponding to the current tail entity name; use the head-tail entity association relationship of the current information group as the entity relationship R corresponding to the current entities A and B; and form a corresponding first triple from the entity A, entity relationship R, and entity B corresponding to the current information group.

[0229] Step E5, and form a corresponding first triple subset from all the first triples obtained this time.

[0230] In the embodiments of the present invention, the number of image attribute prediction models is equal to the number of medical examination images, and the image attribute prediction models and the medical examination images are in one-to-one correspondence. In another specific implementation manner of the embodiments of the present invention, each image attribute prediction model is specifically configured to perform image attribute prediction on the first medical image input to the model according to the attribute analysis requirements of the corresponding first analysis attribute set and output the corresponding first image attribute vector. Here, the first medical image in the embodiments of the present invention is a type of medical examination image; the first image attribute vector includes multiple first attribute prediction data; and the first attribute prediction data corresponds one-to-one to the first attribute names in the first analysis attribute set.

[0231] As Figure 5 shown, the image attribute prediction model in the embodiments of the present invention is composed of a first feature extraction network and a first prediction network; it should be noted that: the first feature extraction network is implemented based on a type of image encoder model; the first prediction network is implemented based on a type of non-linear regression prediction model; the image encoder model at least includes a CNN network and a residual network, that is, the image encoder model in the embodiments of the present invention can be implemented at least based on a CNN network or a residual network; the non-linear regression prediction model at least includes an MLP model, a decision tree model, an SVR model, and a random forest model, that is, the non-linear regression prediction model in the embodiments of the present invention can be implemented at least based on an MLP model, a decision tree model, an SVR model, or a random forest model.

[0232] The connection relationship of each component of the image attribute prediction model in the embodiments of the present invention is: the input end of the first feature extraction network is connected to the input end of the image attribute prediction model, and the output end is connected to the input end of the first prediction network; the output end of the first prediction network is connected to the output end of the image attribute prediction model.

[0233] The functions of each component of the image attribute prediction model in the embodiments of the present invention are:

[0234] 1) The first feature extraction network is used to perform feature extraction processing on the radiomics features of the first medical image input to the model to obtain the corresponding first image feature tensor and send it to the first prediction network;

[0235] 2) The first prediction network is used to predict the attribute values corresponding to each first attribute name specified in the first analysis attribute set according to the first image feature tensor to obtain the corresponding first attribute prediction data; and all the obtained first attribute prediction data are used to form the corresponding first image attribute vector.

[0236] In the embodiment of the present invention, the number of audio attribute prediction models is the same as the number of medical examination audios, and the audio attribute prediction models correspond to the medical examination audios one by one. In another specific implementation manner of the embodiment of the present invention, each audio attribute prediction model is used to perform audio attribute prediction on the first medical audio input to the model according to the attribute analysis requirements of the corresponding second analysis attribute set and output the corresponding first audio attribute vector. Here, the first medical audio in the embodiment of the present invention is a type of medical examination audio; the first audio attribute vector includes multiple second attribute prediction data; the second attribute prediction data corresponds to the second attribute names in the second analysis attribute set one by one.

[0237] As Figure 5 shown, the audio attribute prediction model in the embodiment of the present invention is composed of a second feature extraction network and a second prediction network; it should be noted that: the second feature extraction network is implemented based on a type of audio encoder model; the second prediction network is implemented based on a type of non-linear regression prediction model; the audio encoder model at least includes a CNN network, an RNN network, a residual network, an LSTM model, a Transformer model, that is, the audio encoder model in the embodiment of the present invention can be implemented at least based on a CNN network, an RNN network, a residual network, an LSTM model or a Transformer model.

[0238] The connection relationship of each component of the audio attribute prediction model in the embodiment of the present invention is: the input end of the second feature extraction network is connected to the input end of the audio attribute prediction model, and the output end is connected to the input end of the second prediction network; the output end of the second prediction network is connected to the output end of the audio attribute prediction model.

[0239] The functions of each component of the audio attribute prediction model in the embodiment of the present invention are:

[0240] 1) The second feature extraction network is used to perform feature extraction processing on some or all of the Mel cepstrum coefficient features, linear prediction coefficient features, short-time energy features, short-time zero-crossing rate features and frequency domain features of the first medical audio input to the model to obtain the corresponding first audio feature tensor and send it to the second prediction network;

[0241] 2) The second prediction network is used to predict the attribute values corresponding to each second attribute name specified by the second analysis attribute set based on the first audio feature tensor to obtain corresponding second attribute prediction data; and all the obtained second attribute prediction data are used to form a corresponding first audio attribute vector.

[0242] In the embodiment of the present invention, the number of video attribute prediction models is equal to the number of medical examination videos, and the video attribute prediction models correspond to the medical examination videos one by one. In another specific implementation manner of the embodiment of the present invention, each video attribute prediction model is used to perform video attribute prediction on the input first medical video according to the attribute analysis requirements of the corresponding third analysis attribute set and output a corresponding first video attribute vector. Here, the first medical video in the embodiment of the present invention is a type of medical examination video; the first video attribute vector includes multiple third attribute prediction data; the third attribute prediction data corresponds to the third attribute names in the third analysis attribute set one by one.

[0243] As Figure 5 shown, the video attribute prediction model in the embodiment of the present invention is composed of a third feature extraction network and a third prediction network; it should be noted that: the third feature extraction network is implemented based on a type of video encoder model; the third prediction network is implemented based on a type of non-linear regression prediction model; the video encoder model at least includes a 3D CNN network, a TCN network, an RNN network, an LSTM model, and a Transformer model, that is, the audio encoder model in the embodiment of the present invention can be implemented at least based on a 3D CNN network, a TCN network, an RNN network, an LSTM model, or a Transformer model.

[0244] The connection relationship of each component of the video attribute prediction model in the embodiment of the present invention is: the input end of the third feature extraction network is connected to the input end of the video attribute prediction model, and the output end is connected to the input end of the third prediction network; the output end of the third prediction network is connected to the output end of the video attribute prediction model.

[0245] The functions of each component of the video attribute prediction model in the embodiment of the present invention are:

[0246] 1) The third feature extraction network is used to perform feature extraction processing on some or all of the temporal features of physiological activities, the temporal features of physiological indicators, and the temporal deformation of physiological structures in the input first medical video of the model to obtain a corresponding first video feature tensor and send it to the third prediction network;

[0247] 2) The third prediction network is used to predict the attribute values corresponding to each third attribute name specified by the third analysis attribute set based on the first video feature tensor to obtain corresponding third attribute prediction data; and all the obtained third attribute prediction data are used to form a corresponding first video attribute vector.

[0248] It should also be noted that the training methods of the image attribute prediction model, audio attribute prediction model, and video attribute prediction model in the embodiments of the present invention are generally similar to the traditional supervised training methods: 1) Before training, a large amount of training data is collected based on the model input data format of the image attribute prediction model / audio attribute prediction model / video attribute prediction model, and corresponding label data is set for each training data based on the model output data format of the image attribute prediction model / audio attribute prediction model / video attribute prediction model. A corresponding training data set is composed of a large number of training-label data pairs, and a model training loss function is set based on a conventional regression prediction loss function, such as a conventional L1 loss function, L2 loss function, cross-entropy loss function, etc.; 2) Then, the image attribute prediction model / audio attribute prediction model / video attribute prediction model is trained as a whole based on the training data set and the model training loss function.

[0249] In another specific implementation manner of the embodiment of the present invention, when the atlas processing module 12 is specifically used to perform medical knowledge triple recognition on the first description and the first image of the current preprocessed data and the preset image attribute prediction model corresponding to various medical examination images to obtain a corresponding first triple subset:

[0250] Step F1, extract the corresponding disease information, examination category information, and examination description information from the first description of the current preprocessed data;

[0251] Step F2, and extract the corresponding disease name from the disease information as a corresponding entity name; and use the disease entity type corresponding to the current disease name as a corresponding entity type; and set an empty entity attribute set for the current disease name; and form a corresponding entity A from the entity name, entity type, and entity attribute set corresponding to the current disease name.

[0252] Step F3, extract the corresponding inspection content theme from the inspection description information as a new entity name; and use the medical inspection type entity type corresponding to the inspection category information as a new entity type; and confirm the type of the medical inspection image corresponding to the first image according to the inspection category information and the inspection description information, and select the corresponding first analysis attribute set and image attribute prediction model as the corresponding current analysis attribute set and current image attribute prediction model based on the confirmation result; and use the first image of the current preprocessed data as the corresponding first medical image to input the current image attribute prediction model for image attribute prediction to obtain the corresponding first image attribute vector; and use each first attribute name of the current analysis attribute set and the first attribute prediction data corresponding to the current first attribute name in the first image attribute vector as a group of corresponding attribute name and attribute value to form a corresponding entity attribute; and form a new entity attribute set from all the obtained entity attributes; and form a corresponding entity B from the new entity name, entity type and entity attribute set;

[0253] Step F4, set a corresponding entity relationship R based on the association relationship between the disease entity - medical inspection type entity;

[0254] Step F5, and form a corresponding first triple from the entity A, entity relationship R and entity B obtained this time; and form a corresponding first triple subset from the first triple obtained this time.

[0255] Another specific implementation manner of the embodiment of the present invention, the atlas processing module 12 is specifically used for when identifying the medical knowledge triple to obtain the corresponding first triple subset based on the first description and the first audio of the current preprocessed data and the preset audio attribute prediction model corresponding to various medical inspection audios:

[0256] Step G1, extract the corresponding disease information, inspection category information, and inspection description information from the first description of the current preprocessed data;

[0257] Step G2, extract the corresponding disease name from the disease information as a corresponding entity name; and use the disease entity type corresponding to the current disease name as a corresponding entity type; and set an empty entity attribute set for the current disease name; and form a corresponding entity A from the entity name, entity type and entity attribute set corresponding to the current disease name;

[0258] Step G3, extract the corresponding inspection content theme from the inspection description information as a new entity name; and use the medical inspection type entity type corresponding to the inspection category information as a new entity type; and confirm the type of the medical inspection audio corresponding to the first audio according to the inspection category information and the inspection description information, and select the corresponding second analysis attribute set and audio attribute prediction model as the corresponding current analysis attribute set and current audio attribute prediction model based on the confirmation result; and use the first audio of the current preprocessed data as the corresponding first medical audio to input the current audio attribute prediction model for audio attribute prediction to obtain the corresponding first audio attribute vector; and use each second attribute name of the current analysis attribute set and the second attribute prediction data corresponding to the current second attribute name in the first audio attribute vector as a group of corresponding attribute names and attribute values to form a corresponding entity attribute; and form a new entity attribute set from all the obtained entity attributes; and form a corresponding entity B from the new entity name, entity type and entity attribute set;

[0259] Step G4, set a corresponding entity relationship R based on the association relationship between the disease entity - medical inspection type entity;

[0260] Step G5, and form a corresponding first triple from the entity A, entity relationship R and entity B obtained this time; and form a corresponding first triple subset from the first triple obtained this time.

[0261] Another specific implementation manner of the embodiment of the present invention, the atlas processing module 12 is specifically used for when identifying medical knowledge triples based on the first description and the first video of the current preprocessed data and the preset video attribute prediction models corresponding to various medical inspection videos to obtain the corresponding first triple subset:

[0262] Step H1, extract the corresponding disease information, inspection category information, and inspection description information from the first description of the current preprocessed data;

[0263] Step H2, extract the corresponding disease name from the disease information as a corresponding entity name; and use the disease entity type corresponding to the current disease name as a corresponding entity type; and set an empty entity attribute set for the current disease name; and form a corresponding entity A from the entity name, entity type and entity attribute set corresponding to the current disease name;

[0264] Step H3, extract the corresponding inspection content theme from the inspection description information as a new entity name; and use the medical inspection type entity type corresponding to the inspection category information as a new entity type; and confirm the type of the medical inspection video corresponding to the first video according to the inspection category information and the inspection description information, and select the corresponding third analysis attribute set and video attribute prediction model as the corresponding current analysis attribute set and current video attribute prediction model based on the confirmation result; and use the first video of the current preprocessed data as the corresponding first medical video to input the current video attribute prediction model for video attribute prediction to obtain the corresponding first video attribute vector; and use each third attribute name in the current analysis attribute set and the third attribute prediction data corresponding to the current third attribute name in the first video attribute vector as a group of corresponding attribute names and attribute values to form a corresponding entity attribute; and form a new entity attribute set from all the obtained entity attributes; and form a corresponding entity B from the new entity name, entity type and entity attribute set;

[0265] Step H4, set a corresponding entity relationship R based on the association relationship between the disease entity - medical inspection type entity;

[0266] Step H5, form a corresponding first triple from the entity A, entity relationship R and entity B obtained this time; and form a corresponding first triple subset from the first triple obtained this time.

[0267] In another specific implementation manner of the embodiment of the present invention, the atlas processing module 12 is specifically used for when updating the medical knowledge atlas according to the first triple set and the atlas database 13:

[0268] Step I1, perform a round of traversal on all the first triples in the first triple set; and during this round of traversal, use the currently traversed first triple as the corresponding current triple; and use the entity names, entity types and entity attribute sets of the entities A and B of the current triple as the corresponding names A and B, types A and B, and attribute sets A and B; and use the entity relationship R of the current triple as a corresponding A - B association relationship; and use the first node record tables corresponding to types A and B in the atlas database 13 as the corresponding node record tables A and B; and use the first node record in the node name field of the node record tables A and B that matches the corresponding names A and B as the corresponding matching node records A and B; and use the first edge record in the first edge record table of the atlas database 13 whose association type field matches the A - B association relationship, and whose first node identification field has a mapping relationship with the node identification field of the matching node record A, and whose tail node identification field has a mapping relationship with the node identification field of the matching node record B as the corresponding matching edge record C;

[0269] And identify the matching node records A and B;

[0270] If both the matching node records A and B are not empty, then take the matching node records A and B as the corresponding current matching node records in sequence, and perform old node update processing based on the current matching node records and the current triple. When the matching edge record C is empty, perform single-edge addition processing based on the matching node records A and B and the A-B association relationship;

[0271] If one of the matching node records A and B is empty, then take the non-empty matching node record A or B as the corresponding current matching node record, and take the entity A or B corresponding to the empty matching node record A or B as the corresponding current newly added entity. Perform old node update processing based on the current matching node record and the current triple, and perform single-node addition processing based on the current newly added entity to obtain the corresponding newly added node record for this time. Then, form a new pair of matching node records A and B from the current matching node record and the newly added node record for this time, and perform single-edge addition processing based on the new matching node records A and B and the A-B association relationship;

[0272] If both the matching node records A and B are empty, then take the entities A and B of the current triple as the corresponding current newly added entities in sequence, and perform single-node addition processing based on the current newly added entities to obtain the corresponding newly added node record for this time. Then, form a new pair of matching node records A and B from the two newly added node records corresponding to the entities A and B, and perform single-edge addition processing based on the new matching node records A and B and the A-B association relationship;

[0273] Step I2. After the current round of traversal of all the first triples in the first triple set, construct the latest first node set based on all the first node record tables and first attribute record tables on the graph database 13; construct the latest first edge set based on the first edge record table on the graph database 13; and form the latest medical knowledge graph from the obtained first node set and first edge set.

[0274] In another specific implementation manner of the embodiment of the present invention, the graph processing module 12 is specifically configured to, when performing old node update processing based on the current matching node record and the current triple:

[0275] Step J1. Take the attribute set A or B corresponding to the current matching node record in the current triple as the corresponding current attribute set; and take the first node feature vector record in the first node feature vector record table of the graph database 13 whose node vector identification field has a mapping relationship with the node vector field of the current matching node record as the corresponding current node feature record;

[0276] Step J2, and perform a round of traversal on all entity attributes of the current attribute set; during this round of traversal, use the currently traversed entity attribute as the corresponding current entity attribute; and use the attribute field corresponding to the attribute name of the current entity attribute in the current matching node record as the corresponding current attribute field;

[0277] And identify whether the current attribute field is an empty field;

[0278] If the current attribute field is not an empty field, use the first attribute record in the first attribute record table of the graph database 13 whose attribute identifier field has a mapping relationship with the current attribute field as the corresponding current attribute record, and use the attribute value of the latest sequence element in the attribute value sequence field of the current attribute record as the corresponding old attribute value, and confirm whether the attribute value of the current entity attribute matches the old attribute value. If the confirmation shows a mismatch, add a new sequence element to the attribute value sequence stored in the current attribute value sequence field as the corresponding current new element, and set the addition time and attribute value of the current new element based on the current time and the attribute value of the current entity attribute;

[0279] If the current attribute field is an empty field, add a new first attribute record in the first attribute record table and record it as the corresponding current new attribute record, assign a unique identifier to the current new attribute record as the corresponding current attribute record identifier, set the attribute identifier field and attribute name field of the current new attribute record based on the current attribute record identifier and the attribute name of the current entity attribute, initialize an empty sequence in the attribute value sequence field of the current new attribute record as the corresponding current attribute value sequence, add a new sequence element to the current attribute value sequence as the corresponding current new element, and set the addition time and attribute value of the current new element based on the current time and the attribute value of the current entity attribute. And establish a corresponding foreign key-primary key mapping relationship between the current attribute field of the current matching node record and the attribute identifier field of the current new attribute record by setting the current attribute field of the current matching node record to the current attribute record identifier;

[0280] Step J3. After the current round of traversal of all entity attributes in the current attribute set, use the node name field recorded by the current matching node as the corresponding current name data; use the entity type corresponding to the current matching node as the corresponding current type data; perform a round of traversal on all attribute fields recorded by the current matching node; during this round of traversal, use the currently traversed attribute field as the corresponding current attribute field; identify whether the current attribute field is empty; if so, form a corresponding first attribute key-value pair consisting of the attribute name corresponding to the current attribute field and an empty attribute value; if not, form a corresponding first attribute key-value pair consisting of the attribute name corresponding to the current attribute field and the attribute value of the sequence element with the latest addition time in the attribute value sequence field mapped by the current attribute field; at the end of this round of traversal, form a corresponding current attribute sequence by sorting all the obtained first attribute key-value pairs in order; form a corresponding current node data vector from the current type data, current name data, and current attribute sequence corresponding to the current matching node; encode the current node data vector based on a preset node feature encoding algorithm to obtain a corresponding current node feature encoding vector; and reset the node feature vector field of the current node feature record based on the current node feature encoding vector.

[0281] It should be noted here that the node feature encoding algorithm in the embodiments of the present invention is a preset encoding algorithm, which can be a conventional word vector embedding encoding algorithm (such as Word2vec embedding encoding algorithm, etc.), or an embedding translation encoding algorithm for embedding and translating the node features of a knowledge graph (such as RotatE algorithm, TransE algorithm, etc.), or other encoding algorithms customized based on application requirements. If the node feature encoding algorithm in the embodiments of the present invention is a conventional word vector embedding encoding algorithm, the corresponding current node feature encoding vector can be directly obtained by encoding the previous node data vector based on the corresponding word vector embedding encoding algorithm; if it is an embedding translation encoding algorithm (such as RotatE algorithm, TransE algorithm, etc.), it is necessary to first obtain all node data vectors and edge data vectors of the current medical knowledge graph to form a corresponding node data vector set and edge data vector set, and then based on the current embedding translation encoding algorithm (such as RotatE algorithm, TransE algorithm, etc.), perform an overall translation on the node and edge feature encoding vectors of all nodes and all edges of the current medical knowledge graph according to the node data vector set and edge data vector set to obtain a corresponding node feature encoding vector set and edge feature encoding vector set, and then use the node feature encoding vector in the node feature encoding vector set corresponding to the current node data vector as the corresponding current node feature encoding vector.

[0282] In another specific implementation manner of the embodiment of the present invention, when the graph processing module 12 performs unilateral addition processing based on the matching node records A, B, and the A-B association relationship:

[0283] Step K1, create a new first edge record in the first edge record table of the graph database 13 as the corresponding current new edge record; and create a new first edge feature vector record in the first edge feature vector record table of the graph database 13 as the corresponding current new vector record; and assign a unique record identifier to the current new edge record as the corresponding current edge identifier; and assign a unique record identifier to the current new vector record as the corresponding current edge vector identifier;

[0284] Step K2, extract the node identifier fields of the matching node records A and B as the corresponding node identifiers A and B; and sequentially splice the A-B association relationship, node identifier A, and node identifier B into a corresponding current edge data vector; and encode the current edge data vector based on a preset edge feature encoding algorithm to obtain a corresponding current edge feature encoding vector; and set the edge vector identifier field and edge feature vector field of the current new vector record to the corresponding current edge vector identifier and current edge feature encoding vector;

[0285] It should be noted here that the edge feature encoding algorithm of the embodiment of the present invention is a preset encoding algorithm, which can be a conventional word vector embedding encoding algorithm (such as Word2vec embedding encoding algorithm, etc.), or an embedding translation encoding algorithm for embedding and translating the node features of a knowledge graph (such as RotatE algorithm, TransE algorithm, etc.), or other encoding algorithms customized based on application requirements; if the edge feature encoding algorithm of the embodiment of the present invention is a conventional word vector embedding encoding algorithm, the current edge data vector can be directly encoded based on the corresponding word vector embedding encoding algorithm to obtain a corresponding current edge feature encoding vector; if it is an embedding translation encoding algorithm (such as RotatE algorithm, TransE algorithm, etc.), it is necessary to first obtain all the node data vectors and edge data vectors of the current medical knowledge graph to form a corresponding node data vector set and edge data vector set, and then based on the current embedding translation encoding algorithm (such as RotatE algorithm, TransE algorithm, etc.), the node and edge feature encoding vectors of all nodes and all edges of the current medical knowledge graph are translated as a whole to obtain a corresponding node feature encoding vector set and edge feature encoding vector set, and then the edge feature encoding vector in the edge feature encoding vector set corresponding to the current edge data vector is used as the corresponding current edge feature encoding vector;

[0286] Step K3, set the edge identification field and the association type field of the currently newly added edge record to the corresponding current edge identification and the A-B association relationship; and establish a corresponding foreign key-primary key mapping relationship between the head and tail node identification fields of the currently newly added edge record and the node identification fields of the matching node records A and B by setting the head and tail node identification fields of the currently newly added edge record to the corresponding node identification fields of the matching node records A and B; and establish a corresponding foreign key-primary key mapping relationship between the edge vector field of the currently newly added edge record and the edge vector identification field of the currently newly added vector record by setting the edge vector field of the currently newly added edge record to the edge vector identification field of the currently newly added vector record.

[0287] In another specific implementation manner of the embodiment of the present invention, the graph processing module 12 is specifically configured to, when performing single-node addition processing based on the currently newly added entity to obtain the corresponding newly added node record for the current time:

[0288] Step L1, use the first node record table corresponding to the entity type of the currently newly added entity in the graph database 13 as the corresponding current node record table;

[0289] Step L2, and add a first node record with all empty field contents in the current node record table as the corresponding newly added node record for the current time; and assign a unique record identifier to the newly added node record for the current time as the corresponding newly added node identifier for the current time; and set the node identification field and the node name field of the newly added node record for the current time to the corresponding newly added node identifier for the current time and the entity name of the currently newly added entity;

[0290] Step L3, and perform a round of traversal on all entity attributes of the entity attribute set of the currently newly added entity; and during this round of traversal process, use the currently traversed entity attribute as the corresponding current entity attribute; and add a first attribute record in the first attribute record table of the graph database 13 as the corresponding currently newly added attribute record; and assign a unique identifier to the currently newly added attribute record as the corresponding current attribute record identifier; and set the attribute identifier field and the attribute name field of the currently newly added attribute record based on the current attribute record identifier and the attribute name of the current entity attribute; and initialize an empty sequence in the attribute value sequence field of the currently newly added attribute record as the corresponding current attribute value sequence; and add a sequence element to the current attribute value sequence as the corresponding currently newly added element; and set the addition time and the attribute value of the currently newly added element based on the current time and the attribute value of the current entity attribute; and use the attribute field corresponding to the attribute name of the current entity attribute in the newly added node record for the current time as the corresponding current attribute field; and establish a corresponding foreign key-primary key mapping relationship between the current attribute field of the newly added node record for the current time and the attribute identifier field of the currently newly added attribute record by setting the current attribute field of the newly added node record for the current time to the corresponding current attribute record identifier;

[0291] Step L4. After the current round of traversal of all entity attributes in the entity attribute set of the newly added entity, use the node name field recorded in the newly added node at this time as the corresponding current name data; use the entity type corresponding to the newly added node record at this time as the corresponding current type data; perform a round of traversal on all attribute fields recorded in the newly added node at this time; during this round of traversal, use the currently traversed attribute field as the corresponding current attribute field; identify whether the current attribute field is empty; if so, form a corresponding second attribute key-value pair consisting of the attribute name corresponding to the current attribute field and an empty attribute value; if not, form a corresponding second attribute key-value pair consisting of the attribute name corresponding to the current attribute field and the attribute value of the sequence element with the latest addition time in the attribute value sequence field of the first attribute record mapped by the current attribute field; at the end of this round of traversal, form a corresponding current attribute sequence by sorting all the obtained second attribute key-value pairs in order; form a corresponding current node data vector from the current type data, current name data, and current attribute sequence corresponding to the newly added node record at this time; and encode the previous node data vector based on a preset node feature encoding algorithm to obtain a corresponding current node feature encoding vector;

[0292] Step L5. Add a first node feature vector record in the first node feature vector record table of the graph database 13 as the corresponding current node feature vector record; assign a unique identifier to the current node feature vector record as the corresponding current identifier; set the node vector identifier field and node feature vector field of the current node feature vector record to the corresponding current identifier and current node feature encoding vector; and establish a corresponding foreign key-primary key mapping relationship between the node vector field of the newly added node record at this time and the node vector identifier field of the current node feature vector record by setting the node vector field of the newly added node record at this time to the corresponding current identifier;

[0293] Step L6. Output the newly added node record with the settings completed as the result of this processing.

[0294] In another specific implementation manner of the embodiment of the present invention, the graph processing module 12 is specifically used for regularly predicting whether there are potential new connection relationships on the latest medical knowledge graph and performing new edge complementation processing on the current medical knowledge graph based on the prediction result to obtain the latest medical knowledge graph:

[0295] Step M1. Regularly use the first node set of the latest medical knowledge graph as the corresponding node set V;

[0296] Here, the node set V of the embodiment of the present invention consists of multiple nodes v. Each node v is assigned a unique integer value as the corresponding node index. The node v corresponds one-to-one with the first node in the first node set. The node feature of the node v consists of the first node name, the first node type, and the first node attribute set of the corresponding first node;

[0297] Step M2, and construct a virtual edge e with direction features and association relationship features between every two nodes v in the node set V, and form a corresponding virtual edge set E composed of all the obtained virtual edges e; and initialize the direction features and association relationship features of all the virtual edges e in the virtual edge set E to the corresponding invalid directions and invalid relationships;

[0298] Here, the edge feature of each virtual edge e of the embodiment of the present invention at least includes a direction feature and an association relationship feature; the direction feature includes forward, reverse, and invalid direction; if the node with a larger node index among the two nodes v corresponding to each virtual edge e is denoted as the large-index node, and the node with a smaller node index is denoted as the small-index node, then, when the direction feature is forward, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the large-index node to the small-index node, when it is reverse, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the small-index node to the large-index node, and when it is invalid direction, it indicates that the direction of the directed edge corresponding to the current virtual edge e is unknown; the association relationship feature includes all medical entity association types and invalid relationships of the medical knowledge graph;

[0299] Step M3, and perform a round of traversal on all the first edges in the first edge set of the latest medical knowledge graph; and during this round of traversal, regard the currently traversed first edge as the corresponding current entity edge; and regard the virtual edge e in the virtual edge set E corresponding to the current entity edge as the corresponding current virtual edge; and reset the direction feature of the current virtual edge based on the node indexes corresponding to the first nodes identified by the head and tail nodes of the current entity edge, and reset the association relationship feature of the current virtual edge based on the first association type of the current entity edge; and after the end of this round of traversal, denote the latest virtual edge set E as the corresponding initial edge set E ini ;

[0300] Step M4, and form a graph structure data from the node set V and the initial edge set E ini and denote it as the corresponding initial graph; and input the initial graph into a preset knowledge graph edge prediction model, and the knowledge graph edge prediction model predicts the direction type and association relationship type of each virtual edge e of the input initial graph to obtain the corresponding predicted edge set E * ;

[0301] Here, the knowledge graph edge prediction model of the embodiment of the present invention is a prediction model with a graph neural network model as the core encoder, such as Figure 6As shown; the predicted edge set E output by the model * consists of multiple predicted edges e * ; the predicted edge set E * has a predicted edge e * that corresponds one-to-one with the virtual edge e of the initial edge set E ini ; each predicted edge e * has edge features that at least include a direction feature and an association relationship feature;

[0302] Step M5, mark the virtual edges e in the initial edge set E ini whose direction features and association relationship features are corresponding invalid directions and invalid relationships as the corresponding initial invalid edges; and mark the predicted edges e * in the predicted edge set E whose direction features are not invalid directions, association relationship features are not invalid relationships, and that correspond to an initial invalid edge * as the corresponding newly added valid edges;

[0303] Step M6, and when the total number of the obtained newly added valid edges is not zero, perform a round of traversal on all the obtained newly added valid edges; and during this round of traversal, regard the currently traversed newly added valid edge as the corresponding currently newly added edge; and based on the direction feature of the currently newly added edge, perform corresponding head and tail node markings on the two first nodes corresponding to the currently newly added edge; and record the two first node records on the graph database 13 corresponding to the current head and tail nodes as a pair of new matching node records A and B; and regard the association relationship feature of the currently newly added edge as a new A-B association relationship; and perform single-edge addition processing based on the matching node records A, B and the A-B association relationship;

[0304] Step M7, and after the end of this round of traversal of all the newly added valid edges, add all the newly added valid edges to the current medical knowledge graph to obtain the latest medical knowledge graph.

[0305] In a specific implementation manner of the embodiment of the present invention, the knowledge graph edge prediction model is specifically used to perform prediction processing according to the direction types and association relationship types of the respective virtual edges e of the initial graph input to the model to obtain the corresponding predicted edge set E * .

[0306] As Figure 6 shown, the knowledge graph edge prediction model of the embodiment of the present invention consists of a graph embedding encoding module, a graph feature encoder, a direction prediction head, a relationship prediction head, and an output module; it should be noted that: the graph feature encoder is implemented based on a type of graph neural network model, and the graph neural network model at least includes a GCN model, a GNN model, and an ApeGNN model, that is, the graph neural network model of the embodiment of the present invention can be implemented at least based on the GCN model, the GNN model, or the ApeGNN model; the direction prediction head and the relationship prediction head are each implemented based on an MLP model.

[0307] In the embodiments of the present invention, the connection relationships of the components of the knowledge graph edge prediction model are as follows: The input end of the graph embedding encoding module is connected to the input end of the knowledge graph edge prediction model, and the output end is connected to the input end of the graph feature encoder; the output end of the graph feature encoder is respectively connected to the input ends of the direction and relationship prediction heads; the output ends of the direction and relationship prediction heads are respectively connected to one input end of the output module; the output end of the output module is connected to the output end of the knowledge graph edge prediction model.

[0308] In the embodiments of the present invention, the functions of the components of the knowledge graph edge prediction model are as follows:

[0309] 1) The graph embedding encoding module performs embedding encoding processing on the node set V and the initial edge set E of the initial graph input to the model according to the embedding encoding rules of the model input end of the graph neural network model corresponding to the graph feature encoder ini to obtain the corresponding node set embedding encoding tensor and edge set embedding encoding tensor respectively; and forms the corresponding graph embedding encoding tensor from the node set embedding encoding tensor and the edge set embedding encoding tensor and sends it to the graph feature encoder;

[0310] Among them, the node set embedding encoding tensor is composed of multiple node embedding encoding vectors, and the node embedding encoding vectors correspond one by one to the nodes v in the node set V; the edge set embedding encoding tensor is composed of multiple edge embedding encoding vectors, and the edge embedding encoding vectors correspond one by one to the ini virtual edges e in the initial edge set E;

[0311] 2) The graph feature encoder is used to perform edge feature extraction processing according to the graph embedding encoding tensor to obtain the corresponding edge set feature encoding vectors and send them to the direction and relationship prediction heads respectively;

[0312] Among them, the edge set feature encoding vector is composed of multiple edge feature encoding vectors, and the edge feature encoding vectors correspond one by one to the ini virtual edges e in the initial edge set E;

[0313] 3) The direction prediction head is used to predict the direction types of each edge according to the edge set feature encoding vector to obtain the corresponding edge set direction prediction vector and send it to the output module;

[0314] Among them, the edge set direction prediction vector is composed of multiple edge direction prediction vectors, and the edge direction prediction vectors correspond one by one to the ini virtual edges e in the initial edge set E; the edge direction prediction vector includes three edge direction prediction probabilities; each edge direction prediction probability corresponds to a type of edge direction type; the edge direction types include forward, reverse, and invalid directions;

[0315] 4) The relationship prediction head is used to predict the association relationship types of each edge according to the edge set feature encoding vector to obtain the corresponding edge set relationship prediction vector and send it to the output module;

[0316] Among them, the edge set relationship prediction vector is composed of multiple edge relationship prediction vectors, and the edge relationship prediction vectors correspond one-to-one to the virtual edges e of the initial edge set E ini ; the edge relationship prediction vector includes multiple edge relationship prediction probabilities; each edge relationship prediction probability corresponds to a type of edge relationship; the edge relationship types include all medical entity association types and invalid relationships in the medical knowledge graph;

[0317] 5) The output module is used to receive the edge set direction prediction vector and the edge set relationship prediction vector; and use the edge direction type corresponding to the maximum edge direction prediction probability in the edge direction prediction vectors corresponding to each virtual edge e as a corresponding first prediction direction; and use the edge relationship type corresponding to the maximum edge relationship prediction probability in the edge relationship prediction vectors corresponding to each virtual edge e as a corresponding first prediction relationship; and set a corresponding predicted edge e for each virtual edge e * , and set the direction feature and association relationship feature of the predicted edge e corresponding to the current virtual edge e based on the first prediction direction and the first prediction relationship corresponding to each virtual edge e * ; and the obtained all predicted edges e * constitute the corresponding predicted edge set E * .

[0318] It should also be noted that the training method of the knowledge graph edge prediction model in the embodiment of the present invention is generally similar to the self-supervised training method implemented by the traditional graph neural network model based on the masking mechanism: 1) Before training, first collect a large amount of data and construct a complete sample knowledge graph based on all prediction types and graph structures corresponding to the knowledge graph edge prediction model, and set the corresponding model training loss function based on the loss function of the conventional multi-classification prediction model, such as the conventional multi-classification cross-entropy loss function, etc.; 2) Then, use different masking ratios (such as 5%, 10%, 20%, 30%, 40%... 80%, 90%, 95%, etc.) to randomly mask the edge features (including direction features and association relationship features) of multiple node edges (such as 5%, 10%, 20%, 30%, 40%... 80%, 90%, 95% of the node edges) on the sample knowledge graph (such as setting all or part of the features of the node edges to incorrect values) to obtain multiple masked knowledge graphs; 3) Then, use each masked knowledge graph as the training graph and the sample knowledge graph as the label graph, and perform overall training on the knowledge graph edge prediction model based on each training-label graph pair and the model training loss function.

[0319] (III) Graph Database 13:

[0320] The atlas database 13 is used to store a series of data tables of the medical knowledge atlas. Here, the series of data tables in the embodiments of the present invention at least include a first resource record table, a first edge record table, a plurality of first node record tables, a first attribute record table, a first edge feature vector record table, and a first node feature vector record table, as Figure 7 shown.

[0321] Here, the first resource record table in the embodiments of the present invention is used to store information about the original data. When the first resource record table is not empty, it consists of one or more first resource records, as Figure 7 shown; the first resource record includes a resource identifier field, a resource type field, a resource name field, a content brief field, and a storage address field; the resource identifier field therein is set as the primary key.

[0322] The first edge feature vector record table in the embodiments of the present invention is used to store the feature vectors of all the first edges of the medical knowledge atlas. When the first edge feature vector record table is not empty, it consists of one or more first edge feature vector records, as Figure 7 shown; the first edge feature vector record corresponds to the first edge of the medical knowledge atlas one by one; the first edge feature vector record includes an edge vector identifier field and an edge feature vector field; the edge vector identifier field therein is set as the primary key.

[0323] The first node feature vector record table in the embodiments of the present invention is used to store the feature vectors of all the first nodes of the medical knowledge atlas. When the first node feature vector record table is not empty, it consists of one or more first node feature vector records, as Figure 7 shown; the first node feature vector record corresponds to the first node of the medical knowledge atlas one by one; the first node feature vector record includes a node vector identifier field and a node feature vector field; the node vector identifier field therein is set as the primary key.

[0324] The first attribute record table in the embodiments of the present invention is used to store the node attributes of all the first nodes of the medical knowledge atlas. When the first attribute record table is not empty, it consists of one or more first attribute records, as Figure 7 shown; the first attribute record includes an attribute identifier field, an attribute name field, and an attribute value sequence field; the attribute value sequence field is used to store an attribute value sequence; when the attribute value sequence is not empty, it is composed of one or more sequence elements sorted in chronological order, and the sequence element includes an addition time and an attribute value; the attribute identifier field therein is set as the primary key.

[0325] The first node record table in the embodiment of the present invention corresponds one-to-one with the first node type of the medical knowledge graph, that is, the entity type. Each first node record table is used to store information about all first nodes of a specified node type, that is, a specified entity type. When each first node record table is not empty, it is composed of one or more first node records, as Figure 7 shown; the first node record corresponds one-to-one with the first node of the medical knowledge graph; the first node record includes a node identifier field, a node name field, multiple attribute fields, and a node vector field; the total number and types of the attribute fields are consistent with the total number and types of the entity attributes corresponding to the entity type corresponding to the current first node record table; the node identifier field therein is set as the primary key field, and each attribute field and the node vector field are set as foreign key fields; each attribute field forms a one-to-one foreign key-primary key mapping relationship with the attribute identifier field of a first attribute record, and the node vector field forms a one-to-one foreign key-primary key mapping relationship with the node vector identifier field of a first node feature vector record.

[0326] The first edge record table in the embodiment of the present invention is used to store information about all first edges of the medical knowledge graph. When the first edge record table is not empty, it is composed of one or more first edge records, as Figure 7 shown; the first edge record corresponds one-to-one with the first edge of the medical knowledge graph; the first edge record includes an edge identifier field, an association type field, a head node identifier field, a tail node identifier field, and an edge vector field; the total number and type range of the association type fields are consistent with the total number and relationship range of the entity relationship R; the edge identifier field therein is set as the primary key field, and the head node identifier field, the tail node identifier field, and the edge vector field are set as foreign key fields; the head node identifier field forms a one-to-one foreign key-primary key mapping relationship with the node identifier field of a first node record; the tail node identifier field forms a one-to-one foreign key-primary key mapping relationship with the node identifier field of another first node record; the edge vector field forms a one-to-one foreign key-primary key mapping relationship with the edge vector identifier field of a first edge feature vector record.

[0327] (IV) Information retrieval module 14:

[0328] The information retrieval module 14 is configured to receive a first retrieval request sent by the client 2; extract the corresponding first user identifier and the first retrieval text from the first retrieval request; perform retrieval keyword extraction processing on the first retrieval text to obtain a corresponding first keyword set; perform user behavior data addition processing according to the first user identifier, the first keyword set, and the user behavior database 15; and perform knowledge graph retrieval processing according to the first user identifier, the first keyword set, the user behavior database 15, and the graph database 13 to obtain a corresponding first retrieval graph and send it back to the client 2. Here, the first retrieval request in the embodiment of the present invention includes a first user identifier and a first retrieval text; the first keyword set is composed of one or more first keywords.

[0329] In another specific implementation manner of the embodiment of the present invention, when the information retrieval module 14 is specifically configured to perform user behavior data addition processing according to the first user identifier, the first keyword set, and the user behavior database 15:

[0330] Take the first user record in the first user record table of the user behavior database 15 whose user identifier field matches the first user identifier as the corresponding current user record; and when the current user record is not empty, take the first behavior record table corresponding to the behavior table index field of the current user record as the corresponding current behavior record table; and when the current behavior record table is not empty, add a new first behavior record in the current behavior record table as the corresponding current behavior record; and assign a unique identifier to the current behavior record as the corresponding current record identifier; and set the behavior record identifier field and the behavior time field of the current behavior record to the corresponding current record identifier and the current time; and set the behavior type field of the current behavior record to retrieval; and sequentially splice all the first keywords in the first keyword set into a single sentence text denoted as the corresponding first keyword text; and identify the preset text triple extraction model set. If the text triple extraction model set is the first model set, use the preset first entity naming model to identify the entity type range involved in the first keyword text to obtain one or more corresponding entity types. If the text triple extraction model set is the second model set, use the preset second entity naming model to identify the entity type range involved in the first keyword text to obtain one or more corresponding entity types; and form a corresponding entity type sequence from all the entity types identified this time, and set the knowledge type field of the current behavior record to the corresponding entity type sequence.

[0331] In another specific implementation manner of the embodiment of the present invention, when the information retrieval module 14 is specifically configured to perform knowledge graph retrieval processing according to the first user identifier, the first keyword set, the user behavior database 15, and the graph database 13 to obtain a corresponding first retrieval graph and send it back to the client 2:

[0332] Step N1, use the first user record corresponding to the first user identifier in the first user record table of the user behavior database 15 as the corresponding current user record; and extract the personalized feature field of the current user record as the corresponding current user feature set;

[0333] Step N2, and obtain the first keyword vector corresponding to the first keyword set; and use a preset node feature encoding algorithm to encode the first keyword vector to obtain the corresponding first feature vector; and use the node feature vector fields of each first node feature vector record in the first node feature vector record table in the atlas database 13 as the corresponding first comparison vectors; and calculate the vector similarity between the first feature vector and each first comparison vector to obtain the corresponding second similarity; and record the first node feature vector record corresponding to the second similarity that exceeds the preset second similarity threshold as the candidate feature record; and use the first node record having a mapping relationship with each candidate feature record as the corresponding primary selection node record;

[0334] Here, the second similarity threshold is a preset threshold parameter;

[0335] Step N3, and identify the current user feature set; if the current user feature set is empty, record all the primary selection node records as the corresponding secondary selection node records; if the current user feature set is not empty, only record the primary selection node records that match each personalized feature of the current user feature set as the corresponding secondary selection node records;

[0336] Step N4, and perform a round of traversal on all the secondary selection node records; and during this round of traversal, use the currently traversed secondary selection node record as the corresponding current node record; and use the first node corresponding to the current node record in the medical knowledge atlas as the corresponding current node; and select a sub-atlas centered on the current node in the medical knowledge atlas as the corresponding first sub-atlas;

[0337] Here, the maximum node distance between each first node and the current node in the first sub-atlas of the embodiment of the present invention is K; where K is a preset positive integer, and the node distance is the total number of other nodes between two directly or indirectly connected first nodes;

[0338] Step N5, and at the end of this round of traversal of all the secondary selection node records, send back the corresponding first retrieval atlas composed of all the first sub-atlases to the client 2.

[0339] (V) User behavior database 15:

[0340] The user behavior database 15 is used to store the first user record table and multiple first behavior record tables, such as Figure 8As shown below. Here, the first row record table in the embodiment of the present invention corresponds to a user one by one and is used to store all the behavior data of a specified user. Each first row record table corresponds to a unique table index; when the first row record table is not empty, it consists of one or more first row records, such as Figure 8 shown; the first row record includes a behavior record identification field, a behavior time field, a behavior type field, and a knowledge type field; the behavior type field at least includes retrieval, browsing, liking, forwarding, and collection; the knowledge type field at least includes all entity types; the record identification field is set as the primary key.

[0341] The first user record table in the embodiment of the present invention is used to store the user information of all users. When the first user record table is not empty, it consists of one or more first user records, such as Figure 8 shown; the first user record corresponds to a user one by one; the first user record includes a user identification field, a name field, an age field, a gender field, a user type field, a personalized feature field, and a behavior table index field; the user type field at least includes multiple types of medical practitioners and multiple types of medical students; the personalized feature field is used to store a personalized feature set, and when the personalized feature set is not empty, it consists of one or more personalized features, and the personalized feature is a type of entity type; the user identification field is set as the primary key; the behavior table index field is the table index of the first row record table corresponding to the user corresponding to the current first user record.

[0342] (6) Behavior acquisition module 16:

[0343] The behavior acquisition module 16 is used to collect user behavior data through the client 2 and perform user behavior data addition processing based on the collected data and the user behavior database 15.

[0344] The behavior acquisition module 16 is also used to perform behavior analysis regularly according to the user behavior database 15.

[0345] In another specific implementation manner of the embodiment of the present invention, the behavior acquisition module 16 is specifically used when collecting user behavior data through the client 2 and performing user behavior data addition processing based on the collected data and the user behavior database 15:

[0346] Step O1, monitor the user's information browsing, information liking, information forwarding, and information collection behaviors through the client 2; and when the user completes an information browsing, information liking, information forwarding, or information collection operation each time, use the current operation time as a corresponding behavior collection time, use the information theme of the current browsing, liking, forwarding, or collecting information as a corresponding behavior collection theme, use the current browsing, liking, forwarding, or collecting behavior type as a corresponding behavior collection type, and form a corresponding behavior collection data from the behavior collection time, behavior collection theme, and behavior collection type corresponding to the current operation.

[0347] Step O2, and when obtaining a behavior collection data of a user through the client 2 each time, use the first user record in the first user record table of the user behavior database 15 that matches the current user as the corresponding current user record; and when the current user record is not empty, use the first behavior record table corresponding to the behavior table index field of the current user record as the corresponding current behavior record table; and when the current behavior record table is not empty, add a new first behavior record to the current behavior record table as the corresponding current behavior record; and assign a unique identifier to the current behavior record as the corresponding current record identifier; and set the behavior record identifier field of the current behavior record to the corresponding current record identifier; and set the behavior time field of the current behavior record to the behavior collection time of the current behavior collection data; and set the behavior type field of the current behavior record to the behavior collection type of the current behavior collection data; use the behavior collection theme of the current behavior collection data as the corresponding first theme text; and identify the preset text triple extraction model set. If the text triple extraction model set is the first model set, use the preset first entity naming model to identify the entity type involved in the first theme text to obtain a corresponding entity type. If the text triple extraction model set is the second model set, use the preset second entity naming model to identify the entity type involved in the first theme text to obtain a corresponding entity type; and set the knowledge type field of the current behavior record based on the entity type identified this time.

[0348] In another specific implementation manner of the embodiment of the present invention, the behavior collection module 16 is specifically used for: when performing behavior analysis according to the user behavior database 15 regularly

[0349] Step P1, regularly use each first user record in the first user record table of the user behavior database 15 as the corresponding current user record;

[0350] Step P2, and extract the age field, gender field, and user type field of the current user record as the corresponding user age, user gender, and user type to form a corresponding user basic parameter;

[0351] Step P3, and use the first row of the record table corresponding to the current user record as the corresponding current behavior record table; extract the last updated X first row records in the current behavior record table and sort them in chronological order to form the corresponding first record sequence; extract the behavior time field, behavior type, and knowledge type field of each first row record in the first record sequence as the corresponding first time, first behavior type, and first entity type sequence to form a corresponding first behavior vector; and form a corresponding user behavior sequence by sorting all the obtained first behavior vectors in chronological order.

[0352] Here, X is a preset positive integer.

[0353] Step P4, and input the user basic parameters and the user behavior sequence into a preset user behavior prediction model. The user behavior prediction model performs personalized feature prediction processing based on the user basic parameters and the user behavior sequence to obtain the corresponding prediction feature set.

[0354] Here, the user behavior prediction model in the embodiment of the present invention is a classification prediction model implemented based on a deep learning model framework, as Figure 9 shown; when the prediction feature set output by the model is not empty, it is composed of one or more prediction features, and each prediction feature is a type of entity type.

[0355] Step P5, and when the current prediction feature set is not empty, reset the personalized feature field of the current user record based on the current prediction feature set.

[0356] In another specific implementation manner of the embodiment of the present invention, the user behavior prediction model is specifically used to perform personalized feature prediction processing based on the user basic parameters and the user behavior sequence input into the model to obtain the corresponding prediction feature set; here, as shown above, when the prediction feature set output by the model is not empty, it is composed of one or more prediction features, and each prediction feature is a type of entity type.

[0357] As Figure 9 shown, the user behavior prediction model in the embodiment of the present invention is composed of a behavior feature extraction network, a basic feature encoding module, a feature fusion module, a personalized prediction head, and a prediction output module. It should be noted that: the behavior feature extraction network includes at least an LSTM model and a bi-LSTM model, that is, the behavior feature extraction network of the embodiment of the present invention can be implemented at least based on the LSTM model or the bi-LSTM model; the personalized prediction head is implemented based on a type of multi-classification prediction model.

[0358] In the embodiment of the present invention, the connection relationships of the components of the user behavior prediction model are as follows: The input ends of the behavior feature extraction network and the basic feature encoding module are respectively connected to one model input end of the user behavior prediction model, and the output ends of the behavior feature extraction network and the basic feature encoding module are respectively connected to one input end of the feature fusion module; the output end of the feature fusion module is connected to the input end of the personalized prediction head; the output end of the personalized prediction head is connected to the input end of the prediction output module; the output end of the prediction output module is connected to the output end of the user behavior prediction model.

[0359] In the embodiment of the present invention, the functions of the components of the user behavior prediction model are as follows:

[0360] 1) The behavior feature extraction network is used to perform feature encoding processing on the user behavior sequence input to the model to obtain a corresponding first encoded vector and send it to the feature fusion module;

[0361] 2) The basic feature encoding module is used to extract the corresponding user age, user gender, and user type from the user basic parameters input to the model; and perform normalized encoding on the user age to obtain a corresponding age encoding, perform one-hot encoding on the user gender to obtain a corresponding gender encoding, perform one-hot encoding on the user type to obtain a corresponding type encoding, and form a corresponding second encoded vector composed of the obtained age encoding, gender encoding, and type encoding and send it to the feature fusion module;

[0362] 4) The feature fusion module is used to splice the first and second encoded vectors and use the obtained spliced vector as a corresponding third encoded vector and send it to the personalized prediction head;

[0363] 5) The personalized prediction head is used to perform classification prediction processing on the third encoded vector to obtain a corresponding personalized prediction vector and send it to the prediction output module;

[0364] Among them, the personalized prediction vector is composed of multiple personalized prediction probabilities; each personalized prediction probability corresponds to an entity type;

[0365] 6) The prediction output module is used to record the personalized prediction probabilities in the personalized prediction vector whose probability values exceed the preset probability threshold as corresponding candidate prediction probabilities; and identify the total number of the obtained candidate prediction probabilities; if the total number is zero, set the corresponding prediction feature set to be empty; if the total number is greater than zero, use the entity types corresponding to each candidate prediction probability as a corresponding prediction feature, and form a corresponding prediction feature set from all the obtained prediction features. The preset probability threshold here is a threshold parameter set in advance.

[0366] It should also be noted that the training method of the user behavior prediction model in the embodiments of the present invention is implemented by a supervised model training method. Specifically: 1) Before training, a large amount of training data is collected based on the model input data format of the user behavior prediction model, and corresponding label data is set for each training data based on the model output data format of the user behavior prediction model. A corresponding training data set is composed of a large number of training-label data pairs, and the model training loss function is set based on a conventional classification prediction loss function, such as a conventional cross-entropy loss function or a cross-entropy loss function with a regularization penalty term, etc.; 2) Then, the user behavior prediction model is trained based on the training data set and the model training loss function.

[0367] It should also be noted that it should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; it is also possible that some modules are implemented in the form of software called by a processing element, and some modules are implemented in the form of hardware. For example, the data receiving module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module is called and executed by a certain processing element of the above system. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each method step of the foregoing method or each module processing step of the foregoing system can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0368] For example, these modules of the above system can be one or more integrated circuits configured to implement the foregoing method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module of the above system is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0369] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The above available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0370] An embodiment of the present invention provides a management system for a medical knowledge graph. The system includes: a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database, and a behavior collection module; wherein, the graph processing module is used for extracting medical knowledge triples, updating the medical knowledge graph, and complementing the medical knowledge graph; the graph database is used for storing a series of data tables of the medical knowledge graph; the information retrieval module is used for performing personalized graph retrieval in combination with user behavior characteristics; the user behavior database is used for storing user information; the behavior collection module is used for collecting user behavior data and analyzing user behavior characteristics. The embodiment of the present invention can achieve the technical purpose of managing the specialized medical knowledge in a certain disease field or the general medical knowledge in the whole field based on the knowledge graph technology. Based on the embodiment of the present invention, not only the management efficiency and retrieval efficiency of knowledge are improved, but also a personalized knowledge retrieval function is provided for users.

[0371] Those skilled in the art should also be able to further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0372] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0373] The specific embodiments described above have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A management system for a medical knowledge graph, characterized in that The system includes: a data receiving module, a graph processing module, a graph database, an information retrieval module, a user behavior database, and a behavior collection module; The data receiving module is connected to the graph processing module; the graph database is respectively connected to the graph processing module and the information retrieval module; the user behavior database is respectively connected to the information retrieval module and the behavior collection module; The data receiving module is configured to receive a first original data packet sent by a client; perform data preprocessing on the first original data packet to obtain a corresponding first preprocessed data packet; and send the first preprocessed data packet and the first original data packet to the graph processing module; The graph processing module is configured to perform duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet, and the graph database; perform medical knowledge triple extraction processing on the first preprocessed data packet to obtain a corresponding first triple set; and update the medical knowledge graph according to the first triple set and the graph database; The graph processing module is further configured to periodically predict whether there are potential new connection relationships on the latest medical knowledge graph and perform new edge complement processing on the current medical knowledge graph based on the prediction result to obtain the latest medical knowledge graph; The graph database is used to store a series of data tables of the medical knowledge graph; the series of data tables at least includes a first resource record table, a first edge record table, multiple first node record tables, a first attribute record table, a first edge feature vector record table, and a first node feature vector record table; The information retrieval module is configured to receive a first retrieval request sent by the client; extract a corresponding first user identifier and a first retrieval text from the first retrieval request; perform retrieval keyword extraction processing on the first retrieval text to obtain a corresponding first keyword set; perform user behavior data addition processing according to the first user identifier, the first keyword set, and the user behavior database; and perform knowledge graph retrieval processing according to the first user identifier, the first keyword set, the user behavior database, and the graph database to obtain a corresponding first retrieval graph and send it back to the client; the first retrieval request includes the first user identifier and the first retrieval text; the first keyword set consists of one or more first keywords; The user behavior database is used to store a first user record table and multiple first behavior record tables; The behavior collection module is configured to collect user behavior data through the client and perform user behavior data addition processing based on the collected data and the user behavior database; The behavior collection module is further configured to perform behavior analysis according to the user behavior database periodically.

2. The management system of the medical knowledge graph according to claim 1, wherein The first original data packet consists of one or more first file packets; each of the first file packets consists of a first medical file and a corresponding first file parameter set; the first file parameter set includes file type, file name, and content summary; the file type includes at least PDF files, table files, image files, audio files, and video files; when the file type is a PDF file, the corresponding first medical file is a medical literature material in a certain medical knowledge field, and the corresponding content summary includes the file source information, release time information, and content abstract information corresponding to the current first medical file; when the file type is a table file, the corresponding first medical file is a formatted medical data table; when the file type is an image file, audio file, or video file, the corresponding first medical file is a medical examination image, medical examination audio, or medical examination video generated by a certain type of medical examination corresponding to a certain disease, and the corresponding content summary includes the disease information, examination category information, and examination description information corresponding to the current first medical file; the medical examination images include at least X-ray images, CT images, ultrasound images, magnetic resonance images, and nuclear medicine images; the medical examination audio includes at least heart sound auscultation audio, lung auscultation audio, and abdominal auscultation audio; the medical examination videos include at least endoscopic examination videos and dynamic medical imaging videos; The first preprocessing data packet consists of one or more first preprocessing data; the first preprocessing data corresponds one-to-one with the first file packet; when the file type of the first file packet is a PDF file, the corresponding first preprocessing data includes a first data type and a first sentence sequence, and the first data type is specifically the first type, and the first sentence sequence is formed by sorting multiple first sentences in sequence; when the file type of the first file packet is a table file, the corresponding first preprocessing data includes the first data type and a first data table, and the first data type is specifically the second type; when the file type of the first file packet is an image file, the corresponding first preprocessing data includes the first data type, a first description, and a first image, and the first data type is the third type; when the file type of the first file packet is an audio file, the corresponding first preprocessing data includes the first data type, the first description, and a first audio, and the first data type is the fourth type; when the file type of the first file packet is a video file, the corresponding first preprocessing data includes the first data type, the first description, and a first video, and the first data type is the fifth type; The first triple set includes a plurality of first triples; the first triples are composed of corresponding entity A, entity relationship R, and entity B according to the knowledge triple structure of entity-relationship-entity; the entity parameters of entity A and B are both composed of entity name, entity type, and entity attribute set; the entity type includes multiple medical entity types, including at least multiple drug entity types, multiple disease entity types, multiple disease symptom entity types, multiple disease treatment plan entity types, multiple disease clinical record entity types, multiple disease clinical experiment entity types, multiple biochemical gene entity types, multiple medical examination type entity types, multiple medical test type entity types, and the entity type is consistent with the type range of the knowledge graph node type; the entity attribute set corresponds one-to-one with the entity type, the entity attribute set includes multiple entity attributes, the entity attribute includes an attribute name and an attribute value, and the number and types of the entity attributes corresponding to each type of medical entity type are fixed; the entity relationship R includes multiple medical entity association types, and the entity relationship R is consistent with the relationship range of the knowledge graph edge association relationship. The medical knowledge graph is composed of a first node set and a first edge set; the first node set includes a plurality of first nodes; the first edge set includes a plurality of first edges; the node parameters of the first node include a first node identifier, a first node name, a first node type, and a first node attribute set, the first node attribute set includes multiple first node attributes, and the first node attribute is composed of an attribute name and an attribute value; the edge parameters of the first edge include a first edge identifier, a first association type, and a first association node group; the first association node group includes a head node identifier and a tail node identifier. When the first resource record table is not empty, it is composed of one or more first resource records; the first resource record includes a resource identifier field, a resource type field, a resource name field, a content summary field, and a storage address field; the resource identifier field is set as the primary key. When the first edge feature vector record table is not empty, it is composed of one or more first edge feature vector records; the first edge feature vector record corresponds one-to-one with the first edge of the medical knowledge graph; the first edge feature vector record includes an edge vector identifier field and an edge feature vector field; the edge vector identifier field is set as the primary key. When the first node feature vector record table is not empty, it is composed of one or more first node feature vector records; the first node feature vector record corresponds one-to-one with the first node of the medical knowledge graph; the first node feature vector record includes a node vector identifier field and a node feature vector field. The node vector identifier field is set as the primary key. When the first attribute record table is not empty, it consists of one or more first attribute records; the first attribute record includes an attribute identification field, an attribute name field, and an attribute value sequence field; the attribute value sequence field is used to store an attribute value sequence; when the attribute value sequence is not empty, it is composed of one or more sequence elements sorted in chronological order, and the sequence element includes an addition time and an attribute value; the attribute identification field is set as the primary key; The first node record table corresponds one-to-one with the first node type of the medical knowledge graph, that is, the entity type; when the first node record table is not empty, it consists of one or more first node records; the first node record corresponds one-to-one with the first node of the medical knowledge graph; the first node record includes a node identification field, a node name field, multiple attribute fields, and a node vector field; the total number and types of the attribute fields are consistent with the total number and types of the entity attributes corresponding to the entity type corresponding to the current first node record table; the node identification field is set as the primary key field; each of the attribute fields and the node vector field is set as a foreign key field; each of the attribute fields forms a one-to-one foreign key-primary key mapping relationship with the attribute identification field of a first attribute record; the node vector field forms a one-to-one foreign key-primary key mapping relationship with the node vector identification field of a first node feature vector record; When the first edge record table is not empty, it consists of one or more first edge records; the first edge record corresponds one-to-one with the first edge of the medical knowledge graph; the first edge record includes an edge identification field, an association type field, a head node identification field, a tail node identification field, and an edge vector field; the total number and range of types of the association type field are consistent with the total number and range of relationships of the entity relationship R; the edge identification field is set as the primary key field; the head node identification field, the tail node identification field, and the edge vector field are all set as foreign key fields; the head node identification field forms a one-to-one foreign key-primary key mapping relationship with the node identification field of a first node record; the tail node identification field forms a one-to-one foreign key-primary key mapping relationship with the node identification field of another first node record; the edge vector field forms a one-to-one foreign key-primary key mapping relationship with the edge vector identification field of a first edge feature vector record; The first behavior record table corresponds one-to-one with the user; each first behavior record table corresponds to a unique table index; when the first behavior record table is not empty, it consists of one or more first behavior records; the first behavior record includes a behavior record identification field, a behavior time field, a behavior type field, and a knowledge type field; the behavior type field at least includes retrieval, browsing, liking, forwarding, and collection; the knowledge type field at least includes all the entity types; the record identification field is set as the primary key; When the first user record table is not empty, it consists of one or more first user records; the first user records correspond to users one by one; the first user record includes a user identification field, a name field, an age field, a gender field, a user type field, a personalized feature field, and a behavior table index field; the user type field includes at least multiple types of medical practitioner types and multiple types of medical student types; the personalized feature field is used to store a personalized feature set, which consists of one or more personalized features when not empty, and the personalized feature is a type of the entity type; The user identification field is set as the primary key; the behavior table index field is the table index of the first behavior record table corresponding to the user corresponding to the current first user record.

3. The management system of the medical knowledge graph according to claim 2, wherein The graph processing module is specifically used for when performing duplicate resource screening and new resource storage according to the first original data packet, the first preprocessed data packet, and the graph database: Taking each of the first file packets of the first original data packet as a corresponding current file packet one by one; and extracting the corresponding first medical file and the first file parameter set from the current file packet as the corresponding current medical file and current file parameter set; and extracting the corresponding file type, file name, and content summary from the current file parameter set as the corresponding current type, current name, and current description; And performing word embedding encoding on the current name based on the bag-of-words encoding algorithm to obtain a corresponding first name encoding vector; and performing word segmentation processing on the current description to obtain a corresponding current word segmentation sequence, and then performing word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain a corresponding first description encoding vector; and performing vector splicing on the obtained first name encoding vector and the first description encoding vector to obtain a corresponding first splicing vector; Initialize the duplicate file check status as not duplicate; mark all the first resource records in the first resource record table of the atlas database whose resource type fields match the current type as corresponding records to be traversed; perform a round of traversal on all the records to be traversed; during this round of traversal, use the currently traversed record to be traversed as the corresponding current record; perform word embedding encoding on the resource name field of the current record based on the bag-of-words encoding algorithm to obtain the corresponding second name encoding vector; perform word segmentation on the content summary field of the current record to obtain the corresponding current word segmentation sequence, and then perform word embedding encoding on the current word segmentation sequence based on the Word2Vec encoding algorithm to obtain the corresponding second description encoding vector; splice the obtained second name encoding vector and the second description encoding vector to obtain the corresponding second splicing vector; calculate the vector similarity of the first and second splicing vectors based on the cosine vector similarity algorithm to obtain the corresponding current similarity; identify whether the current similarity exceeds the preset first similarity threshold; if so, set the duplicate file check status to duplicate and stop this round of traversal; if not, go to the next record to be traversed and continue traversing until the last record to be traversed is traversed; Identify the obtained duplicate file check status; If the duplicate file check status is not duplicate, add a new first resource record in the first resource record table as the corresponding current new record; store the current medical file and use the storage address as the corresponding current storage address; set a unique identifier for the current new record as the corresponding current record identifier; Set the corresponding resource identifier field, resource type field, resource name field, content summary field, and storage address field in the current new record based on the current record identifier, current type, current name, current description, and current storage address; If the duplicate file check status is duplicate, delete the current file package from the first original data packet and synchronously delete the corresponding first preprocessed data in the first preprocessed data packet for the current file package.

4. The management system of the medical knowledge atlas according to claim 2, wherein The atlas processing module is specifically used for when performing medical knowledge triple extraction processing on the first preprocessed data packet to obtain the corresponding first triple set: Take each of the first preprocessed data packets of the first preprocessed data packet as the corresponding current preprocessed data; and use the first data type of the current preprocessed data as the corresponding current data type; Identify the current data type; If the current data type is the first type, identify the preset text triple extraction model set; if the text triple extraction model set is the first model set, perform medical knowledge triple identification on the first sentence sequence of the current preprocessed data based on the preset first entity naming model, first relation extraction model, and first attribute extraction models corresponding to various entity types to obtain the corresponding first triple subset; If the text triple extraction model set is the second model set, perform medical knowledge triple identification on the first sentence sequence of the current preprocessed data based on the preset second entity naming model, second relation extraction model, and second attribute extraction models corresponding to various entity types to obtain the corresponding first triple subset; The text triple extraction model set includes a first model set and a second model set; the first entity naming model, the first relation extraction model, and each of the first attribute extraction models are each implemented based on a stacking model framework with a multi-class logistic regression model as the base model; the second entity naming model, the second relation extraction model, and each of the second attribute extraction models are each implemented based on an NLP model framework with a BERT model as the core encoder; the first triple subset includes one or more of the first triples; If the current data type is the second type, perform knowledge triple identification on the first data table of the current preprocessed data based on the preset formatted data table - medical knowledge triple conversion template to obtain the corresponding first triple subset; If the current data type is the third type, perform medical knowledge triple identification on the first description and the first image of the current preprocessed data and the preset image attribute prediction models corresponding to various medical examination images to obtain the corresponding first triple subset; each of the image attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network; If the current data type is the fourth type, perform medical knowledge triple identification on the first description and the first audio of the current preprocessed data and the preset audio attribute prediction models corresponding to various medical examination audio to obtain the corresponding first triple subset; each of the audio attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network; If the current data type is the fifth type, perform medical knowledge triple identification on the first description and the first video of the current preprocessed data and the preset video attribute prediction models corresponding to various medical examination videos to obtain the corresponding first triple subset; each of the video attribute prediction models is implemented based on a deep learning model framework composed of a feature extraction network and an attribute prediction network; Merge all the first triple subsets obtained from the first preprocessed data packet to obtain a corresponding first set; remove duplicates from the first triples in the first set, and use the first set after deduplication as the corresponding first triple set.

5. The management system of the medical knowledge graph according to claim 4, wherein the first entity naming model is used to perform entity classification processing on each word segment in the first text input to the model according to a preset entity classification set and output a corresponding first entity classification vector; the entity classification set includes all the entity types; the first text is a single-sentence text; the first entity classification vector is composed of multiple first word segment classification vectors, and the first word segment classification vectors correspond one-to-one with the word segments in the first text; the first word segment classification vector is composed of multiple first entity classification probabilities, and the first entity classification probabilities correspond one-to-one with the entity types; The first entity naming model consists of a first preprocessing unit, Na parallel base models BM-1 a and a meta-model MM-1; all the base models BM-1 a are implemented based on the principle of the multi-class logistic regression model, and all the base models BM-1 a have the same model structure but different hyperparameters. Na is the number of all hyperparameter combinations, and 1 ≤ base model index a ≤ Na; the meta-model MM-1 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameters are one of all hyperparameter combinations of the multi-class logistic regression model; the first relationship extraction model is used to perform relationship classification processing on the first entity pair features input to the model according to a preset association relationship set and output a corresponding first relationship classification vector; the association relationship set includes all the medical entity association types corresponding to all the entity relationships R; the first entity pair features include a first context word segment sequence, a head entity feature, and a tail entity feature; the first context word segment sequence is sorted by multiple first word segment texts; the head entity feature includes a head entity index and a head entity type; the tail entity feature includes a tail entity index and a tail entity type; the head and tail entity indexes are the word segment text indexes of the corresponding head and tail entities in the first context word segment sequence; the head and tail entity types are one type of the entity types; the first relationship classification vector is composed of multiple first relationship classification probabilities, and each of the first relationship classification probabilities corresponds to one type of the medical entity association types; The first relation extraction model consists of a second preprocessing unit, Nb parallel base models BM-2 b and a meta-model MM-2; all the base models BM-2 b are implemented based on the principle of the multi-class logistic regression model. All the base models BM-2 b have the same model structure but different hyperparameters. Nb is the number of all hyperparameter combinations, and 1 ≤ base model index b ≤ Nb; the meta-model MM-2 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameters are one of all hyperparameter combinations of the multi-class logistic regression model. The number of models of the first attribute extraction model is equal to the number of entity types, and the first attribute extraction model corresponds to the entity type one by one; each of the first attribute extraction models is used to perform attribute classification processing on the basis of all attribute type categories of the corresponding single-entity attribute type set according to the first entity features input to the model and output the corresponding first attribute classification vector; the single-entity attribute type set corresponds to one of the entity types and consists of multiple attribute types, and each attribute type matches the attribute name of a category of entity attributes corresponding to the current entity type; the first entity features include a second context word segmentation sequence and a first entity index; the second context word segmentation sequence is formed by sorting multiple second word segmentation texts; the first entity index is a word segmentation text index in the second context word segmentation sequence; the first attribute classification vector is composed of multiple first word segmentation attribute vectors, and the first word segmentation attribute vectors correspond to the second word segmentation texts in the second context word segmentation sequence one by one; the vector length of the first word segmentation attribute vector matches the total number of attribute types of the corresponding single-entity attribute type set and consists of multiple first attribute classification probabilities, and the first attribute classification probabilities correspond to the attribute types of the corresponding single-entity attribute type set one by one; The first attribute extraction model consists of a third preprocessing unit, Nc parallel base models BM-3 c and a meta-model MM-3; all the base models BM-3 c are implemented based on the principle of the multi-class logistic regression model, and all the base models BM-3 c have the same model structure but different hyperparameters. Nc is the number of all hyperparameter combinations, and 1 ≤ base model index c ≤ Nc; the meta-model MM-3 is also implemented based on the principle of the multi-class logistic regression model, and its hyperparameters are one of all hyperparameter combinations of the multi-class logistic regression model.

6. The management system of the medical knowledge graph according to claim 4, wherein The second entity naming model is used to perform entity classification processing on each word segmentation in the second text input to the model according to the preset entity classification set and output the corresponding second entity classification vector; the second text is a single-segment text and consists of one or more single-sentence texts; the second entity classification vector is composed of multiple second word segmentation classification vectors, and the second word segmentation classification vectors correspond to the word segmentations in the second text one by one; the second word segmentation classification vector is composed of multiple second entity classification probabilities, and the second entity classification probabilities correspond to the entity types one by one; The second entity naming model is composed of a fourth preprocessing unit, a first BERT model, a first linear network, and a first Softmax layer; the first BERT model is one of a pre-trained basic BERT model, BioBERT, and ClinicalBERT; the first linear network is implemented based on one or more fully connected layers; The second relation extraction model is used to perform entity pair relation classification processing according to the preset association relation set, the second word segmentation sequence and the second entity classification vector input by the model, and output the corresponding first word pair classification matrix; the second word segmentation sequence is the word segmentation sequence generated by the fourth preprocessing unit of the second entity naming model, and the second entity classification vector is the entity classification vector output by the second entity naming model; each row or each column of the first word pair classification matrix corresponds one by one to the second word segments of the second word segmentation sequence; each matrix unit not on the diagonal in the first word pair classification matrix is a second relation classification vector of a word pair, and each matrix unit on the diagonal of the first word pair classification matrix is an invalid classification vector with a vector length consistent with that of the second relation classification vector and all vector data being zero or all being negative values; the second relation classification vector is composed of multiple second relation classification probabilities, and each of the second relation classification probabilities corresponds to one type of the medical entity association types; The second relation extraction model is composed of a fifth preprocessing unit, a second BERT model, a second linear network and a second Softmax layer; the model structures of the first and second BERT models are the same and the model parameters are the same; The second linear network is implemented based on one or more fully connected layers; The number of the second attribute extraction models is equal to the number of the entity types, and the second attribute extraction models correspond to the entity types one by one; each of the second attribute extraction models is used to perform attribute classification processing according to all the attribute type categories of the corresponding single entity attribute type set, based on the second word segmentation sequence and the first entity position input by the model, and output the corresponding second attribute classification vector; the second word segmentation sequence is the word segmentation sequence generated by the fourth preprocessing unit of the second entity naming model, and the first entity position is the word segmentation index corresponding to one of the second word segments in the second word segmentation sequence; the second attribute classification vector is composed of multiple second word segment attribute vectors, and the second word segment attribute vectors correspond to the second word segments in the second word segmentation sequence one by one; the vector length of the second word segment attribute vector matches the total number of the attribute types of the corresponding single entity attribute type set and is composed of multiple second attribute classification probabilities, and the second attribute classification probabilities correspond to the attribute types of the corresponding single entity attribute type set one by one; The second attribute extraction model is composed of a sixth preprocessing unit, a third BERT model, a third linear network and a third Softmax layer; the model structures of the first and third BERT models are the same and the model parameters are the same; the third linear network is implemented based on one or more fully connected layers.

7. The management system of the medical knowledge graph according to claim 4, wherein Each type of the medical examination images corresponds to a preset first analysis attribute set, and each of the first analysis attribute sets consists of one or more preset first analysis attributes; each of the first analysis attributes corresponds to a first attribute name and a first attribute value range; Each type of the medical examination audios corresponds to a preset second analysis attribute set, and each of the second analysis attribute sets consists of one or more preset second analysis attributes; each of the second analysis attributes corresponds to a second attribute name and a second attribute value range; Each type of the medical examination videos corresponds to a preset third analysis attribute set, and each of the third analysis attribute sets consists of one or more preset third analysis attributes; each of the third analysis attributes corresponds to a third attribute name and a third attribute value range; The number of the image attribute prediction models is equal to the number of the medical examination images, and the image attribute prediction models correspond to the medical examination images one by one; each of the image attribute prediction models is used to perform image attribute prediction on the input first medical image according to the attribute analysis requirements of the corresponding first analysis attribute set and output a corresponding first image attribute vector; the first medical image is a type of the medical examination images; the first image attribute vector includes a plurality of first attribute prediction data; the first attribute prediction data corresponds to the first attribute name of the first analysis attribute set one by one; The image attribute prediction model consists of a first feature extraction network and a first prediction network; the first feature extraction network is implemented based on a type of image encoder model; the first prediction network is implemented based on a type of non-linear regression prediction model; the image encoder model at least includes a CNN network and a residual network; the non-linear regression prediction model at least includes an MLP model, a decision tree model, an SVR model, and a random forest model; The number of the audio attribute prediction models is equal to the number of the medical examination audios, and the audio attribute prediction models correspond to the medical examination audios one by one; each of the audio attribute prediction models is used to perform audio attribute prediction on the input first medical audio according to the attribute analysis requirements of the corresponding second analysis attribute set and output a corresponding first audio attribute vector; the first medical audio is a type of the medical examination audios; the first audio attribute vector includes a plurality of second attribute prediction data; the second attribute prediction data corresponds to the second attribute name of the second analysis attribute set one by one; The audio attribute prediction model consists of a second feature extraction network and a second prediction network; the second feature extraction network is implemented based on a type of audio encoder model; the second prediction network is implemented based on a type of the non-linear regression prediction model; the audio encoder model at least includes a CNN network, an RNN network, a residual network, an LSTM model, and a Transformer model; The number of models of the video attribute prediction model is equal to the number of the medical examination videos, and the video attribute prediction models correspond to the medical examination videos one by one; each of the video attribute prediction models is used to perform video attribute prediction on the input first medical video according to the attribute analysis requirements of the corresponding third analysis attribute set and output a corresponding first video attribute vector; the first medical video is a type of the medical examination videos; the first video attribute vector includes a plurality of third attribute prediction data; the third attribute prediction data corresponds to the third attribute names of the third analysis attribute set one by one; The video attribute prediction model is composed of a third feature extraction network and a third prediction network; the third feature extraction network is implemented based on a type of video encoder model; the third prediction network is implemented based on a type of the non-linear regression prediction model; the video encoder model includes at least a 3D CNN network, a TCN network, an RNN network, an LSTM model, and a Transformer model; The input end of the third feature extraction network is connected to the input end of the video attribute prediction model, and the output end is connected to the input end of the third prediction network; the output end of the third prediction network is connected to the output end of the video attribute prediction model.

8. The management system of the medical knowledge graph according to claim 2, wherein, The graph processing module is specifically used for when updating the medical knowledge graph according to the first triple set and the graph database: Performing a round of traversal on all the first triples in the first triple set; and in the process of this round of traversal, taking the currently traversed first triple as the corresponding current triple; and taking the entity names, entity types, and entity attribute sets of the entities A and B in the current triple as the corresponding names A and B, types A and B, and attribute sets A and B; and taking the entity relationship R of the current triple as a corresponding A-B association relationship; And taking the first node record tables corresponding to the types A and B in the graph database as the corresponding node record tables A and B; and taking the first node record whose node name field in the node record tables A and B matches the corresponding names A and B as the corresponding matching node records A and B; and taking the first edge record in the first edge record table of the graph database whose association type field matches the A-B association relationship, and the head node identification field has a mapping relationship with the node identification field of the matching node record A, and the tail node identification field has a mapping relationship with the node identification field of the matching node record B as the corresponding matching edge record C; and identifying the matching node records A and B; If both of the matching node records A and B are not empty, then use the matching node records A and B as the corresponding current matching node records in sequence, and perform old node update processing based on the current matching node records and the current triple. When the matching edge record C is empty, perform single-edge addition processing based on the matching node records A and B and the A-B association relationship; If one of the matching node records A and B is empty, then use the non-empty matching node record A or B as the corresponding current matching node record, and use the entity A or B corresponding to the empty matching node record A or B as the corresponding current newly added entity. Perform old node update processing based on the current matching node record and the current triple, and perform single-node addition processing based on the current newly added entity to obtain the corresponding newly added node record for this time. Then, form a new pair of the matching node records A and B from the current matching node record and the newly added node record for this time, and perform single-edge addition processing based on the new matching node records A and B and the A-B association relationship; If both of the matching node records A and B are empty, then use the entities A and B of the current triple as the corresponding current newly added entities in sequence, and perform single-node addition processing based on the current newly added entities to obtain the corresponding newly added node record for this time. Then, form a new pair of the matching node records A and B from the two newly added node records corresponding to the entities A and B, and perform single-edge addition processing based on the new matching node records A and B and the A-B association relationship; After the current round of traversal of all the first triples in the first triple set is completed, construct the latest first node set based on all the first node record tables and the first attribute record tables on the graph database; and construct the latest first edge set based on the first edge record table on the graph database; And form the latest medical knowledge graph from the obtained first node set and first edge set.

9. The management system of the medical knowledge graph according to claim 2, wherein When the graph processing module is specifically used for regularly predicting whether there are still potential new connection relationships on the latest medical knowledge graph and performing new edge complement processing on the current medical knowledge graph based on the prediction result to obtain the latest medical knowledge graph: Regularly use the first node set of the latest medical knowledge graph as the corresponding node set V; the node set V consists of multiple nodes v, and each node v is assigned a unique integer value as the corresponding node index. The node v corresponds one-to-one with the first node in the first node set, and the node feature of the node v consists of the first node name, the first node type, and the first node attribute set of the corresponding first node; A virtual edge e with direction features and association relationship features is constructed between every two nodes v in the node set V, and all the obtained virtual edges e form a corresponding virtual edge set E; the direction features and the association relationship features of all the virtual edges e in the virtual edge set E are initialized to the corresponding invalid directions and invalid relationships; the edge features of each virtual edge e at least include the direction features and the association relationship features; the direction features include forward, reverse, and invalid directions; if the node with the larger node index among the two nodes v corresponding to each virtual edge e is denoted as the large-index node and the node with the smaller node index is denoted as the small-index node, then when the direction feature is forward, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the large-index node to the small-index node, when it is reverse, it indicates that the direction of the directed edge corresponding to the current virtual edge e is from the small-index node to the large-index node, and when it is an invalid direction, it indicates that the direction of the directed edge corresponding to the current virtual edge e is unknown; the association relationship features include all the medical entity association types and invalid relationships of the medical knowledge graph. Perform a round of traversal on all the first edges in the first edge set of the latest medical knowledge graph; during this round of traversal, take the currently traversed first edge as the corresponding current entity edge; take the virtual edge e in the virtual edge set E corresponding to the current entity edge as the corresponding current virtual edge; reset the direction feature of the current virtual edge based on the node index of the first node corresponding to the head and tail node identifiers of the current entity edge, and reset the association relationship feature of the current virtual edge based on the first association type of the current entity edge; after this round of traversal, record the latest virtual edge set E as the corresponding initial edge set E ini ; and the node set V and the initial edge set E ini form a graph structure data denoted as the corresponding initial graph; and input the initial graph into a preset knowledge graph edge prediction model, and the knowledge graph edge prediction model predicts the direction type and the association relationship type of each virtual edge e of the initial graph input to the model to obtain a corresponding predicted edge set E * ; the knowledge graph edge prediction model is a prediction model with a graph neural network model as the core encoder; the predicted edge set E * is composed of multiple predicted edges e * ; the predicted edge set E * of the predicted edge e * corresponds one-to-one with the virtual edge e of the initial edge set E ini ; the edge feature of each predicted edge e * at least includes the direction feature and the association relationship feature; Denote all the virtual edges \(e\) in the initial edge set \(E\) ini where the direction features and the association relationship features are the corresponding invalid directions and invalid relationships as the corresponding initial invalid edges; and denote each prediction edge \(e\) in the prediction edge set \(E\) * where the direction feature is not an invalid direction, the association relationship feature is not an invalid relationship, and which corresponds to an initial invalid edge * as the corresponding newly added valid edge; When the total number of the obtained newly added valid edges is not zero, a round of traversal is performed on all the obtained newly added valid edges; during this round of traversal, the currently traversed newly added valid edge is used as the corresponding currently added edge; and based on the direction feature of the currently added edge, the corresponding head and tail node markings are performed on the two first nodes corresponding to the currently added edge. The records of the two first nodes corresponding to the current head and tail nodes on the graph database are used as a pair of new matching node records A and B; the association relationship feature of the currently added edge is used as a new A-B association relationship; and a single-edge addition process is performed based on the matching node records A and B and the A-B association relationship. After the round of traversal of all the newly added valid edges is completed, all the newly added valid edges are added to the current medical knowledge graph to obtain the latest medical knowledge graph.

10. The management system of the medical knowledge graph according to claim 9, wherein The knowledge graph edge prediction model is used to perform prediction processing based on the direction type and association relationship type of each virtual edge e of the initial graph input to the model to obtain a corresponding predicted edge set E * ; The knowledge graph edge prediction model is composed of a graph embedding encoding module, a graph feature encoder, a direction prediction head, a relationship prediction head, and an output module; the graph feature encoder is implemented based on a type of graph neural network model, and the graph neural network model at least includes a GCN model, a GNN model, and an ApeGNN model; the direction prediction head and the relationship prediction head are each implemented based on an MLP model.

11. The management system of the medical knowledge graph according to claim 2, wherein The information retrieval module is specifically used for, when sending back the corresponding first retrieval graph obtained by performing knowledge graph retrieval processing on the basis of the first user identifier, the first keyword set, the user behavior database, and the graph database to the client: Take the first user record corresponding to the first user identifier in the first user record table of the user behavior database as the corresponding current user record; and extract the personalized feature field of the current user record as the corresponding current user feature set; And the first keyword vector corresponding to the first keyword set; and use a preset node feature encoding algorithm to encode the first keyword vector to obtain the corresponding first feature vector; and use the node feature vector fields of each of the first node feature vector records in the graph database as the corresponding first comparison vectors; and calculate the vector similarity between the first feature vector and each of the first comparison vectors to obtain the corresponding second similarity; and record the first node feature vector record corresponding to the second similarity that exceeds the preset second similarity threshold as the candidate feature record; and take the first node record having a mapping relationship with each of the candidate feature records as the corresponding preliminary selection node record; And identify the current user feature set; if the current user feature set is empty, record all the preliminary selection node records as the corresponding secondary selection node records; If the current user feature set is not empty, only record the preliminary selection node records that match each of the personalized features of the current user feature set as the corresponding secondary selection node records; And perform a round of traversal on all the secondary selection node records; And during this round of traversal, take the currently traversed secondary selection node record as the corresponding current node record; And take the first node corresponding to the current node record in the medical knowledge graph as the corresponding current node; And select a sub-graph centered on the current node in the medical knowledge graph as the corresponding first sub-graph; the maximum node distance between each of the first nodes in the first sub-graph and the current node is K, where K is a preset positive integer, and the node distance is the total number of other nodes between two directly or indirectly connected first nodes; And at the end of this round of traversal of all the secondary selection node records, send back the corresponding first retrieval graph composed of all the first sub-graphs obtained to the client.

12. The management system of the medical knowledge graph according to claim 2, wherein, The behavior collection module is specifically used for when performing behavior analysis on the user behavior database regularly: Regularly take each of the first user records in the first user record table of the user behavior database as the corresponding current user record; And extract the age field, the gender field and the user type field of the current user record as the corresponding user age, user gender and user type to form a corresponding user basic parameter; And use the first row record table corresponding to the current user record as the corresponding current behavior record table; extract the last X first row records in the current behavior record table and sort them in chronological order to form a corresponding first record sequence; extract the behavior time field, the behavior type, and the knowledge type field of each first row record in the first record sequence as a corresponding first time, first behavior type, and first entity type sequence to form a corresponding first behavior vector; and form a corresponding user behavior sequence by sorting all the obtained first behavior vectors in chronological order; X is a preset positive integer; And input the user basic parameters and the user behavior sequence into a preset user behavior prediction model, and the user behavior prediction model performs personalized feature prediction processing according to the user basic parameters and the user behavior sequence to obtain a corresponding prediction feature set; the user behavior prediction model is a classification prediction model implemented based on a deep learning model framework; when the prediction feature set is not empty, it consists of one or more prediction features, and the prediction feature is a type of the entity type; And when the current prediction feature set is not empty, reset the personalized feature field of the current user record based on the current prediction feature set.

13. The management system of the medical knowledge graph according to claim 12, wherein The user behavior prediction model is used to perform personalized feature prediction processing according to the user basic parameters and the user behavior sequence input into the model to obtain the corresponding prediction feature set; The user behavior prediction model is composed of a behavior feature extraction network, a basic feature encoding module, a feature fusion module, a personalized prediction head, and a prediction output module; the behavior feature extraction network at least includes an LSTM model and a bi-LSTM model; the personalized prediction head is implemented based on a type of multi-classification prediction model.

Citation Information

Patent Citations

  • Knowledge graph platform

    CN110795567A

  • Knowledge graph-based user hobbies and interests determination method and system

    CN112328645A

  • Knowledge graph completion method based on context awareness

    CN115905568A

  • Information processing apparatus, information processing method and program

    JP2023144408A