Similar case retrieval method and device, electronic equipment and computer readable medium

Through structured processing and the utilization of multimodal information, a similar case search model is generated, which solves the problem of low similar case search performance under the multimodal problem of medical data, and improves the ability and retrieval performance of processing complex case data.

CN120108752APending Publication Date: 2025-06-06HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510169869.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

Smart Images

  • Figure CN120108752A_ABST
    Figure CN120108752A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a similar case retrieval method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the steps of obtaining a pre-stored case data information group; performing structured processing on the case data information group to obtain each initial case node; generating each case node and each case relation graph; generating each sampling path information group; generating each coding vector group and each keyword feature vector; generating a heterogeneous graph of each sampling case; generating a similar case retrieval model; generating a case feature vector library; and in response to the received to-be-processed case data information, generating each retrieval result based on the to-be-processed case data information, the similar case retrieval model and the case feature vector library. According to the embodiment, the capability of processing complex case data and the similar case retrieval performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a similar case retrieval method, device, electronic device, and computer-readable medium. Background Art

[0002] Automated similar case retrieval can provide doctors with more references and treatment ideas, and improve the efficiency of diagnosis and treatment. The commonly used similar case retrieval method is to extract key attributes based on text information and calculate the similarity score under each attribute to screen similar cases.

[0003] However, when using the above method to search for similar cases, the following technical problems often occur:

[0004] The explosive growth of medical data, including data in multiple modalities such as text and images, relies only on text information to extract key attributes, resulting in less data support for similar case retrieval, weaker ability to process complex case data, and lower performance of similar case retrieval.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose similar case retrieval methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a similar case retrieval method, the method comprising: obtaining a pre-stored case data information group, wherein each case data information in the case data information group includes each case attribute; performing structured processing on the case data information group to obtain each initial case node, wherein each initial case node in the initial case nodes corresponds to a node type; based on a pre-constructed medical knowledge graph ontology library and the initial case nodes, generating each case node and each case relationship graph; based on the each case node and the case relationship graph, generating each sampling path information group; based on the case data information group and the initial case nodes, generating each sampling path information group; For each case node, generate each coding vector group and each keyword feature vector; based on the above-mentioned each sampling path information group, generate each sampling case heterogeneous graph; based on the above-mentioned each case node, each coding vector group, each keyword feature vector and each sampling case heterogeneous graph, generate a similar case retrieval model; based on the above-mentioned each case node and the above-mentioned similar case retrieval model, generate a case feature vector library, wherein the above-mentioned case feature vector library includes each case data information and each case feature vector; in response to receiving the case data information to be processed, generate each retrieval result based on the above-mentioned case data information to be processed, the above-mentioned similar case retrieval model and the above-mentioned case feature vector library.

[0009] In a second aspect, some embodiments of the present disclosure provide a similar case retrieval device, the device comprising: an acquisition unit, configured to acquire a pre-stored case data information group, wherein each case data information in the above case data information group includes various case attributes; a structured processing unit, configured to perform structured processing on the above case data information group to obtain various initial case nodes, wherein each of the above initial case nodes corresponds to a node type; a first generation unit, configured to generate various case nodes and various case relationship diagrams based on a pre-constructed medical knowledge graph ontology library and the above initial case nodes; a second generation unit, configured to generate various sampling path information groups based on the above case nodes and the above case relationship diagrams; a third generation unit, configured to generate various sampling path information groups based on the above case data information The fourth generation unit is configured to generate each sampling case heterogeneous graph based on the above-mentioned sampling path information groups and the above-mentioned each case node, generate each coding vector group and each keyword feature vector; the fifth generation unit is configured to generate a similar case retrieval model based on the above-mentioned each case node, the above-mentioned each coding vector group, the above-mentioned each keyword feature vector and the above-mentioned each sampling case heterogeneous graph; the sixth generation unit is configured to generate a case feature vector library based on the above-mentioned each case node and the above-mentioned similar case retrieval model, wherein the above-mentioned case feature vector library includes each case data information and each case feature vector; the seventh generation unit is configured to generate each retrieval result based on the above-mentioned case data information to be processed, the above-mentioned similar case retrieval model and the above-mentioned case feature vector library in response to receiving the case data information to be processed.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the similar case retrieval method of some embodiments of the present disclosure, the ability to process complex case data and the performance of similar case retrieval are improved. Specifically, the reason for the low ability to process complex case data and the low performance of similar case retrieval is that the explosive growth of medical data, including data of multiple modalities such as text and images, only relies on text information to extract key attributes, so that the data support for similar case retrieval is less, resulting in weak ability to process complex case data and low performance of similar case retrieval. Based on this, the similar case retrieval method of some embodiments of the present disclosure, first, obtains a pre-stored case data information group, wherein each case data information in the above case data information group includes each case attribute. Thus, various types of case data information can be obtained. Then, the above case data information group is structured to obtain each initial case node, wherein each initial case node in the above initial case nodes corresponds to a node type. Thus, the data can be structured to obtain each initial case node. Then, based on the pre-constructed medical knowledge graph ontology library and the above initial case nodes, each case node and each case relationship diagram are generated. Thus, each initial case node can be completed to obtain each case node and each case relationship graph between each case node attribute. Secondly, based on each case node and each case relationship graph, each sampling path information group is generated. Thus, each sampling path information group corresponding to each case relationship graph can be obtained. Then, based on the case data information group and each case node, each coding vector group and each keyword feature vector are generated. Thus, each coding vector group and each keyword feature vector corresponding to each case node can be obtained for the generation of subsequent similar case retrieval model. Then, based on each sampling path information group, each sampling case heterogeneous graph is generated. Thus, each sampling case heterogeneous graph generated by each sampling path information group can be obtained. Secondly, based on each case node, each coding vector group, each keyword feature vector and each sampling case heterogeneous graph, a similar case retrieval model is generated. Thus, a similar case retrieval model can be generated for case retrieval. Then, based on each case node and the similar case retrieval model, a case feature vector library is generated, wherein the case feature vector library includes each case data information and each case feature vector. Thus, each case feature vector corresponding to each type of case data information can be generated, thereby forming a case feature vector library. Finally, in response to receiving the case data information to be processed, each search result is generated based on the case data information to be processed, the similar case search model and the case feature vector library. Thus, each search result can be obtained by performing a similar case search on the newly received case data information to be processed.Because each case node and each coding vector group and each keyword feature vector corresponding to each case node are generated through various types of case data information, they are used to generate a subsequent similar case retrieval model. Therefore, the generated similar case retrieval model can make better use of multimodal information, thereby improving the data support for similar case retrieval and the ability to process complex case data, thereby improving the performance of similar case retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flow chart of some embodiments of the similar case retrieval method according to the present disclosure;

[0015] Figure 2 is a schematic diagram of the structure of some embodiments of the similar case retrieval device according to the present disclosure;

[0016] Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 The process 100 of some embodiments of the similar case retrieval method according to the present disclosure is shown. The similar case retrieval method comprises the following steps:

[0024] Step 101, obtaining a pre-stored case data information group.

[0025] In some embodiments, the execution subject (such as a computing device) of the similar case retrieval method can obtain a pre-stored case data information group. Among them, each case data information in the above case data information group can be text information or image information used to characterize the patient's case. Each case data information in the above case data information group includes various case attributes. Each case attribute of the above case attributes can be text information or image information used to characterize various components of the case. For example, the above case attributes can be but are not limited to "diagnostic medical records", "test results", "CT images", "ultrasound images", "X-ray films". The above execution subject can be a server for storing and managing the above case data information group.

[0026] Step 102, structurally process the case data information group to obtain each initial case node.

[0027] In some embodiments, the execution subject may perform structured processing on the case data information group to obtain each initial case node. Each of the initial case nodes may be a node corresponding to the case data information in the knowledge graph. Each of the initial case nodes may have a corresponding node type. The node type may be text information for characterizing the type of the initial case node. For example, the node type may be, but is not limited to, a "case node", a "clinical feature node", or a "treatment method node".

[0028] In some optional implementations of some embodiments, the execution subject may perform structured processing on the case data information group through the following steps to obtain each initial case node:

[0029] The first step is to perform the following steps for each case data information in the above case data information group:

[0030] The first sub-step is to classify the case data information including the case attributes to obtain text case attributes and image case attributes. Among them, each of the text case attributes can be used to characterize the case attribute as text information. Each of the image case attributes can be used to characterize the case attribute as image information. In practice, for each of the case attributes, first, the execution subject can determine the case attribute as a text case attribute in response to the case attribute being text information. Then, the execution subject can determine the case attribute as an image case attribute in response to the case attribute being image information.

[0031] In the second sub-step, for each of the above-mentioned text case attributes, perform the following steps:

[0032] Sub-step 1: Perform text extraction processing on the above-mentioned text case attributes to obtain text case information. The above-mentioned text case information can be text information converted into a structured table form after text extraction processing. In practice, the above-mentioned execution subject can perform natural language processing on the above-mentioned text case attributes to obtain text case information. The above-mentioned natural language processing can include but is not limited to: named entity recognition, information extraction.

[0033] Sub-step 2: Integrate the above text case attributes and the above text case information to obtain case text information. In practice, the above execution entity may combine the above text case attributes and the above text case information into case text information.

[0034] The third sub-step is to perform the following steps for each of the above-mentioned image case attributes:

[0035] Sub-step 1: Display the image case attributes. In practice, the execution subject may display the image case attributes via a display device connected to the execution subject. The display device may be a display.

[0036] Sub-step 2: receiving the image case information input by the target person. The image case information may be text information annotated by the target person on the image case attributes. The target person may be a technician who uses and manages the similar case retrieval method. In practice, the execution subject may receive the image case information input by the target person.

[0037] Sub-step three: integrating the above-mentioned image case attributes and the above-mentioned image case information to obtain case image information. In practice, the above-mentioned execution subject may combine the above-mentioned image case attributes and the above-mentioned image case information into case image information.

[0038] The second step is to integrate the obtained case text information and the obtained case image information to obtain an initial case node. In practice, the above execution entity can combine the obtained case text information and the obtained case image information into an initial case node.

[0039] Step 103, based on the pre-built medical knowledge graph ontology library and each initial case node, generate each case node and each case relationship graph.

[0040] In some embodiments, the execution subject may generate each case node and each case relationship graph based on the pre-built medical knowledge graph ontology library and each initial case node. Among them, the medical knowledge graph ontology library may be a database composed of a knowledge graph pre-built for storing and managing professional system knowledge in the medical field. The medical knowledge graph ontology library includes each medical knowledge graph. Each medical knowledge graph in each medical knowledge graph includes each medical knowledge entity information. The medical knowledge entity information may be, but is not limited to, entities such as clinical manifestations of diseases, examination methods, detection indicators, disease types, treatment methods, and drug types. The medical knowledge graph may be a knowledge graph for characterizing the relationship between each medical knowledge entity information. Each case node in each case node may be an initial case node supplemented by the medical knowledge graph ontology library. Each case relationship graph in each case relationship graph may be a graph structure composed of each case attribute included in the case node corresponding to the case relationship graph. There is a one-to-one correspondence between each case node and each case relationship graph.

[0041] In some optional implementations of some embodiments, the execution subject may generate each case node and each case relationship graph based on the pre-built medical knowledge graph ontology library and each initial case node through the following steps:

[0042] The first step is to perform the following steps for each of the above initial case nodes:

[0043] The first sub-step is to generate a case node based on the above-mentioned initial case node and the pre-stored standard case node, wherein the above-mentioned case node includes each case text information and each case image information. The above-mentioned standard case node can be a default text information used to characterize each case text information and each case image information in the case node. The above-mentioned standard case node can be text information pre-set by the target person, which is not limited here. In practice, first, the above-mentioned execution subject can add each case text information and each case image information included in the above-mentioned initial case node to each case text information and each case image information corresponding to the standard case node to update the above-mentioned standard case node. Then, the above-mentioned execution subject can determine the updated standard case node as the case node.

[0044] The second sub-step is to add the node type corresponding to the above initial case node to the above case node to update the above case node.

[0045] The third sub-step is to perform entity matching processing on the above-mentioned medical knowledge graph ontology library including each medical knowledge entity information and the above-mentioned case node including each case text information and each case image information, and obtain the case relationship diagram corresponding to the above-mentioned case node. In practice, the above-mentioned execution subject can perform entity matching processing on the above-mentioned case node including each case text information and each case image information with the above-mentioned medical knowledge graph ontology library including each medical knowledge entity information, and obtain the case relationship diagram corresponding to the above-mentioned case node. Among them, the algorithm used for the above-mentioned entity matching processing can be but is not limited to: Levenshtein distance algorithm, cosine similarity (Cosine Similarity).

[0046] Step 104, generating each sampling path information group based on each case node and each case relationship graph.

[0047] In some embodiments, the execution subject may generate each sampling path information group based on each case node and each case relationship graph. Each sampling path information included in each sampling path information group may be path information composed of each case attribute included in the case node corresponding to the case relationship graph.

[0048] In some optional implementations of some embodiments, the execution subject may generate each sampling path information group based on each case node and each case relationship graph through the following steps:

[0049] In the first step, for each of the above case relationship diagrams, perform the following steps:

[0050] The first sub-step is to generate case relationship topology information based on the case relationship graph. The case relationship topology information may be a topology structure diagram of the case relationship graph. In practice, the execution subject may determine the topology structure diagram of the case relationship graph as the case relationship topology information.

[0051] The second sub-step is to generate each meta-path information based on the case relationship topology information and the case relationship graphs, wherein each meta-path information in the meta-path information includes each meta-path node. Each meta-path node in the meta-path nodes may be a case text information or a case image information text information in each case text information and each case image information included in the case node. Each meta-path information in the meta-path information may be a path composed of each meta-path node. For example, the meta-path information may be "[patient→(implementation of examination relationship) examination item→(implementation of examination relationship_rev) patient]". Wherein, R_rev represents the reverse path of relationship R. The meta-path nodes may be "patient", "(implementation of examination relationship) examination item", "(implementation of examination relationship_rev) patient". In practice, the execution subject may process the case relationship topology information and the case relationship graphs through a meta-path algorithm to obtain each meta-path information. Wherein, the meta-path algorithm may be metapath.

[0052] The third sub-step is to generate sampling path information for each meta-path information in the above-mentioned meta-path information based on the above-mentioned meta-path information and the above-mentioned case nodes, wherein the above-mentioned sampling path information includes the various sampling nodes. Each of the above-mentioned sampling nodes can be a case text information or a case image information used to characterize the various case text information and the various case image information included in the above-mentioned case node. The above-mentioned sampling path information can be a path composed of the above-mentioned various sampling nodes. In practice, for each meta-path information in the above-mentioned meta-path information, the above-mentioned execution subject can determine the case text information or case image information corresponding to the above-mentioned meta-path node in the above-mentioned case nodes as a sampling node for each meta-path node in the above-mentioned meta-path information. Then, the above-mentioned execution subject can combine the determined sampling nodes into sampling path information.

[0053] The fourth sub-step is to determine the generated individual sampling path information as a sampling path information group.

[0054] Step 105, based on the case data information group and each case node, generate each encoding vector group and each keyword feature vector.

[0055] In some embodiments, the execution subject may generate each coding vector group and each keyword feature vector based on the case data information group and each case node. Each coding vector included in each coding vector group in each coding vector group may be a feature vector for characterizing each case attribute included in the case node. Each keyword feature vector in each keyword feature vector may be a feature vector for characterizing the case node. There is a one-to-one correspondence between each coding vector group and each keyword feature vector.

[0056] In some optional implementations of some embodiments, the execution subject may generate each encoding vector group and each keyword feature vector based on the case data information group and each case node through the following steps:

[0057] The first step is to perform the following steps for each of the above case nodes:

[0058] The first sub-step is to perform feature extraction processing on each case text information included in the above-mentioned case node to obtain each text encoding vector. Among them, each of the above-mentioned text encoding vectors can be a feature vector for characterizing the case text information. In practice, for each case text information in the above-mentioned case text information, the above-mentioned execution subject can extract the above-mentioned case text information through a feature extraction network to obtain a text encoding vector. Among them, the above-mentioned feature extraction network can be a neural network with case text information as input and text encoding vector as output. For example, the above-mentioned feature extraction network can be a CLIP (Contrastive Language-Image Pre-Training) network.

[0059] The second sub-step is to perform feature extraction processing on each case image information included in the case node to obtain each image coding vector. Each of the above image coding vectors can be a feature vector for characterizing the case image information. The generation method of each of the above image coding vectors can refer to the generation method of each of the above text coding vectors.

[0060] The third sub-step is to integrate the above-mentioned text encoding vectors and the above-mentioned image encoding vectors to obtain an encoding vector group. In practice, the above-mentioned execution subject can combine the above-mentioned text encoding vectors and the above-mentioned image encoding vectors into an encoding vector group.

[0061] The fourth sub-step is to determine the case data information corresponding to the case node in the case data information group as the target case data information.

[0062] The fifth sub-step is to perform word segmentation on the above-mentioned target case data information to obtain each target entry information. Among them, each target entry information in the above-mentioned each target entry information can be text information used to characterize the target case data information. In practice, the above-mentioned execution entity can perform word segmentation on the above-mentioned target case data information to obtain each target entry information. Among them, the above-mentioned word segmentation processing can include but is not limited to at least one of the following: word segmentation, removal of stop words (such as common but meaningless words like "de", "he", "zai", etc.), stemming or lemmatization.

[0063] The sixth sub-step is to perform the following steps on each target entry information in the above-mentioned each target entry information:

[0064] Sub-step one: Based on the pre-constructed high-frequency keyword library and the above-mentioned target entry information, generate a semantic similarity. Among them, the above-mentioned high-frequency keyword library can be a database used to store and manage high-frequency keywords. The above-mentioned high-frequency keyword library can include each high-frequency keyword. The above-mentioned high-frequency keyword can be a keyword used to characterize case data information. In practice, first, the above-mentioned execution entity can determine the similarity between the above-mentioned target entry information and each high-frequency keyword included in the above-mentioned high-frequency keyword library as each similarity to be determined. Then, the above-mentioned execution entity can determine the maximum similarity to be determined among the above-mentioned each similarity to be determined as the semantic similarity.

[0065] Sub-step two: In response to determining that the above-mentioned semantic similarity meets the preset similarity condition, determine the first preset vector information as the target vector information. Among them, the above-mentioned preset similarity condition can be that the above-mentioned semantic similarity is less than the preset similarity threshold. The above-mentioned preset similarity threshold can be a pre-determined value, which is not limited here. The above-mentioned first preset vector information can be information used to characterize that the above-mentioned semantic similarity meets the preset similarity condition. For example, the above-mentioned first preset vector information can be "0".

[0066] Sub-step three: In response to determining that the above-mentioned semantic similarity does not meet the preset similarity condition, determine the second preset vector information as the target vector information. Among them, the above-mentioned second preset vector information can be information used to characterize that the above-mentioned semantic similarity does not meet the preset similarity condition. For example, the above-mentioned second preset vector information can be "1".

[0067] The seventh sub-step is to generate a keyword feature vector based on the determined each target vector information. In practice, the above-mentioned execution entity can input the above-mentioned each target vector information into the first preset formula to obtain the keyword feature vector. Among them, the above-mentioned first preset formula can be X key =[x 1 ,x 2 ,…,x n . X keyis the keyword feature vector. 1 , X 2 , …, X n is each target vector information. n is the number of target vector information.

[0068] Step 106, generating each sampling case heterogeneous graph based on each sampling path information group.

[0069] In some embodiments, the execution subject may generate various sampled case heterogeneous graphs based on the various sampled path information groups, wherein each sampled case heterogeneous graph in the various sampled case heterogeneous graphs may be a heterogeneous graph composed of various case attributes included in the case node corresponding to the sampled path information group.

[0070] In the process of adopting technical solutions to solve the above technical problems, the following problems are often accompanied:

[0071] Retrieving similar cases only through case data information groups and individual case nodes results in weak connections between individual case nodes and weak ability to process complex case data, leading to low performance in similar case retrieval.

[0072] Faced with the above technical problems, we decided to adopt the following solutions:

[0073] In some optional implementations of some embodiments, the execution subject may generate each sampling case heterogeneous graph based on each sampling path information group through the following steps:

[0074] The first step is to perform the following steps on each sampling path information group in the above sampling path information groups:

[0075] The first sub-step is to determine the first sampling path information in the above sampling path information group as the target sampling path information.

[0076] The second sub-step is to determine each sampling node included in the target sampling path information as each heterogeneous graph node.

[0077] The third sub-step is to generate each heterogeneous graph edge based on each sampling node included in the target sampling path information. Each of the heterogeneous graph edges may be information for characterizing the edge composed of the sampling nodes. In practice, for each of the sampling nodes, the execution subject may determine the edge corresponding to the sampling node as a heterogeneous graph edge according to the target sampling path information.

[0078] The fourth sub-step is to integrate the above-mentioned heterogeneous graph nodes and the above-mentioned heterogeneous graph edges to obtain an initial sampled case heterogeneous graph. The above-mentioned initial sampled case heterogeneous graph may be a graph structure used to characterize the relationship between the above-mentioned heterogeneous graph nodes and the above-mentioned heterogeneous graph edges. In practice, the above-mentioned execution subject may combine the above-mentioned heterogeneous graph nodes and the above-mentioned heterogeneous graph edges into an initial sampled case heterogeneous graph.

[0079] The fifth sub-step is to determine the second sampling path information in the above sampling path information group as the sampling path information to be added.

[0080] The sixth sub-step is to execute the following loop steps based on the initial sampled case heterogeneous graph and the sampled path information to be added:

[0081] Sub-step 1: determining each sampling node included in the above-mentioned sampling path information to be added as each sampling node to be added.

[0082] Sub-step 2: Generate each sampling edge to be added based on each sampling node included in the above-mentioned sampling path information to be added. Each of the above-mentioned sampling edges to be added can be information for characterizing the edge composed of the above-mentioned sampling nodes. The generation method of the above-mentioned sampling edges to be added can refer to the generation method of the above-mentioned heterogeneous graph edges.

[0083] Sub-step three, for each of the above-mentioned sampling nodes to be added, in response to determining that the above-mentioned sampling node to be added and the initial sampling case heterogeneous graph meet the preset node adding condition, the above-mentioned sampling node to be added is added to the initial sampling case heterogeneous graph to update the initial sampling case heterogeneous graph. The above-mentioned preset node adding condition may be that the above-mentioned sampling node to be added does not exist in the above-mentioned initial sampling case heterogeneous graph.

[0084] Sub-step 4: for each of the above-mentioned sampling edges to be added, in response to determining that the above-mentioned sampling edge to be added and the updated initial sampling case heterogeneous graph meet the preset edge adding condition, the above-mentioned sampling edge to be added is added to the updated initial sampling case heterogeneous graph to update the updated initial sampling case heterogeneous graph. The above-mentioned preset edge adding condition may be that the above-mentioned sampling edge to be added does not exist in the above-mentioned initial sampling case heterogeneous graph.

[0085] Sub-step five, in response to determining that the sampling path information to be added is not the last sampling path information in the above-mentioned sampling path information group, the next sampling path information of the sampling path information to be added in the above-mentioned sampling path information group is used as the sampling path information to be added to update the sampling path information to be added, and based on the updated initial sampling case heterogeneous graph and the updated sampling path information to be added, the above-mentioned loop step is executed again.

[0086] Sub-step six: in response to determining that the sampling path information to be added is the last sampling path information in the above sampling path information group, determining the updated initial sampling case heterogeneous graph as the sampling case heterogeneous graph.

[0087] The above technical solution and its related contents, as an inventive point of the embodiment of the present disclosure, solve the problem that "only through the case data information group and each case node to retrieve similar cases, the connection between each case node is weak, the ability to process complex case data is weak, and the performance of similar case retrieval is low". The factors that lead to the low performance of similar case retrieval are often as follows: only through the case data information group and each case node to retrieve similar cases, the ability to process complex case data is weak, resulting in weak connection between each case node. If the above factors are solved, the performance of similar case retrieval can be improved. In order to achieve this effect, the present disclosure first performs the following steps for each sampling path information group in the above sampling path information groups: determine the first sampling path information in the above sampling path information group as the target sampling path information. Then, determine the sampling nodes included in the above target sampling path information as each heterogeneous graph node, and then, based on the sampling nodes included in the above target sampling path information, generate each heterogeneous graph edge. Secondly, the above heterogeneous graph nodes and the above heterogeneous graph edges are integrated to obtain the initial sampling case heterogeneous graph. Thus, the sampling case heterogeneous graph corresponding to the sampling path information group can be obtained. Then, the second sampling path information in the above-mentioned sampling path information group is determined as the sampling path information to be added. Then, based on the initial sampling case heterogeneous graph and the sampling path information to be added, the following loop steps are performed: each sampling node included in the above-mentioned sampling path information to be added is determined as each sampling node to be added. Secondly, based on the above-mentioned sampling nodes included in the sampling path information to be added, each sampling edge to be added is generated. Then, for each of the above-mentioned sampling nodes to be added, in response to determining that the above-mentioned sampling node to be added and the initial sampling case heterogeneous graph meet the preset node addition condition, the above-mentioned sampling node to be added is added to the initial sampling case heterogeneous graph to update the initial sampling case heterogeneous graph. Then, for each of the above-mentioned sampling edges to be added, in response to determining that the above-mentioned sampling edge to be added and the updated initial sampling case heterogeneous graph meet the preset edge addition condition, the above-mentioned sampling edge to be added is added to the updated initial sampling case heterogeneous graph to update the updated initial sampling case heterogeneous graph. In this way, the initial sampling case heterogeneous graph can be updated. Secondly, in response to determining that the sampling path information to be added is not the last sampling path information in the above sampling path information group, the next sampling path information of the sampling path information to be added in the above sampling path information group is used as the sampling path information to be added to update the sampling path information to be added, and the above loop step is performed again based on the updated initial sampling case heterogeneous graph and the updated sampling path information to be added. Thus, the initial sampling case heterogeneous graph can be continuously updated through the loop step.Finally, in response to determining that the sampling path information to be added is the last sampling path information in the above sampling path information group, the updated initial sampling case heterogeneous graph is determined as the sampling case heterogeneous graph. Thus, the updated sampling case heterogeneous graph can be obtained. Also, because the heterogeneous graph is constructed by each sampling path information group corresponding to each case node, the connection between each case node is strong, and the ability to process complex case data is improved, so that the performance of similar case retrieval using the sampling case heterogeneous graph is improved.

[0088] Step 107, generating a similar case retrieval model based on each case node, each coding vector group, each keyword feature vector and each sampled case heterogeneous graph.

[0089] In some embodiments, the execution subject may generate a similar case retrieval model based on the case nodes, the coding vector groups, the keyword feature vectors and the sampled case heterogeneous graphs. The similar case retrieval model may be a neural network model that takes the case nodes, the coding vector groups, the keyword feature vectors and the sampled case heterogeneous graphs as inputs and takes the case feature vectors as outputs. The case feature vectors may be feature vectors corresponding to each case node in the case nodes.

[0090] In some optional implementations of some embodiments, the execution subject may generate a similar case retrieval model based on the case nodes, the coding vector groups, the keyword feature vectors and the sampled case heterogeneous graphs through the following steps:

[0091] The first step is to perform the following steps for each of the above case nodes:

[0092] In the first sub-step, the coding vector group corresponding to the case node is determined as the target coding vector group.

[0093] In the second sub-step, the keyword feature vector corresponding to the case node is determined as the target keyword feature vector.

[0094] The third sub-step is to input the above-mentioned target coding vector group and the above-mentioned target keyword feature vector into a pre-trained feature aggregation network to obtain an embedded feature vector. The above-mentioned embedded feature vector may be a feature vector used to characterize the case node corresponding to the above-mentioned target coding feature vector group. The above-mentioned feature aggregation network may be a recurrent neural network that takes the above-mentioned target coding vector group and the above-mentioned target keyword feature vector as input and takes the embedded feature vector as output. In practice, the above-mentioned execution subject may input the above-mentioned target coding vector group and the above-mentioned target keyword feature vector into a feature aggregation network to obtain an embedded feature vector.

[0095] In the second step, based on the pre-stored classification information, the above-mentioned case nodes are classified and processed to obtain various similar case node groups, wherein each similar case node group in the above-mentioned various similar case node groups includes various case nodes. The above-mentioned classification information can be text information for classifying various node types included in each case node. The above-mentioned classification information can be text information predetermined by the target person, which is not limited here.

[0096] Step 3: For each of the above-mentioned similar case node groups, perform the following steps:

[0097] In the first sub-step, each embedded feature vector corresponding to each case node included in the above-mentioned similar case node group is determined as each target embedded feature vector.

[0098] The second sub-step is to input the above-mentioned each target embedded feature vector into a pre-trained similar feature aggregation network to obtain a similar feature vector. Among them, the above-mentioned similar feature vector can be a feature vector used to characterize the similar case node group corresponding to the above-mentioned each target embedded feature vector. The above-mentioned similar feature aggregation network can be a recurrent neural network that takes the above-mentioned each target embedded feature vector as input and the similar feature vector as output. In practice, the above-mentioned execution subject can input each target embedded feature vector into the similar feature aggregation network to obtain a similar feature vector.

[0099] The fourth step is to input the obtained respective similar feature vectors and the obtained respective embedded feature vectors into a pre-trained heterogeneous feature aggregation network to obtain a case feature vector. The case feature vector may be a feature vector used to characterize the case data information corresponding to the above respective similar feature vectors. The above heterogeneous feature aggregation network may be a recurrent neural network that takes the above respective similar feature vectors and the above respective embedded feature vectors as inputs and takes the case feature vector as output. In practice, the above execution entity may input the above respective similar feature vectors and the above respective embedded feature vectors into a heterogeneous feature aggregation network to obtain a case feature vector.

[0100] The fifth step is to integrate the feature aggregation network, the similar feature aggregation network and the heterogeneous feature aggregation network to obtain a similar case retrieval model. In practice, the execution entity can combine the feature aggregation network, the similar feature aggregation network and the heterogeneous feature aggregation network into a similar case retrieval model.

[0101] Step 108, generating a case feature vector library based on each case node and a similar case retrieval model.

[0102] In some embodiments, the execution subject may generate a case feature vector library based on the case nodes and the similar case retrieval model. The case feature vector library includes case data information and case feature vectors. The case feature vector may be a feature vector corresponding to each case node in the case nodes. The case data information and case feature vectors are stored in a one-to-one correspondence.

[0103] Step 109, in response to receiving the case data information to be processed, generating various retrieval results based on the case data information to be processed, a similar case retrieval model and a case feature vector library.

[0104] In some embodiments, the execution subject may generate various search results based on the case data information to be processed, the similar case search model and the case feature vector library in response to receiving the case data information to be processed. The case data information to be processed may be text information or image information of the case characterizing the patient. Each of the various search results may be case data information and case feature vectors in the case feature vector library.

[0105] In some optional implementations of some embodiments, the execution subject may generate various search results based on the case data information to be processed, the similar case search model and the case feature vector library through the following steps:

[0106] The first step is to perform structured processing on the above-mentioned case data information to obtain an initial case node to be processed. The above-mentioned initial case node to be processed can be a node corresponding to the above-mentioned case data information to be processed in the knowledge graph. The method for obtaining the above-mentioned initial case node to be processed can refer to the method for obtaining the above-mentioned initial case node.

[0107] The second step is to generate a pending case node based on the above medical knowledge graph ontology library and the above initial case node. The above pending case node can be obtained by referring to the above case node. The above pending case node includes various case attributes.

[0108] The third step is to generate a pending coding vector group and a pending keyword feature vector based on the pending case data information and the pending case node. The generation method of the pending coding vector group and the pending keyword feature vector can refer to the generation method of the coding vector group and the keyword feature vector. Each pending coding vector included in the pending coding vector group can be a feature vector for characterizing the attributes of each case included in the pending case node. The pending keyword feature vector can be a feature vector for characterizing the pending case node.

[0109] The fourth step is to generate a feature vector of the case to be processed based on the above-mentioned case node to be processed, the above-mentioned coding vector group to be processed, the above-mentioned keyword feature vector to be processed and the above-mentioned similar case retrieval model. The above-mentioned case feature vector to be processed can be a feature vector used to characterize the data information of the case to be processed corresponding to the above-mentioned case node to be processed. The generation method of the above-mentioned case feature vector to be processed can refer to the generation method of the above-mentioned case feature vector.

[0110] In the fifth step, the similarities between each case feature vector included in the case feature vector library and the case feature vector to be processed are determined as each target similarity.

[0111] The sixth step is to sort the above-mentioned target similarities to obtain a target similarity sequence. The above-mentioned target similarity sequence can be a sequence used to characterize that the above-mentioned target similarities are arranged in a preset order. The above-mentioned preset order can be a pre-set order, which is not limited here. For example, the above-mentioned preset order can be "from large to small". In practice, the above-mentioned execution entity can sort the above-mentioned target similarities according to the above-mentioned preset order to obtain a target similarity sequence.

[0112] In the seventh step, each target similarity satisfying the preset similarity condition in the target similarity sequence is determined as each similarity to be processed. The preset similarity condition may be the target similarities belonging to the first preset number in the target similarity sequence. The preset number may be a preset value, which is not limited here.

[0113] In the eighth step, each case data information corresponding to each of the above similarities to be processed in the above case feature vector library is determined as each search result.

[0114] Step 9: Display the above search results. In practice, the execution entity may display the above search results.

[0115] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the similar case retrieval method of some embodiments of the present disclosure, the ability to process complex case data and the performance of similar case retrieval are improved. Specifically, the reason for the low ability to process complex case data and the low performance of similar case retrieval is that the explosive growth of medical data, including data of multiple modalities such as text and images, only relies on text information to extract key attributes, so that the data support for similar case retrieval is less, resulting in weak ability to process complex case data and low performance of similar case retrieval. Based on this, the similar case retrieval method of some embodiments of the present disclosure, first, obtains a pre-stored case data information group, wherein each case data information in the above case data information group includes each case attribute. Thus, various types of case data information can be obtained. Then, the above case data information group is structured to obtain each initial case node, wherein each initial case node in the above initial case nodes corresponds to a node type. Thus, the data can be structured to obtain each initial case node. Then, based on the pre-constructed medical knowledge graph ontology library and the above initial case nodes, each case node and each case relationship diagram are generated. Thus, each initial case node can be completed to obtain each case node and each case relationship graph between each case node attribute. Secondly, based on each case node and each case relationship graph, each sampling path information group is generated. Thus, each sampling path information group corresponding to each case relationship graph can be obtained. Then, based on the case data information group and each case node, each coding vector group and each keyword feature vector are generated. Thus, each coding vector group and each keyword feature vector corresponding to each case node can be obtained for the generation of subsequent similar case retrieval model. Then, based on each sampling path information group, each sampling case heterogeneous graph is generated. Thus, each sampling case heterogeneous graph generated by each sampling path information group can be obtained. Secondly, based on each case node, each coding vector group, each keyword feature vector and each sampling case heterogeneous graph, a similar case retrieval model is generated. Thus, a similar case retrieval model can be generated for case retrieval. Then, based on each case node and the similar case retrieval model, a case feature vector library is generated, wherein the case feature vector library includes each case data information and each case feature vector. Thus, each case feature vector corresponding to each type of case data information can be generated, thereby forming a case feature vector library. Finally, in response to receiving the case data information to be processed, each search result is generated based on the case data information to be processed, the similar case search model and the case feature vector library. Thus, each search result can be obtained by performing a similar case search on the newly received case data information to be processed.Because each case node and each coding vector group and each keyword feature vector corresponding to each case node are generated through various types of case data information, they are used to generate a subsequent similar case retrieval model. Therefore, the generated similar case retrieval model can make better use of multimodal information, thereby improving the data support for similar case retrieval and the ability to process complex case data, thereby improving the performance of similar case retrieval.

[0116] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a similar case retrieval device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0117] like Figure 2 As shown, some embodiments of the similar case retrieval device 200 include: an acquisition unit 201, a structured processing unit 202, a first generation unit 203, a second generation unit 204, a third generation unit 205, a fourth generation unit 206, a fifth generation unit 207, a sixth generation unit 208 and a seventh generation unit 209. The acquisition unit 201 is configured to acquire a pre-stored case data information group, wherein each case data information in the above case data information group includes each case attribute; the structured processing unit 202 is configured to perform structured processing on the above case data information group to obtain each initial case node, wherein each initial case node in the above initial case nodes corresponds to a node type; the first generation unit 203 is configured to generate each case node and each case relationship graph based on the pre-constructed medical knowledge graph ontology library and the above initial case nodes; the second generation unit 204 is configured to generate each sampling path information group based on the above case nodes and the above case relationship graphs; the third generation unit 205 is configured to generate each sampling path information group based on the above case data information group and the above case nodes. coding vector groups and keyword feature vectors; the fourth generation unit 206 is configured to generate each sampling case heterogeneous graph based on the above-mentioned sampling path information groups; the fifth generation unit 207 is configured to generate a similar case retrieval model based on the above-mentioned each case node, the above-mentioned each coding vector group, the above-mentioned each keyword feature vector and the above-mentioned each sampling case heterogeneous graph; the sixth generation unit 208 is configured to generate a case feature vector library based on the above-mentioned each case node and the above-mentioned similar case retrieval model, wherein the above-mentioned case feature vector library includes each case data information and each case feature vector; the seventh generation unit 209 is configured to generate each retrieval result based on the above-mentioned case data information to be processed, the above-mentioned similar case retrieval model and the above-mentioned case feature vector library in response to receiving the case data information to be processed.

[0118] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be described in detail here.

[0119] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0120] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0121] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0122] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0123] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0124] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0125] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being installed in the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains a pre-stored case data information group, wherein each case data information in the above-mentioned case data information group includes various case attributes; performs structured processing on the above-mentioned case data information group to obtain various initial case nodes, wherein each of the above-mentioned initial case nodes corresponds to a node type; based on the pre-constructed medical knowledge graph ontology library and the above-mentioned initial case nodes, generates various case nodes and various case relationship diagrams; based on the above-mentioned various case nodes and the above-mentioned various case relationship diagrams, generates various sampling path information groups; based on the above-mentioned case Based on the data information group and the above-mentioned case nodes, generate each coding vector group and each keyword feature vector; based on the above-mentioned sampling path information group, generate each sampling case heterogeneous graph; based on the above-mentioned case nodes, the above-mentioned coding vector groups, the above-mentioned keyword feature vectors and the above-mentioned sampling case heterogeneous graph, generate a similar case retrieval model; based on the above-mentioned case nodes and the above-mentioned similar case retrieval model, generate a case feature vector library, wherein the above-mentioned case feature vector library includes each case data information and each case feature vector; in response to receiving the case data information to be processed, generate each retrieval result based on the above-mentioned case data information to be processed, the above-mentioned similar case retrieval model and the above-mentioned case feature vector library.

[0126] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0127] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0128] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The units described may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a structured processing unit, a first generation unit, a second generation unit, a third generation unit, a fourth generation unit, a fifth generation unit, a sixth generation unit, and a seventh generation unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring a pre-stored case data information group".

[0129] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0130] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A similar case retrieval method, comprising: Acquire a pre-stored case data information group, wherein each case data information in the case data information group includes each case attribute; Performing structured processing on the case data information group to obtain various initial case nodes, wherein each of the various initial case nodes corresponds to a node type; Based on the pre-built medical knowledge graph ontology library and the initial case nodes, generating each case node and each case relationship graph; Based on the case nodes and the case relationship graphs, generating each sampling path information group; Based on the case data information group and the case nodes, generate each coding vector group and each keyword feature vector; Based on the respective sampling path information groups, generating respective sampling case heterogeneous graphs; Generate a similar case retrieval model based on each case node, each coding vector group, each keyword feature vector and each sampled case heterogeneous graph; Based on the case nodes and the similar case retrieval model, a case feature vector library is generated, wherein the case feature vector library includes the case data information and the case feature vector; In response to receiving the case data information to be processed, various retrieval results are generated based on the case data information to be processed, the similar case retrieval model and the case feature vector library.

2. The method according to claim 1, wherein: The structural processing of the case data information group to obtain each initial case node includes: For each case data information in the case data information group, perform the following steps: Classify the case data information including each case attribute to obtain each text case attribute and each image case attribute; For each of the text case attributes, perform the following steps: Performing text extraction processing on the text case attributes to obtain text case information; Integrate the text case attributes and the text case information to obtain case text information; For each of the image case attributes, perform the following steps: Displaying the image case attributes; Receive image case information input by the target person; Integrate the image case attributes and the image case information to obtain case image information; The obtained case text information and the obtained case image information are integrated and processed to obtain an initial case node.

3. The method according to claim 1, wherein: The medical knowledge graph ontology library includes various medical knowledge graphs, and each of the various medical knowledge graphs includes various medical knowledge entity information; And the generation of each case node and each case relationship graph based on the pre-built medical knowledge graph ontology library and each initial case node includes: For each of the initial case nodes, perform the following steps: Generate a case node based on the initial case node and the pre-stored standard case node, wherein the case node includes each case text information and each case image information; Adding the node type corresponding to the initial case node to the case node to update the case node; Entity matching processing is performed on the medical knowledge graph ontology library including various medical knowledge entity information and the case nodes including various case text information and various case image information to obtain a case relationship graph corresponding to the case nodes.

4. The method according to claim 1, wherein: The generating each sampling path information group based on each case node and each case relationship graph includes: For each case relationship diagram in the case relationship diagrams, the following steps are performed: Based on the case relationship graph, generating case relationship topology information; Based on the case relationship topology information and the case relationship graphs, generating each meta-path information, wherein each meta-path information in the each meta-path information includes each meta-path node; For each of the meta-path information, based on the meta-path information and the case nodes, generating sampling path information, wherein the sampling path information includes the sampling nodes; The generated pieces of sampling path information are determined as a sampling path information group.

5. The method according to claim 3, wherein: The generating of each coding vector group and each keyword feature vector based on the case data information group and each case node includes: For each case node in the case nodes, perform the following steps: Performing feature extraction processing on each case text information included in the case node to obtain each text encoding vector; Performing feature extraction processing on each case image information included in the case node to obtain each image coding vector; Integrate the text encoding vectors and the image encoding vectors to obtain an encoding vector group; Determine the case data information corresponding to the case node in the case data information group as target case data information; Performing word segmentation processing on the target case data information to obtain each target term information; For each target term information in the target term information, the following steps are performed: Generate semantic similarity based on the pre-built high-frequency keyword library and the target term information; In response to determining that the semantic similarity satisfies a preset similarity condition, determining the first preset vector information as the target vector information; In response to determining that the semantic similarity does not satisfy a preset similarity condition, determining the second preset vector information as the target vector information; Based on the determined information of each target vector, a keyword feature vector is generated.

6. The method according to claim 5, wherein: The generating of a similar case retrieval model based on each case node, each coding vector group, each keyword feature vector and each sampled case heterogeneous graph comprises: For each case node in the case nodes, perform the following steps: Determine the encoding vector group corresponding to the case node as the target encoding vector group; Determine the keyword feature vector corresponding to the case node as the target keyword feature vector; Inputting the target encoding vector group and the target keyword feature vector into a pre-trained feature aggregation network to obtain an embedded feature vector; Based on the pre-stored classification information, the case nodes are classified to obtain the case node groups of the same type, wherein each case node group of the case node groups of the same type includes the case nodes; For each of the similar case node groups, the following steps are performed: Determine each embedded feature vector corresponding to each case node included in the same type of case node group as each target embedded feature vector; Inputting each target embedded feature vector into a pre-trained similar feature aggregation network to obtain a similar feature vector; Inputting each obtained homogeneous feature vector and each obtained embedded feature vector into a pre-trained heterogeneous feature aggregation network to obtain a case feature vector; The feature aggregation network, the similar feature aggregation network and the heterogeneous feature aggregation network are integrated to obtain a similar case retrieval model.

7. The method according to claim 1, wherein: The generating of various search results based on the case data information to be processed, the similar case search model and the case feature vector library includes: Performing structured processing on the data information of the case to be processed to obtain an initial case node to be processed; Based on the medical knowledge graph ontology library and the initial case node, generating a case node to be processed; Based on the case data information to be processed and the case node to be processed, generating a coding vector group to be processed and a keyword feature vector to be processed; Generate a feature vector of the case to be processed based on the case node to be processed, the coding vector group to be processed, the keyword feature vector to be processed and the similar case retrieval model; Determine the similarities between each case feature vector included in the case feature vector library and the case feature vector to be processed as each target similarity; Sorting the target similarities to obtain a target similarity sequence; Determine each target similarity in the target similarity sequence that meets a preset similarity condition as each similarity to be processed; Determine each case data information corresponding to each similarity to be processed in the case feature vector library as each search result; The respective search results are displayed.

8. A similar case retrieval device, comprising: An acquisition unit is configured to acquire a pre-stored case data information group, wherein each case data information in the case data information group includes each case attribute; A structured processing unit is configured to perform structured processing on the case data information group to obtain various initial case nodes, wherein each of the various initial case nodes corresponds to a node type; A first generating unit is configured to generate each case node and each case relationship graph based on a pre-built medical knowledge graph ontology library and each initial case node; A second generating unit is configured to generate each sampling path information group based on each case node and each case relationship graph; A third generating unit is configured to generate each encoding vector group and each keyword feature vector based on the case data information group and each case node; A fourth generating unit is configured to generate each sampling case heterogeneous graph based on each sampling path information group; A fifth generating unit is configured to generate a similar case retrieval model based on each case node, each coding vector group, each keyword feature vector and each sampled case heterogeneous graph; A sixth generating unit is configured to generate a case feature vector library based on each case node and the similar case retrieval model, wherein the case feature vector library includes each case data information and each case feature vector; The seventh generating unit is configured to generate various retrieval results based on the case data information to be processed, the similar case retrieval model and the case feature vector library in response to receiving the case data information to be processed.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Information query method and device, computer equipment and storage medium

    CN120613149A

  • An information query method and device, a computer device, and a storage medium

    CN120613149B

  • Multi-modal case database construction method and device, electronic equipment and medium

    CN120656630A