Article recognition method, device, computer device and storage medium

By constructing a medical knowledge graph and calculating the importance of medical words, we automatically identify key disease topics in medical articles, solving the high cost and low accuracy problems caused by manual labeling, and achieving efficient and accurate topic recognition.

CN112214580BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011213480.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-03
Publication Date
2025-07-18
Estimated Expiration
2040-11-03

AI Technical Summary

Technical Problem

Key disease topics for identifying medical articles in the prior art rely mainly on manual labeling, resulting in high labor costs and low accuracy of identification results.

Method used

By extracting medical word collections from medical articles, building medical knowledge graphs, calculating the importance of medical words, and using word vectors of key disease words to build key topic vectors, we automatically identify key disease topics in medical articles.

Benefits of technology

It realizes automated identification without human participation, saves labor costs and improves the accuracy of identification results, ensuring the accuracy of key disease topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112214580B_ABST
    Figure CN112214580B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an article recognition method, device, computer device, and storage medium based on artificial intelligence technology. The method includes: extracting a medical word set from a target medical article to be recognized, and constructing a medical knowledge graph of the target medical article by using multiple medical words in the medical word set; calculating the importance of the medical words recorded by each node based on the connection relationship between the nodes in the medical knowledge graph; selecting key disease words of the target medical article from the medical word set according to the importance of each medical word, and constructing a key topic vector of the target medical article by using the word vectors of the key disease words, where the key topic vector is used to indicate the key disease topic of the target medical article. The embodiment of the present invention can automatically identify the key disease topic of a medical article, effectively saving labor costs and improving the accuracy of the recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, specifically to the field of computer technologies, and particularly to an article recognition method, an article recognition device, a computer device, and a computer storage medium. Background Art

[0002] With the development of Internet technologies, Internet medical platforms (such as Internet medical APPs (applications), Internet medical websites, etc.) have emerged as the times require; Internet medical platforms can provide users with a large number of medical information articles (abbreviated as medical articles) so that users can obtain relevant medical information by reading these medical information articles. For a medical article, it usually includes the relevant content of one or more diseases, so that the medical article usually has one or more disease themes; then, in order to better classify, retrieve, and recommend a large number of medical articles provided by the Internet medical platform, it is usually necessary to identify the key disease themes (or called main disease themes) of these medical articles.

[0003] Currently, it is usually through the method of manual tagging to identify the key disease themes of each medical article. Such an identification method not only consumes a large amount of human costs, but also results in a low accuracy of the identification results due to the subjectivity of human identification. Based on this, how to better identify the key disease themes of medical articles has become a research hotspot. Summary of the Invention

[0004] Embodiments of the present invention provide an article recognition method, device, computer device, and storage medium, which can automatically identify the key disease themes of medical articles, effectively save human costs, and improve the accuracy of the recognition results.

[0005] On the one hand, embodiments of the present invention provide an article recognition method, and the method includes:

[0006] Extract a medical word set from a target medical article to be recognized, where the medical word set includes a plurality of medical words, and at least one of the plurality of medical words is a disease word;

[0007] Construct a medical knowledge graph of the target medical article by using the plurality of medical words, where the medical knowledge graph includes a plurality of nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article;

[0008] Calculate the importance of the medical words recorded by the respective nodes based on the connection relationships between the respective nodes in the medical knowledge graph;

[0009] Select the key disease words of the target medical article from the medical word set according to the importance of each medical word, and construct the key topic vector of the target medical article by using the word vectors of the key disease words. The key topic vector is used to indicate the key disease topic of the target medical article.

[0010] On the other hand, an embodiment of the present invention provides an article recognition device, which includes:

[0011] An extraction unit, configured to extract a medical word set from a target medical article to be recognized. The medical word set includes multiple medical words, and at least one of the multiple medical words is a disease word;

[0012] A construction unit, configured to construct a medical knowledge graph of the target medical article by using the multiple medical words. The medical knowledge graph includes multiple nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article;

[0013] A processing unit, configured to calculate the importance of the medical words recorded by each node based on the connection relationship between the nodes in the medical knowledge graph;

[0014] The processing unit is further configured to select the key disease words of the target medical article from the medical word set according to the importance of each medical word, and construct the key topic vector of the target medical article by using the word vectors of the key disease words. The key topic vector is used to indicate the key disease topic of the target medical article.

[0015] In one embodiment, when the processing unit is used to construct the key topic vector of the target medical article by using the word vectors of the key disease words, it may specifically be configured to:

[0016] Obtain the relevant non-disease words corresponding to the key disease words from the medical word set. The relevant non-disease words meet the following conditions: in the medical knowledge graph, the nodes for recording the relevant non-disease words are connected to the nodes for recording the key disease words;

[0017] Obtain the word vectors of the key disease words and the word vectors of the relevant non-disease words;

[0018] Fuse the word vectors of the key disease words and the word vectors of the relevant non-disease words to obtain the key topic vector of the target medical article.

[0019] In another embodiment, when the processing unit is used to select the key disease words of the target medical article from the medical word set according to the importance of each medical word, it may specifically be configured to:

[0020] Select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy; at least one candidate disease word is included in the multiple candidate keywords;

[0021] Select the candidate disease word with the greatest importance from the at least one candidate disease word as the key disease word of the target medical article.

[0022] In another implementation manner, the processing unit can also be used for:

[0023] Obtain the word vectors of each candidate keyword, and calculate the vector similarity between the word vectors of each candidate keyword and the key theme vector;

[0024] Select the article keywords of the target medical article from the multiple candidate keywords according to the vector similarity between the word vectors of each candidate keyword and the key theme vector; wherein, the vector similarity between the word vector of the article keyword and the key theme vector is greater than the similarity threshold;

[0025] Associate and store the target medical article and the article keywords so that the target medical article can be processed according to the article keywords.

[0026] In another implementation manner, when the processing unit is used to select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy, it can be specifically used for:

[0027] Select a preset number of medical words from the medical word set as the multiple candidate keywords of the target medical article in descending order of importance; or,

[0028] Select the medical words with importance greater than the importance threshold from the medical word set as the multiple candidate keywords of the target medical article.

[0029] In another implementation manner, the key theme vector is the dominant theme vector of the target medical article, and the key disease theme is the main disease theme of the target medical article; correspondingly, the processing unit can also be used for:

[0030] Select the reference disease word of the target medical article from the at least one candidate disease word, and the importance of the reference disease word is less than the importance of the key disease word;

[0031] Construct the subordinate theme vector of the target medical article by using the word vector of the reference disease word, and the subordinate theme vector of the target medical article is used to indicate the subordinate disease theme of the target medical article;

[0032] Associate and store the target medical article, the key topic vector, and the subordinate topic vector in a storage space so that when there is an article search request, article search processing is performed according to the key topic vector and the subordinate topic vector.

[0033] In another implementation, the storage space further includes at least one other medical article, and each other medical article has a corresponding dominant topic vector and a corresponding subordinate topic vector; correspondingly, the processing unit can also be used to:

[0034] When there is an article search request, obtain the information vector of the article search information carried by the article search request;

[0035] Obtain each medical article in the storage space and at least one topic vector of each medical article; wherein, each medical article has a recommended weight value, and at least one topic vector of each medical article includes the dominant topic vector and the subordinate topic vector of each medical article;

[0036] Calculate the matching degree between each topic vector of each medical article and the information vector respectively, and update the recommended weight value of each medical article according to the calculated matching degree;

[0037] Arrange the medical articles in descending order according to the updated recommended weight value of each medical article; select the medical article ranked first as the medical article to be recommended for output.

[0038] In another implementation, when the processing unit is used to update the recommended weight value of each medical article according to the calculated matching degree, it can specifically be used to:

[0039] For any medical article, determine the topic vector with the largest matching degree from at least one topic vector of the any medical article according to the matching degree between each topic vector of the any medical article and the information vector;

[0040] If the topic vector with the largest matching degree is the dominant topic vector of the any medical article, increase the recommended weight value of the any medical article;

[0041] If the topic vector with the largest matching degree is the subordinate topic vector of the any medical article, decrease the recommended weight value of the any medical article.

[0042] In another implementation, when the construction unit is used to construct the medical knowledge graph of the target medical article by using the multiple medical words, it can specifically be used to:

[0043] Construct an initial knowledge graph of the target medical article using the multiple medical terms. The initial knowledge graph includes multiple nodes, and each node records a medical term.

[0044] Select at least one pair of co-occurring word pairs from the multiple medical terms. A co-occurring word pair refers to a word pair composed of two medical terms that have a co-occurrence relationship in the target medical article.

[0045] According to the at least one pair of co-occurring word pairs, determine at least one node group from the initial knowledge graph. Any node group includes: two nodes that respectively record the two medical terms in a pair of co-occurring word pairs.

[0046] Connect the two nodes in each node group in the initial knowledge graph to obtain the medical knowledge graph of the target medical article.

[0047] In another implementation manner, when the construction unit is used to select at least one pair of co-occurring word pairs from the multiple medical terms, it may specifically be used to:

[0048] Determine the first distribution position of the first medical term in the target medical article. The first medical term is any medical term in the multiple medical terms.

[0049] Obtain a second medical term from the multiple medical terms according to the first distribution position of the first medical term. The distance between the second distribution position and the first distribution position of the second medical term in the target medical article is less than the position distance threshold.

[0050] Calculate the semantic distance value between the first medical term and the second medical term. The semantic distance value is used to indicate the semantic similarity between the first medical term and the second medical term.

[0051] If the semantic distance value between the first medical term and the second medical term is greater than the semantic threshold, determine that the first medical term and the second medical term have the co-occurrence relationship in the target medical article, and construct a pair of co-occurring word pairs using the first medical term and the second medical term.

[0052] In another implementation manner, when the processing unit is used to calculate the importance of the medical terms recorded by the respective nodes based on the connection relationships between the respective nodes in the medical knowledge graph, it may specifically be used to:

[0053] For the medical term recorded by any node, based on the connection relationships between the respective nodes in the medical knowledge graph, determine at least one associated node connected to the any node.

[0054] Calculate the semantic distance values between the medical term recorded by the any node and the medical terms recorded by the respective associated nodes.

[0055] Calculate the importance degree of the medical terms recorded in any node according to the calculated semantic distance value and the importance degrees of the medical terms recorded in each associated node.

[0056] In another embodiment, when the extraction unit is used to extract a medical term set from a target medical article to be recognized, it can be specifically used for:

[0057] Perform word segmentation on the target medical article to be recognized to obtain an initial word set, where the initial word set includes a plurality of initial words;

[0058] Screen out a plurality of intermediate words from the initial word set according to at least one medical dictionary, where the intermediate words refer to the initial words existing in the at least one medical dictionary;

[0059] Construct the medical term set of the target medical article by using the plurality of intermediate words.

[0060] In another embodiment, each initial word in the initial word set has a part of speech; correspondingly, when the extraction unit is used to screen out a plurality of intermediate words from the initial word set according to at least one medical dictionary, it can be specifically used for:

[0061] Screen out the initial words of the target part of speech from the initial word set, where the target part of speech includes at least one of the following: noun, verb, and adjective;

[0062] Screen out a plurality of intermediate words from the initial words of the target part of speech according to at least one medical dictionary.

[0063] In another embodiment, when the extraction unit is used to construct the medical term set of the target medical article by using the plurality of intermediate words, it can be specifically used for:

[0064] If there are intermediate words that meet the bonding condition among the plurality of intermediate words, perform bonding processing on the intermediate words that meet the bonding condition; add the phrases after the bonding processing and the intermediate words that have not been subjected to the bonding processing to the medical term set of the target medical article as medical terms;

[0065] If there are no intermediate words that meet the bonding condition among the plurality of intermediate words, add each intermediate word to the medical term set of the target medical article as a medical term;

[0066] Wherein, the bonding condition includes: being adjacent in the distribution position in the target medical article and existing in the same medical dictionary.

[0067] On the other hand, an embodiment of the present invention provides a computer device, the terminal includes an input interface and an output interface, and the terminal further includes:

[0068] A processor, adapted to implement one or more instructions; and,

[0069] A computer storage medium storing one or more instructions, the one or more instructions being adapted to be loaded and executed by the processor to perform the following steps:

[0070] Extract a medical word set from a target medical article to be recognized, the medical word set including a plurality of medical words, and at least including disease words among the plurality of medical words;

[0071] Construct a medical knowledge graph of the target medical article using the plurality of medical words, the medical knowledge graph including a plurality of nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article;

[0072] Calculate the importance of the medical words recorded by the respective nodes based on the connection relationships between the respective nodes in the medical knowledge graph;

[0073] Select key disease words of the target medical article from the medical word set according to the importance of each medical word, and construct a key theme vector of the target medical article using the word vectors of the key disease words, the key theme vector being used to indicate the key disease theme of the target medical article.

[0074] In another aspect, an embodiment of the present invention provides a computer storage medium storing one or more instructions, the one or more instructions being adapted to be loaded and executed by a processor to perform the following steps:

[0075] Extract a medical word set from a target medical article to be recognized, the medical word set including a plurality of medical words, and at least including disease words among the plurality of medical words;

[0076] Construct a medical knowledge graph of the target medical article using the plurality of medical words, the medical knowledge graph including a plurality of nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article;

[0077] Calculate the importance of the medical words recorded by the respective nodes based on the connection relationships between the respective nodes in the medical knowledge graph;

[0078] Select key disease words of the target medical article from the medical word set according to the importance of each medical word, and construct a key theme vector of the target medical article using the word vectors of the key disease words, the key theme vector being used to indicate the key disease theme of the target medical article.

[0079] In the embodiments of the present invention, for a target medical article to be recognized, multiple medical words can be first extracted from the target medical article, and a medical knowledge graph of the target medical article can be constructed using the multiple medical words. Since the medical words recorded by any two connected nodes in the medical knowledge graph have a co-occurrence relationship in the target medical article, and the medical words with more co-occurrence relationships are usually more important; therefore, based on the connection relationships between the nodes in the medical knowledge graph, the importance of the medical words recorded by each node can be calculated more accurately. Then, the key disease words of the target medical article can be selected from the medical word set according to the importance of each medical word, and a key theme vector for indicating the key disease theme of the target medical article can be constructed using the word vectors of the key disease words. It can be seen that the embodiments of the present invention can effectively improve the accuracy of the key disease words by improving the accuracy of the importance of each medical word, thereby improving the accuracy of the key disease theme; moreover, the entire theme recognition process does not require manual participation by the user, and the key disease theme of the target medical article can be automatically recognized, effectively saving labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0081] Figure 1a is a system architecture diagram of an identification system provided by an embodiment of the present invention;

[0082] Figure 1b is a system architecture diagram of another identification system provided by an embodiment of the present invention;

[0083] Figure 2 is a schematic flowchart of an article identification method provided by an embodiment of the present invention;

[0084] Figure 3a is a schematic diagram of sliding a sliding window on a target medical article provided by an embodiment of the present invention;

[0085] Figure 3b is a schematic diagram of constructing a medical knowledge graph of a target medical article using multiple medical words provided by an embodiment of the present invention;

[0086] Figure 4 is a schematic flowchart of an article identification method provided by another embodiment of the present invention;

[0087] Figure 5 is a model structure diagram of a word vector generation model provided by an embodiment of the present invention;

[0088] Figure 6 It is a schematic structural diagram of an article recognition device provided by an embodiment of the present invention;

[0089] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Specific embodiments

[0090] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0091] With the continuous development of Internet technology, AI (Artificial Intelligence) technology has also been better developed. The so-called AI refers to the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science; it mainly produces a new intelligent machine that can react in a way similar to human intelligence by understanding the essence of intelligence, so that the intelligent machine has multiple functions such as perception, reasoning and decision-making. Correspondingly, AI technology is an interdisciplinary subject, which mainly includes several major directions such as natural language processing (NLP) technology, computer vision technology (CV), speech technology, and machine learning / deep learning. Among them, natural language processing technology is an important direction in the field of computer science and artificial intelligence. The so-called natural language processing refers to a science that integrates linguistics, computer science and mathematics; natural language processing technology usually includes technologies such as knowledge graphs, text processing, semantic understanding, machine translation and robot question answering.

[0092] Based on the natural language processing technology in the above-mentioned artificial intelligence technology, the embodiments of the present invention propose an article recognition scheme for medical articles (i.e., medical information articles) to enable more accurate recognition of the key disease topics of medical articles. The key disease topics mentioned here refer to the most important or main disease topics among one or more disease topics in a medical article; the so-called disease topic refers to a word or sentence that can summarize the main content related to a certain disease involved in the medical article. In a specific implementation, the article recognition scheme can be executed by a computer device, which can be a terminal device (hereinafter referred to as a terminal) or a server; specifically, the general principle of the article recognition scheme is as follows: For the target medical article to be recognized, the computer device can obtain at least one medical dictionary. The so-called medical dictionary refers to a dictionary composed of a large number of medical words in the medical field published by an authoritative medical institution. The medical words mentioned here refer to words used to describe disease information, such as disease words composed of disease names, non-disease words composed of disease symptoms or drug names, etc. After obtaining the target medical article to be recognized, the computer device can extract a plurality of medical words included in the target medical article according to at least one medical dictionary. Then, according to the co-occurrence relationship of these multiple medical words in the target medical article, a medical knowledge graph can be constructed using these multiple medical words; and the key disease words of the target medical article can be selected from the multiple medical words according to the medical knowledge graph, so that the key disease topic of the target medical article can be determined based on the selected key disease words. Optionally, after determining the key disease topic of the target medical article, the computer device can also identify the article keywords of the target medical article based on the key disease topic.

[0093] To better implement the above article recognition scheme, the embodiments of the present invention also provide a related topic recognition system; the topic recognition system can at least include: a computer device 11 and a medical dictionary provider 12. The medical dictionary provider 12 mentioned here refers to a terminal, a client (i.e., an APP) or a server that can be used to provide at least one medical dictionary for the computer device 11. As can be seen from the foregoing, the computer device 11 can be a terminal or a server; then when the computer device 11 is a terminal, the system architecture of the topic recognition system can be seen in Figure 1a shown; when the computer device 11 is a server, the system architecture of the topic recognition system can be seen in Figure 1b shown. See Figure 1bAs shown, the subject recognition system in this case may further include at least one terminal 13, which is used to send the target medical article to be recognized to the computer device 11 (i.e., the server); moreover, the terminal 13 and the computer device 11 (i.e., the server) can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present invention do not limit this here. It should be noted that the above-mentioned terminal can be a smart phone, a tablet computer, a laptop computer, a smart watch, a desktop computer, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, etc.

[0094] As can be seen from the above description, through the subject recognition system and article recognition solution proposed by the embodiments of the present invention, the key disease subject of the target medical article can be automatically recognized; the entire subject recognition process does not require manual participation by the user, which can effectively save labor costs. Moreover, the subjectivity of manual recognition can be reduced through the automatic recognition method, avoiding the influence of the subjectivity of manual recognition on the accuracy of the recognition result, and effectively improving the accuracy of the recognition result.

[0095] Based on the above description, an article recognition method is proposed in the embodiments of the present invention, and this article main recognition method can be executed by the above-mentioned computer device. Please refer to Figure 2 , and this article recognition method may include the following steps S201 - S204:

[0096] S201, extract a medical word set from the target medical article to be recognized.

[0097] Studies have shown that for any medical article, the medical article usually includes a large number of medical terms, and one or more disease terms may be included in these large numbers of medical terms; the so-called disease terms refer to the terms composed of disease names, such as "gastric ulcer", "acute gastritis", "chronic gastritis", and so on. And optionally, one or more non-disease terms may also be included in these large numbers of medical terms; the so-called non-disease terms refer to the terms composed of other disease-related information other than disease names. The other disease-related information here may include but is not limited to: disease symptoms, disease affected parts, disease treatment means (disease treatment methods, disease treatment equipment), drug names of therapeutic drugs for treating diseases, and so on. Correspondingly, non-disease terms can be terms composed of disease symptoms (such as "abdominal distension and pain"), terms composed of disease affected parts (such as "stomach"), terms composed of disease treatment means (such as "gastroscope", "color Doppler ultrasound"), terms composed of drug names of therapeutic drugs (such as "Domperidone"), and so on.

[0098] Since these large numbers of medical terms are all used to describe the disease information (such as disease names, disease symptoms, disease affected parts, etc.) involved in the medical article, and the key disease theme of the medical article is used to summarize the relevant content of the main disease (i.e., the key disease) involved in the medical article; it can be seen that there is a certain correlation between the medical terms in the medical article and the key disease theme of the medical article. Therefore, the embodiments of the present invention propose an identification concept that can determine the key disease theme of the medical article through the medical terms included in the medical article. Based on this identification concept, when there is a need to identify the disease theme of the target medical article, the computer device can obtain the target medical article to be identified. Then, the medical term set can be extracted from the target medical article through step S201; the medical term set here may include multiple medical terms, and at least one disease term is included in the multiple medical terms. In the specific implementation process, the computer device can perform word segmentation processing on the target medical article and match the initial word set obtained by the word segmentation processing with at least one medical dictionary to match and extract the initial words included in the initial word set and located in at least one medical dictionary. Then, the medical terms in the target medical article are determined according to the initial words extracted by the matching, so as to construct the medical term set of the target medical article.

[0099] S202, construct a medical knowledge graph of the target medical article using multiple medical terms.

[0100] In an embodiment of the present invention, a medical knowledge graph of a target medical article may include multiple nodes; one node records one medical term, and the medical terms recorded by any two connected nodes have a co-occurrence relationship in the target medical article. That is, the nodes corresponding to two medical terms having a co-occurrence relationship in the target medical article are connected to each other in the medical knowledge graph. Among them, the co-occurrence relationship mentioned here may include any of the following meanings:

[0101] In one implementation, the co-occurrence relationship mentioned above may refer to the relationship that two medical terms appear in the sliding window simultaneously during the process of sliding a sliding window over the target medical article. Among them, the window length of the sliding window can be set according to an empirical value or the maximum sentence length in the target medical article (i.e., the number of characters included in the longest sentence); for example, if the maximum sentence length in the target medical article is 20 characters, the window length can be set to be less than or equal to 20 characters. For example, assume that a sliding window with a window length of 5 characters is used to slide over the target medical article; and assume that multiple medical terms include: medical term A, medical term B, medical term C, medical term D..., and the distribution positions of each medical term in the target medical article can be seen in Figure 3a the first figure shown. Since during this sliding process, medical term A and medical term B can appear in the sliding window simultaneously (as shown in the second figure in Figure 3a ), it can be considered that medical term A and medical term B have a co-occurrence relationship in the target medical article. Since during this sliding process, medical term B and medical term C can appear in the sliding window simultaneously (as shown in the third figure in Figure 3a ), it can be considered that medical term B and medical term C have a co-occurrence relationship in the target medical article. Since medical term C and medical term D cannot appear in the sliding window simultaneously (as shown in the fourth figure in Figure 3a ), it can be considered that medical term C and medical term D do not have a co-occurrence relationship in the target medical article, and so on.

[0102] In another implementation, the co-occurrence relationship mentioned above may refer to the relationship that two medical terms appear in the sliding window simultaneously and the semantic distance value between the two medical terms is greater than the semantic threshold during the process of sliding a sliding window over the target medical article. Among them, the semantic distance value between two medical terms can be calculated based on the word vectors of the two medical terms; the semantic distance value between two medical terms can be used to reflect the semantic similarity between the two medical terms, and the semantic distance value and the semantic similarity are proportional; that is, if the semantic distance value between two medical terms is larger, the semantic similarity between the two medical terms is greater. And the semantic threshold can be set according to an empirical value or business requirements. For example, assume that the semantic threshold is K and the semantic distance value between medical term A and medical term B is kAB The semantic distance value between medical term B and medical term C is k BC The semantic distance value between medical term C and medical term D is k CD ; and k AB < K, k BC > K, k CD < K. Still following the above Figure 3a example: Since the semantic distance value between medical term A and medical term B (i.e., k AB ) is less than the semantic threshold (K), although medical term A and medical term B can appear in the sliding window at the same time, the computer device may consider that medical term A and medical term B do not have a co-occurrence relationship in the target medical article. Since the semantic distance value between medical term B and medical term C (i.e., k BC ) is greater than the semantic threshold (K), and medical term B and medical term C can appear in the sliding window at the same time, the computer device may consider that medical term B and medical term C have a co-occurrence relationship in the target medical article, and so on. Thus, it can be seen that when determining whether two medical terms have a co-occurrence relationship in the target medical article, this embodiment not only considers the proximity of the distribution positions of the two medical terms in the target medical article through the sliding window, but also considers the semantic similarity between the two medical terms, which can effectively improve the accuracy of the co-occurrence relationship judgment, thereby improving the accuracy of the medical knowledge graph.

[0103] Based on the above description, during the specific implementation step S202, the computer device may first construct an initial knowledge graph of the target medical article using multiple medical terms; the initial knowledge graph includes multiple nodes, and each node records a medical term. Secondly, the computer device may select at least one pair of co-occurring word pairs from the multiple medical terms, and the co-occurring word pair refers to a word pair composed of two medical terms that have a co-occurrence relationship in the target medical article. Then, the computer device may traverse each pair of co-occurring word pairs; for the current co-occurring word pair being traversed, two nodes for recording the two medical terms in the current co-occurring word pair may be respectively connected in the initial knowledge graph; when all co-occurring word pairs have been traversed, the medical knowledge graph of the target medical article can be obtained. For example, see Figure 3bAs shown below: Suppose there are multiple medical terms, including: medical term A (recorded by node A), medical term B (recorded by node B), medical term C (recorded by node C), medical term D (recorded by node D), medical term E (recorded by node E)…; and there are a total of 5 co-occurrence word pairs among these multiple medical terms, which are: (medical term A, medical term B), (medical term A, medical term D), (medical term B, medical term C), (medical term B, medical term E), and (medical term D, medical term E). Then, the computer device can connect node A and node B, connect node A and node D, connect node B and node C, connect node B and node E, and connect node D and node E in the initial knowledge graph respectively, so as to obtain the medical knowledge graph of the target medical article.

[0104] S203. Calculate the importance of the medical terms recorded by each node based on the connection relationships between the nodes in the medical knowledge graph.

[0105] In a specific implementation, since the co-occurrence relationship refers to the relationship that two medical terms appear in the same sliding window at the same time, or refers to the relationship that two medical terms appear in the same sliding window and the semantic distance value between the two medical terms is greater than the semantic threshold; it can be known that the medical terms with more co-occurrence relationships have a higher appearance frequency in the target medical article. Then, it can be considered that the medical terms with more co-occurrence relationships are more important. Based on this, when the computer device executes step S203, it can count the number of relationship quantities of the co-occurrence relationships possessed by each medical term according to the connection relationships between the nodes in the medical knowledge graph; and according to the principle that the relationship quantity and the importance are positively correlated, determine the importance of each medical term according to the relationship quantity corresponding to each medical term.

[0106] Specifically, the relationship quantity corresponding to each medical term can be directly used as the importance of each medical term. Or, perform normalization processing on the relationship quantity corresponding to each medical term to obtain the importance of each medical term; where the normalization processing refers to the processing of mapping the relationship quantity to the range of 0 to 1. Or, the importance parameter can be used to perform weighted calculation on the relationship quantity corresponding to each medical term to obtain the importance of each medical term; where the importance parameter can be set according to empirical values or business requirements. For example, referring to Figure 3b the shown medical knowledge graph, node A is connected to node B, and node A and node D are connected. Then, it can be statistically determined that the number of relationship quantities of the co-occurrence relationships possessed by the medical term A recorded by node A is 2; then the relationship quantity can be directly used as the importance of medical term A (that is, the importance is 2), or the importance parameter (such as 1.5) is used to perform weighted calculation on the relationship quantity to obtain the importance of medical term A (that is, the importance is 3), and so on.

[0107] In another specific implementation, research has shown that if two medical terms have a co-occurrence relationship in a target medical article, since these two medical terms appear simultaneously, their importance levels usually affect each other. Based on this, when the computer device executes step S203, for the medical term recorded by any node, the importance level of the medical term of any node can be calculated by combining the importance levels of the medical terms recorded by the associated nodes connected to that any node, so as to improve the accuracy of the importance level. Specifically, for the medical term recorded by any node, at least one associated node connected to that any node can be determined based on the connection relationship between each node in the medical knowledge graph; then, according to the importance levels of the medical terms recorded by each associated node, the importance level of the medical term recorded by that any node is calculated.

[0108] Among them, the specific implementation manners of calculating the importance level of the medical term recorded by that any node according to the importance levels of the medical terms recorded by each associated node may include any one of the following:

[0109] Embodiment 1: The computer device can first calculate the initial value of the medical term recorded by any node according to the number of relationship quantities of the co-occurrence relationship possessed by the medical term recorded by any node. Secondly, the number of times that the medical term recorded by that any node and the medical terms recorded by each associated node appear in the sliding window can be respectively counted, and the counted times are respectively normalized to obtain the weight values of each associated node. For example, for the medical term A recorded by node A, node A has two associated nodes, node B and node D; if the number of times that the medical term A and the medical term B recorded by node B appear in the sliding window is 15 times, and the number of times that the medical term A and the medical term D recorded by node D appear in the sliding window is 5 times; then the weight value of node B is 15 / (15 + 5) = 0.75, and the weight value of node D is 5 / (15 + 5) = 0.25. After obtaining the weight values of each associated node, the importance levels of each associated node can be weighted and summed using the weight values of each associated node; for example, assuming the importance level of node B is 0.4 and the importance level of node D is 0.2, then 0.4×0.75 + 0.2×0.25 = 0.35 can be executed. Then, the sum operation can be performed on the value obtained by weighted summation and the initial value of the medical term recorded by that any node to obtain the importance level of the medical term recorded by that any node.

[0110] Embodiment 2: The computer device can also use the calculation formula shown in Equation 1.1 to calculate the importance level of the medical term recorded by that any node according to the importance levels of the medical terms recorded by each associated node. Among them, d in Equation 1.1 is the damping coefficient, and its value range is usually (0, 1). For example, d can be set to be equal to 0.85; V i represents the i-th node in the medical knowledge graph, WS(Vi ) represents the importance of the medical term recorded by the $i$-th node; In(V i ) represents the set of associated nodes corresponding to the $i$-th node, and $\in$ means belongs to; V j represents the $j$-th associated node in the set of associated nodes corresponding to the $i$-th node, that is, the $j$-th associated node connected to the $i$-th node; Out(V j ) represents the set of associated nodes corresponding to the $j$-th associated node connected to the $i$-th node, |Out(V j )| represents the number of associated nodes included in the set of associated nodes corresponding to the $j$-th associated node connected to the $i$-th node; WS(V j ) represents the importance of the medical term recorded by the $j$-th associated node connected to the $i$-th node.

[0111]

[0112] In another specific implementation, since the semantic distance value can be used to indicate the semantic similarity between two medical terms, and research has shown that for any medical term, if the semantic similarity between other medical terms and this any medical term is greater, then the influence of the importance of this other medical term on the importance of this any medical term is usually greater. Based on this, when the computer device executes step S203, for the medical term recorded by any node, it can calculate the importance of the medical term of any node by combining the importance of the medical terms recorded by the associated nodes connected to this any node, and the semantic distance values between the medical term recorded by any node and the medical terms recorded by each associated node, so as to further improve the accuracy of the importance. Specifically, for the medical term recorded by any node, at least one associated node connected to this any node can be determined first based on the connection relationship between the nodes in the medical knowledge graph. Then, the semantic distance values between the medical term recorded by this any node and the medical terms recorded by each associated node can be calculated; and the importance of the medical term recorded by any node can be calculated according to the calculated semantic distance values and the importance of the medical terms recorded by each associated node.

[0113] Among them, the specific implementation manners of calculating the importance of the medical term recorded by any node according to the calculated semantic distance values and the importance of the medical terms recorded by each associated node may include any one of the following:

[0114] Embodiment 1: The computer device can first calculate the initial value of the medical term recorded by any node according to the number of relationships of the medical terms recorded by any node. Secondly, the importance degrees of each associated node can be weighted and summed respectively using each semantic distance value. For example, for the medical term A recorded by node A, node A has two associated nodes, node B and node D; and the importance degree of node B is 0.4, and the importance degree of node D is 0.2. If the semantic distance value between the medical term A and the medical term B recorded by node B is k AB , and the semantic distance value between the medical term A and the medical term D recorded by node D is k AD ; then 0.4×k AB +0.2×k AD can be executed. Then, the sum operation can be performed on the value obtained by weighted summation and the initial value of the medical term recorded by any node to obtain the importance degree of the medical term recorded by any node.

[0115] Embodiment 2: The computer device can also use the calculation formula shown in Equation 1.2 to calculate the importance degree of the medical term recorded by any node according to the calculated semantic distance value and the importance degrees of the medical terms recorded by each associated node. Among them, w ji in Equation 1.2 represents the semantic distance value between the medical term recorded by the i-th node and the medical term recorded by the j-th associated node connected to the i-th node; V k represents the k-th associated node in the associated node set corresponding to the j-th associated node connected to the i-th node, that is, the k-th associated node connected to the j-th associated node connected to the i-th node; w jk represents the semantic distance value between the medical term recorded by the j-th associated node and the medical term recorded by the k-th associated node. It should be noted that the specific interpretations of other parameters in Equation 1.2 can be referred to the relevant descriptions of Equation 1.1 above and will not be elaborated here.

[0116]

[0117] S204. Select the key disease terms of the target medical article from the medical term set according to the importance degrees of each medical term, and construct the key theme vector of the target medical article using the word vectors of the key disease terms. The key theme vector is used to indicate the key disease theme of the target medical article.

[0118] Since the key disease topic of the target medical article is used to summarize the relevant content of the main diseases (i.e., key diseases) involved in the target medical article, after the computer device obtains the importance of each medical word, it can select the disease word with the highest importance from the medical word set as the key disease word of the target medical article according to the importance of each medical word. Then, the key topic vector of the target medical article can be constructed by using the word vector of the key disease word, and the key topic vector is used to indicate the key disease topic of the target medical article.

[0119] Among them, the specific implementation manner for the computer device to select the key disease word from the medical word set can be: first, screen out each disease word included in the medical word set, and then select the disease word with the highest importance from the screened disease words as the key disease word of the target medical article. Alternatively, the computer device can first select multiple candidate keywords of the target medical article from the medical word set according to the keyword selection strategy and the importance of each medical word. Among them, at least one candidate disease word is included in the multiple candidate keywords, and the keyword selection strategy can be set according to business requirements. For example, the keyword selection strategy can be set to indicate selecting a preset number of medical words from the medical word set as candidate keywords in the order of importance from large to small; or the keyword selection strategy can be set to indicate selecting medical words with importance greater than the importance threshold from the medical word set as candidate keywords, and so on.

[0120] In the embodiment of the present invention, for the target medical article to be recognized, multiple medical words can be first extracted from the target medical article, and a medical knowledge graph of the target medical article can be constructed by using the multiple medical words. Since there is a co-occurrence relationship between the medical words recorded by any two connected nodes in the medical knowledge graph, and the medical words with more co-occurrence relationships are usually more important; therefore, the importance of the medical words recorded by each node can be calculated more accurately based on the connection relationship between the nodes in the medical knowledge graph. Then, the key disease word of the target medical article can be selected from the medical word set according to the importance of each medical word, and the key topic vector for indicating the key disease topic of the target medical article can be constructed by using the word vector of the key disease word. It can be seen that the embodiment of the present invention can effectively improve the accuracy of the key disease word by improving the accuracy of the importance of each medical word, thereby improving the accuracy of the key disease topic; and, during the entire topic recognition process, no manual participation of the user is required, and the key disease topic of the medical article can be automatically recognized, effectively saving labor costs.

[0121] Based on the above Figure 2 related description of the embodiment of the article recognition method, the embodiment of the present invention also proposes a flow diagram of another more specific article recognition method, and the main article recognition method can be executed by the computer device mentioned above. Please refer to Figure 4, the article recognition method may include the following steps S401 - S409:

[0122] S401, extract a medical word set from the target medical article to be recognized.

[0123] In the specific implementation process, the computer device may first perform word segmentation on the target medical article to be recognized to obtain an initial word set; the initial word set includes multiple initial words, and each initial word in the initial word set has a part of speech. Specifically, the computer device may use an open - source word segmenter to perform full - segmentation word - segmentation processing on the target medical article to obtain the initial word set. The so - called full - segmentation word - segmentation processing means: performing word - segmentation on the target medical article using multiple segmentation forms; the multiple segmentation forms mentioned here may include but are not limited to: segmentation forms based on word frequency statistics, segmentation forms based on thesaurus matching, segmentation forms based on knowledge understanding, and so on. Among them, the segmentation form based on word frequency statistics is used to indicate: counting the frequency of any two characters appearing simultaneously in the article, and if the counted frequency is greater than the frequency threshold, then cutting the two characters into one word; the segmentation form based on thesaurus matching is used to indicate: matching the article with the entries in the thesaurus, and if a certain string in the article can be found in the thesaurus, then cutting the characters in the string into one word; the segmentation form based on knowledge understanding is used to indicate: combining syntactic and grammatical analysis to perform semantic analysis on the context of the article, and cutting the characters in the article based on the analysis results of the information provided by the context.

[0124] After obtaining the initial word set, the computer device can screen out multiple intermediate words from the initial word set according to at least one medical dictionary. The intermediate words mentioned here refer to the initial words existing in at least one medical dictionary, that is, the intermediate words refer to the words that exist in both at least one medical dictionary and the target medical article. In one implementation, the computer device can directly screen out multiple intermediate words from the initial word set according to at least one medical dictionary. Specifically, the computer device can traverse each initial word in the initial word set, and match the currently traversed initial word with at least one medical dictionary to detect whether the currently traversed initial word exists in at least one medical dictionary. If it exists, the currently traversed initial word is used as an intermediate word. In another implementation, since there may be some stop words in the initial word set, the so-called stop words refer to functional words that need to be filtered out and have no actual medical meaning. Specifically, the stop words may include, but are not limited to, determiners for expressing concepts such as location and quantity (such as "one", "these", "there", etc.), modal particles for expressing mood (such as "wow", "awesome", etc.), and so on. In this case, the computer device can first remove these stop words from the initial word set to update the initial word set, so that the updated initial word set (which can be represented by M) only retains words of a specified part of speech. Then, at least one medical dictionary is used to perform a matching process with the updated initial word set, which can effectively reduce the words to be matched, thereby effectively improving the matching efficiency and saving processing resources.

[0125] Based on this, when the computer device screens out multiple intermediate words from the initial word set according to at least one medical dictionary, it can first screen out the initial words of the target part of speech from the initial word set. The target part of speech mentioned here can be specified according to business requirements or empirical values, and it can include at least one of the following: nouns, verbs, and adjectives. Among them, nouns are used to represent the unified names of people, things, or abstract concepts; verbs are used to represent actions or states, and adjectives are used to describe or modify nouns or pronouns, and they are mainly used to represent the nature, state, characteristics, or attributes of people or things. Then, the computer device can screen out multiple intermediate words from the initial words of the target part of speech according to at least one medical dictionary. The specific implementation method is similar to the specific implementation method of the step of "directly screening out multiple intermediate words from the initial word set according to at least one medical dictionary" mentioned above, and will not be elaborated here.

[0126] After screening out multiple intermediate words, a medical word set (which can be denoted by M') of the target medical article can be constructed using the multiple intermediate words. In one implementation, the multiple intermediate words can be directly used to construct the medical word set of the target medical article; in this implementation, the medical words in the medical word set are these intermediate words, and the number of medical words is equal to the number of intermediate words. In another implementation, since there may be some intermediate words that are adjacent and have special meanings in the target medical article among the multiple screened intermediate words; for these intermediate words, they usually appear in the medical dictionary at the same time, and the word formed after they are glued together has more medical meaning than a single intermediate word. For example, for the intermediate words "lower abdomen" and "distending pain", they usually appear in the medical dictionary at the same time, and "lower abdomen distending pain" has more medical meaning than "lower abdomen" and "distending pain". In this case, the computer device can glue these intermediate words together and use the glued word as a medical word to improve the accuracy of subsequent topic recognition. Based on this, when the computer device constructs the medical word set of the target medical article using multiple intermediate words, it can determine whether there are intermediate words that meet the gluing conditions among the multiple intermediate words; the gluing conditions here can include: being adjacent in the distribution position in the target medical article and existing in the same medical dictionary. If there are intermediate words that meet the gluing conditions among the multiple intermediate words, the intermediate words that meet the gluing conditions are subjected to gluing processing; and the glued phrases and the intermediate words that have not been subjected to gluing processing are all added to the medical word set of the target medical article as medical words. If there are no intermediate words that meet the gluing conditions among the multiple intermediate words, each intermediate word can be added to the medical word set of the target medical article as a medical word.

[0127] S402, construct a medical knowledge graph of the target medical article using multiple medical words.

[0128] In a specific implementation, an initial knowledge graph of the target medical article can be constructed first using multiple medical words; the initial knowledge graph includes multiple nodes, and each node records a medical word. Secondly, at least one pair of co-occurring word pairs can be selected from the multiple medical words, and the co-occurring word pair refers to a word pair composed of two medical words that have a co-occurrence relationship in the target medical article. Then, at least one node group can be determined from the initial knowledge graph according to the at least one pair of co-occurring word pairs; among them, any node group can include: two nodes that respectively record the two medical words in a pair of co-occurring word pairs. Then, the two nodes in each node group can be connected respectively in the initial knowledge graph to obtain the medical knowledge graph of the target medical article.

[0129] S403, calculate the importance of the medical words recorded by each node based on the connection relationship between the nodes in the medical knowledge graph.

[0130] S404. According to the keyword selection strategy and the importance of each medical term, select multiple candidate keywords for the target medical article from the medical term set.

[0131] In a specific implementation, the computer device can select a preset number of medical terms from the medical term set as multiple candidate keywords for the target medical article in the order of decreasing importance. In another specific implementation, the computer device can select medical terms with importance greater than the importance threshold from the medical term set as multiple candidate keywords for the target medical article. Among them, at least one candidate disease term is included in the multiple candidate keywords. For the convenience of explanation, hereinafter, it is assumed that the multiple candidate keywords are a preset number of medical terms selected from the medical term set in the order of decreasing importance.

[0132] S405. Select the candidate disease term with the greatest importance from the at least one candidate disease term as the key disease term of the target medical article.

[0133] S406. Construct a key topic vector of the target medical article using the word vector of the key disease term, and this key topic vector is used to indicate the key disease topic of the target medical article.

[0134] In the specific implementation process, the computer device can first call the word vector generation model to obtain the word vector of the key disease term; the important significance of the word vector is to convert natural language into a vector that the computer device can understand, which can capture the context and semantics of the word, and thus can be used to measure the semantic similarity between words. The word vector generation model mentioned here can include but is not limited to: medical Word2vec model, medical bert model, medical glove model, and so on. Among them, the so-called medical Word2vec model refers to a word vector calculation model that can determine the word vector of a medical term by combining the context of the medical term; the so-called medical bert model refers to a word vector calculation model that can determine the word vector of a medical term by combining the meanings of the medical term in different contexts, so that the medical term can have the same word vector in different contexts; the so-called medical glove model refers to a word vector calculation model that uses the co-occurrence matrix and considers both local information and global information to determine the word vector.

[0135] For the convenience of explanation, hereinafter, it is assumed that the word vector generation model is the medical Word2vec model. See Figure 5As shown, the medical Word2vec model can be a three-layer neural network model, which may include an input layer, a hidden layer, and an output layer. Among them, the input layer is used to obtain multiple candidate keywords of the target medical article and construct a vocabulary with the candidate keywords; and to construct an initial sparse vector of the key disease word according to the position of the key disease word in the vocabulary. For example, if there are a total of 5 candidate keywords in the target medical article, then the vocabulary can include these 5 candidate keywords; assuming that the position of the key disease word in the vocabulary is the 3rd, then the initial sparse vector of the key disease word can be [0, 0, 1, 0, 0]. After obtaining the initial sparse vector, the input layer can also be used to pass the initial sparse vector to the hidden layer. The hidden layer can be used to convert the initial sparse vector passed from the input layer into a dense vector, thereby obtaining the word vector of the key disease word; correspondingly, the output layer is used to output the word vector obtained by the hidden layer. It should be noted that under the assumption of the Word2vec model, the input order of each word is not important. And this medical Word2vec model is obtained by a computer device pre-finetuning the benchmark Word2vec model with a large amount of medical information text; the benchmark Word2Vec model here refers to the Word2Vec model obtained by other devices training the model with a large amount of other information text (such as news information text). It can be seen that by obtaining the medical Word2vec model through model fine-tuning, not only can the resources consumed by the computer device for model training be effectively saved, but also the training duration can be shortened to improve the training efficiency.

[0136] After obtaining the word vector of the key disease word, the key topic vector of the target medical article can be constructed using the word vector of the key disease word. Specifically, the relevant non-disease words corresponding to the key disease word can be obtained from the medical word set; the so-called relevant non-disease words meet the following conditions: in the medical knowledge graph, the nodes used to record the relevant non-disease words are connected to the nodes used to record the key disease word. Secondly, the word vector of the key disease word and the word vectors of the relevant non-disease words can be obtained; it should be noted that any word vector mentioned here can be generated by the computer device by calling the word vector generation model, and any word vector can be a P-dimensional vector; the value of P can be set according to empirical values, for example, P = 200. Then, the word vector of the key disease word and the word vectors of the relevant non-disease words can be fused to obtain the key topic vector of the target medical article; specifically, the word vector of the key disease word and the word vectors of the relevant non-disease words can be added up dimension by dimension to obtain the key topic vector in the target medical article. For example, let the key disease word be "gastric ulcer", and its corresponding word vector be V(s 0 )=(s 0 1, s 0 2, …, s 0P ); The related non-disease words include: "lower abdominal pain", "gastroscopy", "color Doppler ultrasound", and "Domperidone", and the corresponding word vectors are as follows: V(s 1 )=(s 1 1, s 1 2,..., s 1 P ), V(s 2 )=(s 2 1, s 2 2,..., s 2 P ), V(s 3 )=(s 3 1, s 3 2,..., s 3 P ), and V(s 4 )=(s 4 1, s 4 2,..., s 4 P ). Then, by performing an accumulation operation on each dimension of the word vectors, the key topic vector can be obtained as V(S)=(s 0 1 + s 1 1 + s 2 1 + s 3 1 + s 4 1, s 0 2 + s 1 2 + s 2 2 + s 3 2 + s 4 2,..., s 0 P + s 1 P + s 2 P + s 3 P + s 4 P ).

[0137] S407, Obtain the word vectors of each candidate keyword, and calculate the vector similarity between the word vectors of each candidate keyword and the key topic vector.

[0138] S408, Select the article keywords of the target medical article from multiple candidate keywords according to the vector similarity between the word vectors of each candidate keyword and the key topic vector.

[0139] In the specific implementation process of steps S407 - S408, the computer device can first call the word vector generation model to represent each candidate keyword as a P - dimensional vector, thereby obtaining the word vectors (i.e., P - dimensional vectors) of each candidate keyword. Secondly, a vector similarity algorithm can be used to calculate the vector similarity between the word vectors of each candidate keyword and the key theme vector; the vector similarity algorithm here can include but is not limited to: cosine similarity algorithm, Hamming distance algorithm, Euclidean Distance algorithm, Manhattan distance algorithm, and so on. Then, according to the vector similarity between the word vectors of each candidate keyword and the key theme vector, candidate keywords with a vector similarity greater than the similarity threshold are selected from multiple candidate keywords as the article keywords of the target medical article. It should be noted that the similarity threshold mentioned here can be set according to business requirements or empirical values, for example, it can be set to 0.8, 0.7, etc.; and the article keywords selected through steps S407 - S408 may include the key disease words mentioned above, or may not include the key disease words mentioned above, and there is no restriction on this.

[0140] It can be seen that through steps S407 - S408, the vector similarity between the word vector of the article keyword and the key theme vector can be greater than the similarity threshold, thereby preventing candidate keywords irrelevant to the key disease theme from being selected as article keywords. For example, if the key disease theme of the target medical article is located in "gynecology" - related diseases, then the candidate keyword "orthopedics" irrelevant to this key disease theme can be prevented from being selected as the article keyword of the target medical article. It can be seen that through steps S407 - S408, the selected article keywords can better reflect the main content of the target medical article, improving the accuracy of the article keywords; and it can also make the article keywords of the target medical article all belong to the same key disease theme, ensuring the theme consistency among the article keywords.

[0141] S409, store the target medical article and the article keywords in an associated manner.

[0142] In specific implementation, after the computer device selects the article keywords of the target medical article through the above - mentioned steps S407 - S408, it can store the target medical article and the article keywords in an associated manner. Specifically, the computer device can directly add and store the article keywords to the keyword list related to the target medical article; or, the computer device can add and store the word vectors of each article keyword to the keyword list related to the target medical article. By storing the target medical article and the article keywords in an associated manner, subsequent business processing can be performed on the target medical article according to the article keywords; the business processing mentioned here can include but is not limited to: article classification processing, keyword reminder processing, keyword search processing, article recommended reading processing, and so on.

[0143] Among them, the article classification process refers to: obtaining relevant medical articles that have the same or similar article keywords as the target medical article, and classifying the obtained relevant medical articles and the target medical article into the same article set. The keyword reminder process refers to: when displaying the target medical article, synchronously displaying the article keywords of the target medical article. The keyword retrieval process refers to: when receiving the input keyword entered by the user, if the similarity between the input keyword and the article keyword is greater than the threshold, outputting the target medical article. The article recommended reading process refers to: when the user browses the current content, if it is recognized that the current content includes the article keyword, recommending the target medical article to the user.

[0144] It should be noted that the above-mentioned key topic vector is the dominant topic vector (i.e., the main topic vector) of the target medical article, and the key disease topic is the main disease topic of the target medical article. In an alternative embodiment, when the number of candidate disease words extracted from the target medical article is multiple, it indicates that the target medical article has multiple disease topics. In this case, the computer device can also select a reference disease word of the target medical article from at least one candidate disease word, and the importance of the reference disease word is less than the importance of the key disease word. Secondly, the word vector of the reference disease word can be used to construct the subordinate topic vector (i.e., the non-main topic vector) of the target medical article, and the subordinate topic vector of the target medical article is used to indicate the secondary disease topic of the target medical article. Then, the target medical article, the key topic vector, and the subordinate topic vector can be associated and stored in the storage space, so that when there is an article search request, article search processing can be performed according to the key topic vector and the subordinate topic vector.

[0145] Specifically, the computer device can also use the above method steps to calculate the leading topic vectors and subordinate topic vectors of other medical articles, and store the calculated topic vectors in the storage space; that is, the storage space also includes at least one other medical article, and each other medical article has a corresponding leading topic vector and a corresponding subordinate topic vector. When there is an article search request, the computer device can obtain the information vector of the article search information carried by the article search request. Secondly, it can obtain each medical article in the storage space and at least one topic vector of each medical article; among them, each medical article has a recommendation weight value, and at least one topic vector of each medical article includes the leading topic vector and the subordinate topic vector of each medical article. Then, it can calculate the matching degree between each topic vector of each medical article and the information vector respectively, and update the recommendation weight value of each medical article according to the calculated matching degree. Specifically, for any medical article, it can determine the topic vector with the largest matching degree from at least one topic vector of any medical article according to the matching degree between each topic vector of any medical article and the information vector. If the topic vector with the largest matching degree is the leading topic vector of any medical article, the recommendation weight value of any medical article can be increased to update the recommendation weight value of any medical article; if the topic vector with the largest matching degree is the subordinate topic vector of any medical article, the recommendation weight value of any medical article can be decreased to update the recommendation weight value of any medical article.

[0146] Based on the above update principle, after updating the recommendation weight values of each medical article, the computer device can sort each medical article in descending order according to the updated recommendation weight value of each medical article; select the medical article ranked first as the medical article to be recommended for output. It can be seen that the embodiment of the present invention can realize that for medical articles with multiple disease topics, according to the matching situation between the search information input by the user and the disease topic, the medical articles can be sorted with increased or decreased weights, so as to output the medical article that best matches the search information for the user; this can effectively improve the accuracy of article retrieval and thus improve the user experience.

[0147] In an embodiment of the present invention, for a target medical article to be recognized, multiple medical words can be first extracted from the target medical article, and a medical knowledge graph of the target medical article can be constructed using the multiple medical words. Since the medical words recorded by any two connected nodes in the medical knowledge graph have a co-occurrence relationship in the target medical article, and the medical words with more co-occurrence relationships are usually more important; therefore, based on the connection relationships between the nodes in the medical knowledge graph, the importance of the medical words recorded by each node can be calculated more accurately. Then, the key disease words of the target medical article can be selected from the medical word set according to the importance of each medical word, and a key theme vector for indicating the key disease theme of the target medical article can be constructed using the word vectors of the key disease words. It can be seen that the embodiment of the present invention can effectively improve the accuracy of the key disease words by improving the accuracy of the importance of each medical word, thereby improving the accuracy of the key disease theme; moreover, the entire theme recognition process does not require manual participation by the user, and the key disease theme of the medical article can be automatically recognized, effectively saving labor costs. In addition, after obtaining the key theme vector, the computer device can also purify the keywords of the target medical article based on the vector similarity between the candidate keywords and the key theme vector to obtain article keywords with theme consistency, which can effectively improve the accuracy of the article keywords.

[0148] Based on the description of the above article recognition method embodiment, an embodiment of the present invention also discloses an article recognition device, and the article recognition device can be a computer program (including program code) running in the above-mentioned computer device. The article recognition device can execute Figure 2 or Figure 4 the method shown. Please refer to Figure 6 , and the article recognition device can operate the following units:

[0149] An extraction unit 601, configured to extract a medical word set from a target medical article to be recognized, where the medical word set includes multiple medical words, and at least one of the multiple medical words is a disease word;

[0150] A construction unit 602, configured to construct a medical knowledge graph of the target medical article using the multiple medical words, where the medical knowledge graph includes multiple nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article;

[0151] A processing unit 603, configured to calculate the importance of the medical words recorded by each node based on the connection relationships between the nodes in the medical knowledge graph;

[0152] The processing unit is further configured to select key disease words of the target medical article from the medical word set according to the importance of each medical word, and construct a key topic vector of the target medical article by using the word vectors of the key disease words, where the key topic vector is used to indicate the key disease topic of the target medical article.

[0153] In one implementation, when the processing unit 603 is configured to construct a key topic vector of the target medical article by using the word vectors of the key disease words, it may be specifically configured to:

[0154] Obtain relevant non-disease words corresponding to the key disease words from the medical word set, where the relevant non-disease words meet the following conditions: in the medical knowledge graph, the nodes for recording the relevant non-disease words are connected to the nodes for recording the key disease words;

[0155] Obtain the word vectors of the key disease words and the word vectors of the relevant non-disease words;

[0156] Fuse the word vectors of the key disease words and the word vectors of the relevant non-disease words to obtain the key topic vector of the target medical article.

[0157] In another implementation, when the processing unit 603 is configured to select key disease words of the target medical article from the medical word set according to the importance of each medical word, it may be specifically configured to:

[0158] Select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to a keyword selection strategy; at least one candidate disease word is included in the multiple candidate keywords;

[0159] Select the candidate disease word with the greatest importance from the at least one candidate disease word as the key disease word of the target medical article.

[0160] In another implementation, the processing unit 603 may also be configured to:

[0161] Obtain the word vectors of each candidate keyword, and calculate the vector similarity between the word vectors of each candidate keyword and the key topic vector;

[0162] Select article keywords of the target medical article from the multiple candidate keywords according to the vector similarity between the word vectors of each candidate keyword and the key topic vector; wherein, the vector similarity between the word vector of the article keyword and the key topic vector is greater than a similarity threshold;

[0163] Associate and store the target medical article and the article keywords so that the target medical article can be processed according to the article keywords.

[0164] In another implementation manner, when the processing unit 603 is used to select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy, it can specifically be used for:

[0165] Select a preset number of medical words from the medical word set as multiple candidate keywords of the target medical article in descending order of importance; or,

[0166] Select medical words with importance greater than the importance threshold from the medical word set as multiple candidate keywords of the target medical article.

[0167] In another implementation manner, the key theme vector is the dominant theme vector of the target medical article, and the key disease theme is the main disease theme of the target medical article; correspondingly, the processing unit 603 can also be used for:

[0168] Select a reference disease word of the target medical article from the at least one candidate disease word, and the importance of the reference disease word is less than the importance of the key disease word;

[0169] Construct a subordinate theme vector of the target medical article using the word vector of the reference disease word, and the subordinate theme vector of the target medical article is used to indicate the secondary disease theme of the target medical article;

[0170] Associate and store the target medical article, the key theme vector, and the subordinate theme vector in the storage space so that when there is an article search request, article search processing can be performed according to the key theme vector and the subordinate theme vector.

[0171] In another implementation manner, the storage space further includes at least one other medical article, and each other medical article has a corresponding dominant theme vector and a corresponding subordinate theme vector; correspondingly, the processing unit 603 can also be used for:

[0172] When there is an article search request, obtain the information vector of the article search information carried by the article search request;

[0173] Obtain each medical article in the storage space, and at least one theme vector of each medical article; wherein, each medical article has a recommended weight value, and the at least one theme vector of each medical article includes the dominant theme vector and the subordinate theme vector of each medical article;

[0174] Calculate the matching degree between each topic vector of each medical article and the information vector respectively, and update the recommended weight value of each medical article according to the calculated matching degree;

[0175] Arrange the medical articles in descending order according to the updated recommended weight value of each medical article; select the medical article ranked first as the medical article to be recommended for output.

[0176] In another implementation manner, when the processing unit 603 is used to update the recommended weight value of each medical article according to the calculated matching degree, it may specifically be used for:

[0177] For any medical article, determine the topic vector with the largest matching degree from at least one topic vector of the medical article according to the matching degree between each topic vector of the medical article and the information vector;

[0178] If the topic vector with the largest matching degree is the dominant topic vector of the medical article, increase the recommended weight value of the medical article;

[0179] If the topic vector with the largest matching degree is the subordinate topic vector of the medical article, decrease the recommended weight value of the medical article.

[0180] In another implementation manner, when the construction unit 602 is used to construct the medical knowledge graph of the target medical article by using the multiple medical words, it may specifically be used for:

[0181] Construct an initial knowledge graph of the target medical article by using the multiple medical words, where the initial knowledge graph includes multiple nodes, and each node records a medical word;

[0182] Select at least one pair of co-occurring word pairs from the multiple medical words, where the co-occurring word pair refers to a word pair composed of two medical words having a co-occurrence relationship in the target medical article;

[0183] According to the at least one pair of co-occurring word pairs, determine at least one node group from the initial knowledge graph, and any node group includes: two nodes respectively recording the two medical words in a pair of co-occurring word pairs;

[0184] Connect the two nodes in each node group in the initial knowledge graph to obtain the medical knowledge graph of the target medical article.

[0185] In another implementation manner, when the construction unit 602 is used to select at least one pair of co-occurring word pairs from the multiple medical words, it may specifically be used for:

[0186] Determine the first distribution position of the first medical term in the target medical article, where the first medical term is any one of the multiple medical terms;

[0187] Obtain a second medical term from the multiple medical terms according to the first distribution position of the first medical term, where the distance between the second distribution position and the first distribution position of the second medical term in the target medical article is less than the position distance threshold;

[0188] Calculate the semantic distance value between the first medical term and the second medical term, where the semantic distance value is used to indicate the semantic similarity between the first medical term and the second medical term;

[0189] If the semantic distance value between the first medical term and the second medical term is greater than the semantic threshold, determine that the first medical term and the second medical term have the co-occurrence relationship in the target medical article, and construct a pair of co-occurrence word pairs using the first medical term and the second medical term.

[0190] In another implementation manner, when the processing unit 603 is used to calculate the importance of the medical terms recorded by each node based on the connection relationship between the nodes in the medical knowledge graph, it may specifically be used for:

[0191] For the medical term recorded by any node, based on the connection relationship between the nodes in the medical knowledge graph, determine at least one associated node connected to the any node;

[0192] Calculate the semantic distance value between the medical term recorded by the any node and the medical terms recorded by each associated node;

[0193] According to the calculated semantic distance value and the importance of the medical terms recorded by each associated node, calculate the importance of the medical term recorded by the any node.

[0194] In another implementation manner, when the extraction unit 601 is used to extract a medical term set from the target medical article to be recognized, it may specifically be used for:

[0195] Perform word segmentation on the target medical article to be recognized to obtain an initial word set, where the initial word set includes multiple initial words;

[0196] Screen out multiple intermediate words from the initial word set according to at least one medical dictionary, where the intermediate words refer to the initial words existing in the at least one medical dictionary;

[0197] Construct the medical term set of the target medical article using the multiple intermediate words.

[0198] In another implementation, each initial word in the initial word set has a part of speech; correspondingly, when the extraction unit 601 is used to screen out a plurality of intermediate words from the initial word set according to at least one medical dictionary, it can be specifically used for:

[0199] Screen out the initial words of the target part of speech from the initial word set, where the target part of speech includes at least one of the following: noun, verb, and adjective;

[0200] Screen out a plurality of intermediate words from the initial words of the target part of speech according to at least one medical dictionary.

[0201] In another implementation, when the extraction unit 601 is used to construct the medical word set of the target medical article by using the plurality of intermediate words, it can be specifically used for:

[0202] If there are intermediate words that meet the bonding condition among the plurality of intermediate words, perform bonding processing on the intermediate words that meet the bonding condition; add the phrases after the bonding processing and the intermediate words that have not been bonded as medical words to the medical word set of the target medical article;

[0203] If there are no intermediate words that meet the bonding condition among the plurality of intermediate words, add each intermediate word as a medical word to the medical word set of the target medical article;

[0204] Among them, the bonding condition includes: adjacent distribution positions in the target medical article and existing in the same medical dictionary.

[0205] According to an embodiment of the present application, Figure 2 or Figure 4 Each step involved in the method shown can be Figure 6 executed by each unit in the article recognition device shown. For example, Figure 2 The steps S201-S202 shown in can be respectively executed by Figure 6 the extraction unit 601 and the construction unit 602 shown in, and the steps S203-S204 can both be executed by Figure 6 the processing unit 603 shown in. Another example, Figure 4 The steps S401-S402 shown in can be respectively executed by Figure 6 the extraction unit 601 and the construction unit 602 shown in, and the steps S403-S409 can all be executed by Figure 6 the processing unit 603 shown in, and so on.

[0206] According to another embodiment of the present application, Figure 6Each unit in the article recognition device shown can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with more specific functions to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, based on the article recognition device, other units can also be included. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.

[0207] According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in Figure 2 or Figure 4 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct the article recognition device shown in Figure 6 and to implement the article recognition method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0208] For the target medical article to be recognized in the embodiments of the present invention, multiple medical words can be first extracted from the target medical article, and a medical knowledge graph of the target medical article can be constructed using the multiple medical words. Since there is a co-occurrence relationship in the target medical article between any two connected nodes recorded in the medical knowledge graph, and the medical words with more co-occurrence relationships are usually more important; thus, based on the connection relationships between the nodes in the medical knowledge graph, the importance degrees of the medical words recorded by each node can be calculated more accurately. Then, the key disease words of the target medical article can be selected from the medical word set according to the importance degrees of the respective medical words, and a key theme vector for indicating the key disease theme of the target medical article can be constructed using the word vectors of the key disease words. It can be seen that the embodiments of the present invention can effectively improve the accuracy of the key disease words by improving the accuracy of the importance degrees of the respective medical words, thereby improving the accuracy of the key disease theme; moreover, during the entire theme recognition process, there is no need for user manual participation, and the key disease theme of the target medical article can be automatically recognized, effectively saving labor costs.

[0209] Based on the descriptions of the above method embodiments and device embodiments, the embodiments of the present invention also provide a computer device. Please refer to Figure 7, the computer device may at least include a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. Among them, the processor 701, the input interface 702, the output interface 703, and the computer storage medium 704 in the computer device may be connected through a bus or other means. The computer storage medium 704 may be stored in the memory of the computer device. The computer storage medium 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is adapted to implement one or more instructions, specifically adapted to load and execute one or more instructions to implement the corresponding method flow or corresponding function.

[0210] In one embodiment, the processor 701 described in the embodiments of the present invention may be used to perform a series of article recognition processes, specifically including: extracting a medical word set from a target medical article to be recognized, where the medical word set includes a plurality of medical words, and at least one of the plurality of medical words includes a disease word; constructing a medical knowledge graph of the target medical article using the plurality of medical words, where the medical knowledge graph includes a plurality of nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article; calculating the importance of the medical words recorded by each node based on the connection relationship between the nodes in the medical knowledge graph; selecting the key disease words of the target medical article from the medical word set according to the importance of each medical word, and constructing a key topic vector of the target medical article using the word vectors of the key disease words, where the key topic vector is used to indicate the key disease topic of the target medical article, and so on.

[0211] An embodiment of the present invention further provides a computer storage medium (Memory). The computer storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer storage medium provides a storage space, and the operating system of the computer device is stored in the storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor 701 are stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.

[0212] In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor 701 to implement the corresponding steps of the method in the above-mentioned embodiment of the article recognition method; in a specific implementation, one or more instructions in the computer storage medium are loaded and executed by the processor 701 to perform the following steps:

[0213] Extract a medical word set from the target medical article to be recognized. The medical word set includes multiple medical words, and at least one of the multiple medical words is a disease word.

[0214] Construct a medical knowledge graph of the target medical article using the multiple medical words. The medical knowledge graph includes multiple nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article.

[0215] Based on the connection relationships between the nodes in the medical knowledge graph, calculate the importance degrees of the medical words recorded by the respective nodes.

[0216] Select the key disease words of the target medical article from the medical word set according to the importance degrees of the respective medical words, and construct a key topic vector of the target medical article using the word vectors of the key disease words. The key topic vector is used to indicate the key disease topic of the target medical article.

[0217] In one implementation manner, when constructing the key topic vector of the target medical article using the word vectors of the key disease words, the one or more instructions can be specifically loaded and executed by the processor 701:

[0218] Obtain relevant non-disease words corresponding to the key disease words from the medical word set, where the relevant non-disease words meet the following conditions: in the medical knowledge graph, the nodes for recording the relevant non-disease words are connected to the nodes for recording the key disease words;

[0219] Obtain the word vectors of the key disease words and the word vectors of the relevant non-disease words;

[0220] Fuse the word vectors of the key disease words and the word vectors of the relevant non-disease words to obtain the key topic vector of the target medical article.

[0221] In another implementation, when selecting the key disease words of the target medical article from the medical word set according to the importance of each medical word, the one or more instructions can be loaded and specifically executed by the processor 701:

[0222] Select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy; at least one candidate disease word is included in the multiple candidate keywords;

[0223] Select the candidate disease word with the greatest importance from the at least one candidate disease word as the key disease word of the target medical article.

[0224] In another implementation, the one or more instructions can also be loaded and specifically executed by the processor 701:

[0225] Obtain the word vectors of each candidate keyword and calculate the vector similarity between the word vectors of each candidate keyword and the key topic vector;

[0226] Select the article keywords of the target medical article from the multiple candidate keywords according to the vector similarity between the word vectors of each candidate keyword and the key topic vector; wherein, the vector similarity between the word vector of the article keyword and the key topic vector is greater than the similarity threshold;

[0227] Associate and store the target medical article and the article keywords so that the target medical article can be processed according to the article keywords.

[0228] In another implementation, when selecting multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy, the one or more instructions can be loaded and specifically executed by the processor 701:

[0229] Select a preset number of medical terms from the medical term set as multiple candidate keywords for the target medical article in descending order of importance; or,

[0230] Select medical terms with importance greater than the importance threshold from the medical term set as multiple candidate keywords for the target medical article.

[0231] In another implementation manner, the key theme vector is the dominant theme vector of the target medical article, and the key disease theme is the main disease theme of the target medical article; correspondingly, the one or more instructions can also be loaded and specifically executed by the processor 701:

[0232] Select a reference disease term for the target medical article from the at least one candidate disease term, and the importance of the reference disease term is less than the importance of the key disease term;

[0233] Construct a subordinate theme vector of the target medical article using the word vector of the reference disease term, and the subordinate theme vector of the target medical article is used to indicate the secondary disease theme of the target medical article;

[0234] Associate and store the target medical article, the key theme vector, and the subordinate theme vector in a storage space so that when there is an article search request, article search processing is performed according to the key theme vector and the subordinate theme vector.

[0235] In another implementation manner, the storage space further includes at least one other medical article, and each other medical article has a corresponding dominant theme vector and a corresponding subordinate theme vector; correspondingly, the one or more instructions can also be loaded and specifically executed by the processor 701:

[0236] When there is an article search request, obtain the information vector of the article search information carried by the article search request;

[0237] Obtain each medical article in the storage space, and at least one theme vector of each medical article; wherein, each medical article has a recommended weight value, and the at least one theme vector of each medical article includes the dominant theme vector and the subordinate theme vector of each medical article;

[0238] Calculate the matching degree between each theme vector of each medical article and the information vector respectively, and update the recommended weight value of each medical article according to the calculated matching degree;

[0239] Arrange the medical articles in descending order according to the updated recommended weight value of each medical article; select the medical article ranked first as the medical article to be recommended for output.

[0240] In another implementation, when updating the recommended weight value of each medical article according to the calculated matching degree, the one or more instructions can be loaded and specifically executed by the processor 701:

[0241] For any medical article, determine the topic vector with the largest matching degree from at least one topic vector of the any medical article according to the matching degree between each topic vector of the any medical article and the information vector;

[0242] If the topic vector with the largest matching degree is the dominant topic vector of the any medical article, increase the recommended weight value of the any medical article;

[0243] If the topic vector with the largest matching degree is the subordinate topic vector of the any medical article, decrease the recommended weight value of the any medical article.

[0244] In another implementation, when constructing the medical knowledge graph of the target medical article by using the multiple medical words, the one or more instructions can be loaded and specifically executed by the processor 701:

[0245] Construct an initial knowledge graph of the target medical article by using the multiple medical words. The initial knowledge graph includes multiple nodes, and each node records a medical word;

[0246] Select at least one pair of co-occurring word pairs from the multiple medical words. The co-occurring word pair refers to a word pair composed of two medical words having a co-occurrence relationship in the target medical article;

[0247] According to the at least one pair of co-occurring word pairs, determine at least one node group from the initial knowledge graph. Any node group includes: two nodes respectively recording the two medical words in a pair of co-occurring word pairs;

[0248] Connect the two nodes in each node group in the initial knowledge graph to obtain the medical knowledge graph of the target medical article.

[0249] In another implementation, when selecting at least one pair of co-occurring word pairs from the multiple medical words, the one or more instructions can be loaded and specifically executed by the processor 701:

[0250] Determine the first distribution position of the first medical word in the target medical article. The first medical word is any medical word in the multiple medical words;

[0251] Obtain a second medical term from the multiple medical terms according to the first distribution position of the first medical term, where the distance between the second distribution position of the second medical term in the target medical article and the first distribution position is less than a position distance threshold;

[0252] Calculate a semantic distance value between the first medical term and the second medical term, where the semantic distance value is used to indicate the semantic similarity between the first medical term and the second medical term;

[0253] If the semantic distance value between the first medical term and the second medical term is greater than a semantic threshold, determine that the first medical term and the second medical term have the co-occurrence relationship in the target medical article, and construct a pair of co-occurrence word pairs using the first medical term and the second medical term.

[0254] In another implementation manner, when calculating the importance of the medical terms recorded by each node based on the connection relationship between the nodes in the medical knowledge graph, the one or more instructions can be loaded and specifically executed by the processor 701:

[0255] For the medical term recorded by any node, based on the connection relationship between the nodes in the medical knowledge graph, determine at least one associated node connected to the any node;

[0256] Calculate the semantic distance values between the medical term recorded by the any node and the medical terms recorded by each associated node;

[0257] According to the calculated semantic distance values and the importance of the medical terms recorded by each associated node, calculate the importance of the medical term recorded by the any node.

[0258] In another implementation manner, when extracting a medical term set from a target medical article to be recognized, the one or more instructions can be loaded and specifically executed by the processor 701:

[0259] Perform word segmentation on the target medical article to be recognized to obtain an initial word set, where the initial word set includes multiple initial words;

[0260] Screen out multiple intermediate words from the initial word set according to at least one medical dictionary, where the intermediate words refer to the initial words existing in the at least one medical dictionary;

[0261] Construct the medical term set of the target medical article using the multiple intermediate words.

[0262] In another implementation, each initial word in the initial word set has a part of speech; correspondingly, when screening out multiple intermediate words from the initial word set according to at least one medical dictionary, the one or more instructions can be loaded and specifically executed by the processor 701:

[0263] Screen out the initial words of the target part of speech from the initial word set, where the target part of speech includes at least one of the following: noun, verb, and adjective;

[0264] Screen out multiple intermediate words from the initial words of the target part of speech according to at least one medical dictionary.

[0265] In another implementation, when constructing the medical word set of the target medical article using the multiple intermediate words, the one or more instructions can be loaded and specifically executed by the processor 701:

[0266] If there are intermediate words that meet the bonding condition among the multiple intermediate words, perform bonding processing on the intermediate words that meet the bonding condition; add the phrases after the bonding processing and the intermediate words that have not been bonded as medical words to the medical word set of the target medical article;

[0267] If there are no intermediate words that meet the bonding condition among the multiple intermediate words, add each intermediate word as a medical word to the medical word set of the target medical article;

[0268] Wherein, the bonding condition includes: adjacent distribution positions in the target medical article and existing in the same medical dictionary

[0269] In the embodiment of the present invention, for the target medical article to be recognized, multiple medical words can be first extracted from the target medical article, and a medical knowledge graph of the target medical article can be constructed using the multiple medical words. Since the medical words recorded by any two connected nodes in the medical knowledge graph have a co-occurrence relationship in the target medical article, and the medical words with more co-occurrence relationships are usually more important; therefore, based on the connection relationships between the nodes in the medical knowledge graph, the importance of the medical words recorded by each node can be calculated more accurately. Then, the key disease words of the target medical article can be selected from the medical word set according to the importance of each medical word, and the key topic vector indicating the key disease topic of the target medical article can be constructed using the word vectors of the key disease words. It can be seen that the embodiment of the present invention can effectively improve the accuracy of the key disease words by improving the accuracy of the importance of each medical word, thereby improving the accuracy of the key disease topic; moreover, the entire topic recognition process does not require manual participation by the user, and the key disease topic of the target medical article can be automatically recognized, effectively saving labor costs.

[0270] It should be noted that, according to one aspect of the present application, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative ways in the article recognition method embodiments shown above. Figure 2 or Figure 4 The methods provided in the various alternative ways in the article recognition method embodiments shown above.

[0271] Moreover, it should be understood that the above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A method for article recognition, characterized in that, Including: Extract a medical word set from a target medical article to be recognized. The medical word set includes multiple medical words, and at least includes disease words and non-disease words among the multiple medical words; Construct a medical knowledge graph of the target medical article by using the multiple medical words. The medical knowledge graph includes multiple nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article; The co-occurrence relationship means that during the process of sliding a sliding window in the target medical article, two medical words appear in the sliding window at the same time and the semantic distance value between the two medical words is greater than the semantic threshold; Based on the connection relationship between the nodes in the medical knowledge graph, calculate the importance of the medical words recorded by the nodes; Select the key disease words of the target medical article from the medical word set according to the importance of each medical word, and obtain the relevant non-disease words corresponding to the key disease words from the medical word set. The relevant non-disease words and the key disease words have the co-occurrence relationship in the target medical article; Obtain the word vector of the key disease word and the word vector of the relevant non-disease word; fuse the word vector of the key disease word and the word vector of the relevant non-disease word to obtain the key topic vector of the target medical article, and the key topic vector is used to indicate the key disease topic of the target medical article.

2. The method according to claim 1, wherein The selecting the key disease words of the target medical article from the medical word set according to the importance of each medical word includes: Select multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy; at least one candidate disease word is included in the multiple candidate keywords; Select the candidate disease word with the greatest importance from the at least one candidate disease word as the key disease word of the target medical article.

3. The method according to claim 2, wherein The method further includes: Obtain the word vectors of each candidate keyword, and calculate the vector similarity between the word vectors of each candidate keyword and the key topic vector; Select the article keywords of the target medical article from the multiple candidate keywords according to the vector similarity between the word vectors of each candidate keyword and the key topic vector; wherein, the vector similarity between the word vector of the article keyword and the key topic vector is greater than the similarity threshold; Associate and store the target medical article and the article keywords so that the target medical article can be processed according to the article keywords.

4. The method according to claim 2, characterized in that The selecting multiple candidate keywords of the target medical article from the medical word set according to the importance of each medical word according to the keyword selection strategy includes: Select a preset number of medical words from the medical word set as the multiple candidate keywords of the target medical article in the order of importance from large to small; or, Select the medical words with importance greater than the importance threshold from the medical word set as the multiple candidate keywords of the target medical article.

5. The method according to claim 2, wherein The key topic vector is the dominant topic vector of the target medical article, and the key disease topic is the main disease topic of the target medical article; the method further includes: Selecting a reference disease word of the target medical article from the at least one candidate disease word, where the importance of the reference disease word is less than the importance of the key disease word; Constructing a subordinate topic vector of the target medical article by using the word vector of the reference disease word, where the subordinate topic vector of the target medical article is used to indicate the secondary disease topic of the target medical article; Associatively storing the target medical article, the key topic vector, and the subordinate topic vector in a storage space, so that when there is an article search request, article search processing is performed according to the key topic vector and the subordinate topic vector.

6. The method according to claim 5, wherein The storage space further includes at least one other medical article, and each other medical article has a corresponding dominant topic vector and a corresponding subordinate topic vector; the method further includes: When there is an article search request, obtaining an information vector of the article search information carried in the article search request; Obtaining each medical article in the storage space, and at least one topic vector of each medical article; wherein each medical article has a recommendation weight value, and the at least one topic vector of each medical article includes the dominant topic vector and the subordinate topic vector of each medical article; Calculating the matching degree between each topic vector of each medical article and the information vector respectively, and updating the recommendation weight value of each medical article according to the calculated matching degree; Sorting the medical articles in descending order according to the updated recommendation weight value of each medical article; selecting the medical article ranked first as the medical article to be recommended for output.

7. The method according to claim 6, wherein The updating the recommendation weight value of each medical article according to the calculated matching degree includes: For any medical article, determining the topic vector with the largest matching degree from at least one topic vector of the any medical article according to the matching degree between each topic vector of the any medical article and the information vector; If the topic vector with the largest matching degree is the dominant topic vector of the any medical article, increasing the recommendation weight value of the any medical article; If the topic vector with the largest matching degree is the subordinate topic vector of the any medical article, decreasing the recommendation weight value of the any medical article.

8. The method according to claim 1, characterized in that, The constructing the medical knowledge graph of the target medical article by using the multiple medical words includes: Constructing an initial knowledge graph of the target medical article by using the multiple medical words, where the initial knowledge graph includes multiple nodes, and each node records a medical word; Selecting at least one pair of co-occurring word pairs from the multiple medical words, where the co-occurring word pair refers to a word pair composed of two medical words having a co-occurrence relationship in the target medical article; Determining at least one node group from the initial knowledge graph according to the at least one pair of co-occurring word pairs, and any node group includes: two nodes respectively recording the two medical words in a pair of co-occurring word pairs; Connect two nodes in each node group in the initial knowledge graph respectively to obtain the medical knowledge graph of the target medical article.

9. The method according to claim 8, wherein The selecting at least one pair of co-occurring word pairs from the multiple medical words includes: Determine the first distribution position of the first medical word in the target medical article, where the first medical word is any medical word among the multiple medical words; Obtain a second medical word from the multiple medical words according to the first distribution position of the first medical word, and the distance between the second distribution position of the second medical word and the first distribution position in the target medical article is less than the position distance threshold; Calculate the semantic distance value between the first medical word and the second medical word, and the semantic distance value is used to indicate the semantic similarity between the first medical word and the second medical word; If the semantic distance value between the first medical word and the second medical word is greater than the semantic threshold, determine that the first medical word and the second medical word have the co-occurrence relationship in the target medical article, and construct a pair of co-occurring word pairs with the first medical word and the second medical word.

10. The method according to claim 1, characterized in that, Based on the connection relationships between the nodes in the medical knowledge graph, calculate the importance of the medical words recorded by the nodes, including: For the medical word recorded by any node, based on the connection relationships between the nodes in the medical knowledge graph, determine at least one associated node connected to the any node; Calculate the semantic distance value between the medical word recorded by the any node and the medical words recorded by each associated node; Calculate the importance of the medical word recorded by the any node according to the calculated semantic distance value and the importance of the medical words recorded by each associated node.

11. The method according to claim 1, characterized in that, The extracting the medical word set from the target medical article to be recognized includes: Perform word segmentation on the target medical article to be recognized to obtain an initial word set, and the initial word set includes multiple initial words; Screen out multiple intermediate words from the initial word set according to at least one medical dictionary, where the intermediate words refer to the initial words existing in the at least one medical dictionary; Construct the medical word set of the target medical article with the multiple intermediate words.

12. The method according to claim 11, wherein Each initial word in the initial word set has a part of speech; the screening out multiple intermediate words from the initial word set according to at least one medical dictionary includes: Screen out the initial words with the target part of speech from the initial word set, and the target part of speech includes at least one of the following: noun, verb, and adjective; Screen out multiple intermediate words from the initial words with the target part of speech according to at least one medical dictionary.

13. The method according to claim 11, wherein The constructing the medical word set of the target medical article with the multiple intermediate words includes: If there are intermediate words that meet the bonding condition among the multiple intermediate words, perform bonding processing on the intermediate words that meet the bonding condition; add the phrases after the bonding processing and the intermediate words that have not been bonded as medical words to the medical word set of the target medical article; If there are no intermediate words that meet the bonding condition among the multiple intermediate words, add each intermediate word as a medical word to the medical word set of the target medical article; Among them, the bonding conditions include: being adjacent in the distribution positions in the target medical article and existing in the same medical dictionary.

14. An article recognition device, characterized in that, Including: An extraction unit, configured to extract a medical word set from a target medical article to be recognized, where the medical word set includes multiple medical words, and at least a disease word and a non-disease word are included in the multiple medical words; A construction unit, configured to construct a medical knowledge graph of the target medical article by using the multiple medical words, where the medical knowledge graph includes multiple nodes; one node records one medical word, and the medical words recorded by any two connected nodes have a co-occurrence relationship in the target medical article; the co-occurrence relationship means: during the process of sliding a sliding window in the target medical article, two medical words appear in the sliding window at the same time and the semantic distance value between the two medical words is greater than the semantic threshold; A processing unit, configured to calculate the importance of the medical words recorded by the respective nodes based on the connection relationships between the respective nodes in the medical knowledge graph; The processing unit is further configured to select key disease words of the target medical article from the medical word set according to the importance of each medical word, and obtain relevant non-disease words corresponding to the key disease words from the medical word set, where the relevant non-disease words and the key disease words have the co-occurrence relationship in the target medical article; obtain the word vector of the key disease word and the word vector of the relevant non-disease word; fuse the word vector of the key disease word and the word vector of the relevant non-disease word to obtain a key theme vector of the target medical article, and the key theme vector is used to indicate the key disease theme of the target medical article.

15. A computer device, characterized in that, The computer device includes: an input interface, an output interface, a processor, and a computer storage medium; among them, the computer storage medium stores one or more instructions, and the one or more instructions are loaded and executed by the processor to perform the article recognition method according to any one of claims 1-13.

16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded and executed by the processor to perform the article recognition method according to any one of claims 1-13.

17. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by the processor, the article recognition method according to any one of claims 1-13 is implemented.

Citation Information

Patent Citations

  • Text theme output method and device, storage medium and electronic device

    CN110162769A

  • Text keyword automatic extraction method and device based on co-occurrence language network

    CN111680509A