A medical information retrieval method based on an intelligent medical popular science platform

By dividing medical information retrieval into concept and descriptive entities, and utilizing medical knowledge graphs and semantic filtering to generate optimized evaluation values, this method solves the problems of inaccurate retrieval scope and insufficient result evaluation in existing methods, thus achieving efficient and accurate medical information retrieval.

CN120873030BActive Publication Date: 2025-12-09SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383424.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-09
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing medical information retrieval methods cannot accurately distinguish between medical concept entities and related descriptive entities in user search requests, resulting in search scopes that are too broad or too narrow. They fail to effectively filter information that matches user needs and lack scientific evaluation and ranking of search results, affecting search efficiency and user experience.

Method used

By dividing user search requests into medical concept entities and associated descriptive entities, core concept nodes and adjacent extended nodes are extracted using a pre-set medical knowledge graph. Combined with semantic filtering and multi-dimensional feature extraction, a medical information optimization evaluation value is generated, and the search results are output in priority order.

Benefits of technology

It has enabled the accurate capture of user search needs, improved search accuracy and efficiency, ensured that high-quality information is ranked at the top, reduced the possibility of users obtaining incorrect information, and enhanced the intelligence level of the medical science popularization platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873030B_ABST
    Figure CN120873030B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical information retrieval, and discloses a medical information retrieval method based on an intelligent medical popular science platform. The method collects a user retrieval request, divides the user retrieval request into a medical concept entity and an associated description entity, extracts core concept nodes and adjacent expansion nodes in a preset medical knowledge graph based on the medical concept entity, generates an initial retrieval range, performs semantic filtering on the initial retrieval range based on the associated description entity, screens out semantic matching nodes, constructs a target retrieval set, analyzes medical information in the target retrieval set, generates a medical information optimization evaluation value, prioritizes the target retrieval set according to the evaluation value, and outputs a sorted medical information retrieval result. The method improves the matching degree of the retrieval result and the user demand, facilitates the user to quickly obtain high-quality medical information, improves the retrieval experience, and is suitable for the information retrieval scene of the intelligent medical popular science platform.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical information retrieval, in particular to a medical information retrieval method based on an intelligent medical popular science platform. BACKGROUND

[0002] With the improvement of health awareness, more and more users obtain health-related information through medical popular science platforms, and efficient and accurate medical information retrieval has become the key to meet user needs. Currently, the information retrieval methods commonly used by medical popular science platforms rely on keyword matching or simple semantic analysis, and such methods have obvious limitations in processing user retrieval requests.

[0003] When users make retrieval requests, they often cannot accurately use professional medical terminology, and retrieval sentences often contain a mixture of daily expressions and medical concepts. Existing retrieval methods cannot effectively distinguish between medical concept entities and associated description entities in retrieval requests, resulting in retrieval ranges that are either too wide, including a large amount of information unrelated to core medical concepts, or too narrow, missing some relevant extension information. For example, when a user searches for "remedies for frequent headaches", existing methods may only extract the single medical concept "headache", ignoring the associated descriptions "frequently" and "remedies", and then returning a large amount of information about the causes and diagnosis of headaches, rather than the relief measures the user actually needs, increasing the user's burden of screening information.

[0004] Some platforms have introduced medical knowledge graphs to assist retrieval, but most of them are based only on single medical concept entities to extract nodes in the knowledge graph, without targeted filtering of the extracted nodes. Medical knowledge graphs contain a vast number of nodes, covering multiple dimensions such as diseases, symptoms, treatment plans, and drugs. If only the core concept nodes and adjacent nodes are simply extracted, the initial retrieval range will include a large amount of information that does not match the user's actual needs. For example, for the retrieval request "dietary recommendations for diabetic patients", existing methods may extract nodes based on "diabetes" and include nodes related to the pathogenesis and treatment of complications of diabetes, which have low relevance to "dietary recommendations", resulting in a low proportion of effective information in the subsequent retrieval results.

[0005] Existing retrieval methods lack scientific evaluation and sorting mechanisms for retrieval results. After obtaining retrieval results, they are mostly sorted by information publication time or click volume, without considering the professionalism, accuracy, and matching degree of information to user needs. Medical information has special characteristics, and the quality of information from different sources varies greatly, making it difficult for non-professional users to distinguish between true and false information. If only non-professional indicators are used for sorting, low-quality or even incorrect information may be ranked at the top, and users' reliance on such information may have adverse effects on their health. At the same time, when faced with a large amount of unordered retrieval results, users need to spend a lot of time browsing one by one, and cannot quickly obtain key information, seriously affecting retrieval efficiency and user experience. Summary of the Invention

[0006] The purpose of this invention is to provide a medical information retrieval method based on an intelligent medical science popularization platform, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, this invention provides a medical information retrieval method based on an intelligent medical science popularization platform, the method comprising:

[0008] Collect user search requests and divide the user search requests into medical concept entities and associated descriptive entities;

[0009] Based on the medical concept entities, core concept nodes and adjacent extended nodes are extracted from the preset medical knowledge graph to generate an initial search range.

[0010] Based on the associated description entities, semantic filtering is performed on the initial search range to select semantically matching nodes and construct the target search set;

[0011] The medical information within the target retrieval set is analyzed to generate optimized evaluation values ​​for the medical information.

[0012] The target retrieval set is prioritized and sorted according to the medical information optimization evaluation value, and the sorted medical information retrieval results are output.

[0013] Preferably, the step of analyzing the medical information within the target retrieval set and generating a medical information optimization evaluation value includes:

[0014] Multi-dimensional feature extraction is performed on the medical information within the target retrieval set to obtain the authoritative feature value, timeliness feature value, and audience suitability feature value of the medical information;

[0015] The authoritativeness, timeliness, and audience suitability of the medical information are weighted and calculated to generate a comprehensive evaluation value for the medical information.

[0016] The comprehensive evaluation value of the medical information is compared with a preset evaluation threshold range. If the comprehensive evaluation value of the medical information is lower than the lower limit of the evaluation threshold range, a content credibility warning signal is generated.

[0017] Based on the content credibility warning signal, the cross-modal association path density value and node information update frequency value of the semantic matching node are obtained;

[0018] The cross-modal associated path density value and the node information update frequency value are coupled and calculated to generate a dynamic adjustment factor;

[0019] Based on the dynamic adjustment factor, the medical information comprehensive evaluation value is corrected to obtain a medical information optimized evaluation value.

[0020] Preferably, the specific way of generating the initial search range is:

[0021] The main node position of the medical concept entity in the medical knowledge graph is identified, and the first-level associated nodes directly connected to the main node and their attribute labels are extracted;

[0022] The interaction path of the first-level associated nodes is traversed to obtain all second-level associated nodes and their relationship types;

[0023] The main node, first-level associated nodes and second-level associated nodes are merged to form the initial search range.

[0024] Preferably, the specific way of constructing the target search set is:

[0025] The syntax structure of the associated description entity is parsed to extract the keyword weight distribution and logical connection relationship;

[0026] The semantic similarity values of each node in the initial search range and the keyword weight distribution are calculated;

[0027] The nodes with a semantic similarity value greater than a preset similarity threshold value are screened and marked as semantic matching nodes;

[0028] All the semantic matching nodes and their associated edges are aggregated to form the target search set.

[0029] Preferably, the specific way of multi-dimensional feature extraction is:

[0030] The literature citation level and source agency level of the semantic matching nodes are extracted to generate the medical information authority feature value;

[0031] The latest revision timestamp and historical update cycle of the semantic matching nodes are extracted to generate the timeliness feature value;

[0032] The user historical feedback data and reading difficulty coefficient of the semantic matching nodes are extracted to generate the audience adaptation degree feature value.

[0033] Preferably, the specific way of generating the medical information comprehensive evaluation value is:

[0034] The first weight coefficient of the medical information authority feature value, the second weight coefficient of the timeliness feature value, and the third weight coefficient of the audience adaptation degree feature value are set;

[0035] The medical information comprehensive evaluation value is obtained by multiplying the medical information authority feature value by a first weight coefficient, multiplying the timeliness feature value by a second weight coefficient, multiplying the audience adaptation degree feature value by a third weight coefficient, and then superimposing and summing.

[0036] Preferably, the specific way of obtaining the cross-modal correlation path density value and the node information update frequency value is as follows:

[0037] The number of graph data paths, video data paths and clinical case data paths associated with the semantic matching node is counted to calculate the cross-modal correlation path density value.

[0038] The version change number of the semantic matching node in a preset time window is monitored to generate the node information update frequency value.

[0039] Preferably, the specific way of generating the dynamic adjustment factor is as follows:

[0040] The cross-modal correlation path density value is compared with a path density reference value to obtain a modal coverage coefficient.

[0041] The node information update frequency value is logarithmically normalized with an update frequency reference value to obtain a timeliness enhancement coefficient.

[0042] The modal coverage coefficient is multiplied by the timeliness enhancement coefficient to generate the dynamic adjustment factor.

[0043] Preferably, the specific way of correcting the medical information comprehensive evaluation value to obtain a medical information optimized evaluation value is as follows:

[0044] The medical information comprehensive evaluation value is added to the dynamic adjustment factor to obtain the medical information optimized evaluation value.

[0045] Preferably, the specific way of prioritizing the target retrieval set according to the medical information optimized evaluation value is as follows:

[0046] The nodes in the target retrieval set are sorted from high to low according to the medical information optimized evaluation value.

[0047] When the medical information optimized evaluation values are the same, the secondary sorting is performed according to the cross-modal correlation path density values.

[0048] Compared with the prior art, the present application has the following advantages:

[0049] In the retrieval range determination link, the method first divides the user retrieval request into medical concept entities and associated description entities. This division can accurately capture the core and supplementary information of the user's retrieval requirements, avoiding the retrieval deviation caused by the inability to distinguish between the two types of entities in existing methods. Based on the medical concept entity, the core concept node and adjacent expansion node are extracted in the preset medical knowledge graph, which can ensure that the retrieval range is expanded around the core medical concept, and can cover the potential useful information related to the core concept through the adjacent expansion node, avoiding the omission of key content due to too narrow retrieval range. For example, when the user retrieves "nursing measures for children with cold", the medical concept entity "children with cold" and the associated description entity "nursing measures" can be accurately extracted, and based on "children with cold", the core node and adjacent expansion nodes such as "cold symptom nursing" and "diet during cold" are extracted, laying a foundation for subsequent accurate retrieval.

[0050] In the semantic filtering link, the initial retrieval range is filtered based on the associated description entity, which can effectively eliminate nodes that do not match the user's actual needs, further reduce the retrieval range, and improve the retrieval accuracy. Compared with the existing method which only relies on keyword filtering or non-targeted screening, this semantic filtering link fully combines the associated description information in the user's retrieval request, making the target retrieval set after screening more in line with user needs. For example, when the user retrieves "suitable exercise for hypertension patients", the associated description entity "suitable exercise" can filter out irrelevant nodes such as "high blood pressure causes" and "high blood pressure drug treatment" in the initial retrieval range, and only keep relevant nodes such as "high blood pressure patient exercise type" and "high blood pressure patient exercise intensity", reducing the interference of useless information to the user.

[0051] In the information evaluation and sorting link, the medical information in the target retrieval set is analyzed and an optimized evaluation value is generated, which can comprehensively consider factors such as the professionalism, accuracy, completeness and matching degree of the information with the user's needs, rather than relying on a single indicator. Based on the evaluation value, the retrieval results are prioritized, which can make the information with higher quality and more in line with user needs rank in the front row, making it easier for users to quickly access key information and improving retrieval efficiency and user experience. For non-professional users, they do not need to distinguish and screen one by one in a large amount of information, but only need to focus on the information ranked at the top, which can obtain reliable medical knowledge and reduce the possibility of health risks caused by obtaining incorrect or low-quality information.

[0052] The method fully utilizes the structured advantages of the preset medical knowledge graph, combines with the semantic analysis technology, realizes the intelligent processing of the whole process from the retrieval request analysis to the result output, and improves the intelligent level of the information retrieval of the medical popular science platform. Compared with the traditional retrieval method, the method can better adapt to the diversified and personalized retrieval requirements of users. Whether the user uses daily expressions or professional terms to propose a retrieval request, high-quality retrieval results can be provided through accurate entity division, screening and evaluation, further promoting the practical value improvement of the medical popular science platform, and helping more users conveniently and accurately obtain the required medical information. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 A working principle diagram of the medical information retrieval method based on the intelligent medical popular science platform is described in the present application.

[0054] Figure 2 A method flow chart for generating a medical information optimization evaluation value is described.

[0055] Figure 3 A specific way flow chart for generating an initial retrieval range is described.

[0056] Figure 4 A specific way flow chart for multi-dimensional feature extraction is described. DETAILED DESCRIPTION

[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0058] Please refer to Figure 1 The present application provides a medical information retrieval method based on an intelligent medical popular science platform, which comprises:

[0059] The retrieval request submitted by the user through the input interface is collected, and the user retrieval request is parsed and divided into medical concept entities and associated description entities using natural language processing technology. The medical concept entity refers to the core medical terminology in the retrieval, such as the name of a disease or a drug; the associated description entity includes additional modifiers, symptom descriptions or context information. Based on the medical concept entity, node matching is performed in the preset medical knowledge graph, the core concept node corresponding to the medical concept entity is extracted, and the adjacent expansion node directly connected to the core concept node is further extracted to generate an initial retrieval range. The preset medical knowledge graph stores medical concept nodes and node relationships in a graph structure. Based on the associated description entity, a semantic similarity calculation model is used to perform semantic filtering on the nodes in the initial retrieval range, and nodes that are semantically matched with the associated description entity are selected to construct a target retrieval set. The medical information in the target retrieval set is analyzed, and a medical information optimization evaluation value is generated through multi-dimensional feature extraction and calculation. According to the medical information optimization evaluation value, the nodes in the target retrieval set are prioritized, and the sorted medical information retrieval result is finally returned through the output interface.

[0060] Embodiment 1: refer to Figure 2 , which shows the process of analyzing the medical information in the target retrieval set and generating a medical information optimization evaluation value. Through a specific user retrieval example, a clear display can be obtained. Assuming that the user input retrieval request is "the latest drug treatment plan for type 2 diabetes", the system first divides the request into medical concept entity "type 2 diabetes" and associated description entity "the latest drug treatment plan". After the previous steps, the system has extracted related nodes in the medical knowledge graph based on the core concept "type 2 diabetes", and has constructed a target retrieval set containing various drug information, clinical guideline nodes and research literature nodes according to the semantic filtering of "the latest drug treatment plan".

[0061] The analysis of the target search set begins with multi-dimensional feature extraction. The system processes the medical information authority feature value. For a node about "metformin" in the set, the system extracts its literature citation level. The node is cited by the American Diabetes Association clinical guidelines, and the citation level is high. At the same time, the source institution level is extracted. The information comes from a top medical research institution. According to the preset authority scoring rules, the system generates a higher medical information authority feature value for the node. For another node about a traditional Chinese medicine prescription, its literature is only cited by local journals, and the source institution is a general research institute, so the authority feature value generated is relatively low. The system extracts the timeliness feature value. For the "latest drug treatment plan" appeal, timeliness is crucial. The system reads the latest revision timestamp of each node. A node about GLP-1 receptor agonists has a timestamp showing that it was updated within the last three months, while another node about sulfonylurea drugs has a latest revision timestamp showing that it was updated two years ago. At the same time, the system analyzes the historical update cycle of these nodes. The GLP-1 receptor agonist node has frequent update records in the past year, while the sulfonylurea drug node has a longer update cycle. Based on this, the system generates a higher timeliness feature value for the former and a lower timeliness feature value for the latter. The system extracts the audience adaptation feature value. The platform calls the user's historical feedback data. For example, most users without a medical background give high scores and positive feedback to the "metformin" node. The reading difficulty coefficient is also at a medium level, suitable for the general public to read. On the contrary, a node containing a large amount of molecular formula and clinical trial raw data, although professional and accurate, has low click-through rate and poor reading completion rate in user historical feedback, and the reading difficulty coefficient is extremely high. Therefore, the system generates a higher audience adaptation feature value for the former and a lower audience adaptation feature value for the latter.

[0062] After feature extraction is complete, the system performs weighted calculation to generate a comprehensive evaluation value of medical information. In the system's preset weight configuration, considering the special nature of medical information, the authority feature value is given the highest first weight coefficient, the timeliness feature value is given the second weight coefficient, and the audience adaptation feature value is given the third weight coefficient. The three feature values of each node are multiplied by the corresponding weight coefficients, and then the three product results are added to obtain a preliminary comprehensive evaluation value of medical information. In this example, the GLP-1 receptor agonist node has a higher comprehensive evaluation value due to its high authority, high timeliness, and medium audience adaptation. The node that is highly professional but difficult to read, although its authority and timeliness may not be low, has a significantly reduced comprehensive evaluation value after weighted calculation due to its extremely low audience adaptation.

[0063] The system automatically compares the comprehensive evaluation value with the preset evaluation threshold interval. Assuming that a node about a rare herbal therapy has a low authoritative feature value due to unclear information sources, a low timeliness feature value due to an old publication time, and a comprehensive evaluation value of the medical information falling below the preset credibility threshold interval, the system generates a content credibility warning signal for the node.

[0064] Based on the warning signal, the system triggers a further data acquisition and analysis process. The system counts the number of cross-modal data paths associated with the semantic matching node, that is, checks whether the node is associated with various forms of content such as graphic interpretation, popular science videos, or clinical cases. If the node only has a simple text description, the cross-modal association path density value will be very low. At the same time, the system monitors the information update frequency of the node within the recent 24-month time window and finds that the version change frequency is zero, thereby generating a very low node information update frequency value.

[0065] The system performs coupling calculation to generate a dynamic adjustment factor. The system performs ratio operation on the calculated cross-modal association path density value and the average value of all node path densities in the system (path density reference value) to obtain a modality coverage coefficient, which is much less than 1 in this example. At the same time, the node information update frequency value is logarithmically normalized with the median of the update frequency of this type of node (update frequency reference value) to obtain a timeliness enhancement coefficient, which is also very low in this example. Multiplying the two coefficients generates a dynamic adjustment factor less than 1, which reflects the deficiencies of the node in information dimension and timeliness.

[0066] The system uses the dynamic adjustment factor to modify the original low medical information comprehensive evaluation value. The modification method is to add the comprehensive evaluation value and the dynamic adjustment factor. Since the dynamic adjustment factor is positive, this operation will slightly increase the evaluation value, but the increase is limited due to the small factor. The final medical information optimization evaluation value is still at a low level, which objectively reflects the current situation of the node with low information quality. This process ensures that even if some information is initially poorly evaluated, the system will fine-tune it by examining its information dimension and update potential, thereby outputting a more refined and accurate optimization evaluation value for subsequent retrieval result sorting.

[0067] Embodiment 2: see Figure 3, demonstrates the method of generating the initial search scope and constructing the target search set, which can be specifically illustrated through a real user query scenario. Assuming that a user inputs a search request: "inhalation treatment method for children with asthma". After receiving the request, the natural language processing module first parses it and divides it into medical concept entities and associated description entities. In this request, "asthma" is identified as the core medical concept entity, while "children" and "inhalation treatment method" are classified as associated description entities, which limit the core concept in terms of age and treatment method.

[0068] Based on the medical concept entity "asthma", the system starts searching in the preset medical knowledge graph to generate the initial search scope. The knowledge graph is composed of a large number of medical concept nodes and semantic relationship edges between nodes. The system first uses a graph query algorithm to accurately find the location of the main node that completely matches "asthma" in the graph. After locating the main node, the system extracts all first-level associated nodes directly connected to the main node and their attribute labels. These first-level associated nodes may include "bronchodilator", "glucocorticoid", "allergen", and other treatment drugs, pathological factors, and related symptom nodes. Each connection edge has an attribute label describing the relationship type, such as "treatment method", "triggering factor", or "associated symptom". Subsequently, the system starts the traversal process, taking these first-level associated nodes as the starting point to continue exploring along their interaction paths. Through a breadth-first search algorithm, the system obtains all second-level associated nodes directly connected to the first-level associated nodes and their relationship types. For example, from the first-level node "glucocorticoid", it can traverse to the second-level node "budesonide" (a specific drug), with the relationship type "is a"; from the first-level node "allergen", it can traverse to the second-level node "dust mites", with the relationship type "belongs to". This process collects all related nodes within two steps of relationship. Finally, the system combines the main node "asthma", all obtained first-level and second-level associated nodes, and automatically removes duplicate nodes to form an initial search scope that is rich in content and extensive in scope. This scope covers a large number of medical concepts directly and indirectly related to asthma.

[0069] The system then performs semantic filtering on this vast initial search scope based on the associated description entities "children" and "inhalation treatment method" to construct an accurate target search set. The system first analyzes the grammatical structure of these two associated description entities. "Children" is a noun modifier, and "inhalation treatment method" is a noun phrase, with "inhalation" as the manner adverb and "treatment method" as the core noun. Through dependency parsing technology, the system extracts the key words "children" and "inhalation" and determines their logical connection relationship as parallel limitation, which jointly modifies the core concept.

[0070] The system assigns a weight distribution to these keywords, and in the context of this query, "inhaled" as a specific treatment modality, can have a higher weight than the more general "pediatric" age restriction. The system then calculates a semantic similarity value for each node in the initial search scope with respect to this set of keyword weight distribution. This calculation is done through a word embedding model, which vectorizes the text content of each node (e.g. node name, description text, property labels) and the keywords, and measures their cosine similarity in the vector space.

[0071] A node about "adult asthma acute attack management" whose content vector is quite different from the direction of "pediatric" and "inhaled" vector, will have a lower calculated semantic similarity value. On the contrary, a node about "pediatric asthma guideline for inhaled device use" whose content vector is highly consistent with the direction of the keyword vector, will have a very high semantic similarity value. The system has a dynamically adjusted similarity threshold, which is optimized based on the feedback of historical search data. Among all nodes, only those with a semantic similarity value greater than the threshold will be screened out and marked as semantic matching nodes.

[0072] In this case, nodes like "budesonide pediatric dosage" and "pediatric asthma inhaled device selection" will be successfully screened, while irrelevant nodes such as "asthma surgical treatment" or "occupational asthma" will be filtered out. Finally, the system aggregates all marked semantic matching nodes and retains the associated edges between them in the knowledge graph, forming a refined, accurate, and tightly focused target search set around the user's query intent, laying the foundation for subsequent depth analysis and sorting output.

[0073] Embodiment 3: Refer to Figure 4 , which demonstrates the core link of multi-dimensional feature extraction and medical information comprehensive evaluation value generation. This process conducts in-depth analysis on medical information nodes that have passed semantic screening and are included in the target search set, aiming to quantitatively evaluate the information quality and applicability of each node. The following is illustrated through a specific example: Suppose the target search set contains multiple nodes related to "hypertension drug treatment", the system needs to extract features and evaluate these nodes.

[0074] Multi-dimensional feature extraction first targets the medical information authority feature value. The system parses the metadata and associated data sources of each node. For a node describing "angiotensin-converting enzyme inhibitors", the system extracts its literature citation level. By querying integrated external academic databases, it confirms that the node content is cited by multiple randomized controlled trials and European Society of Cardiology guidelines. According to the preset rules, such high-level citations will give the node a higher literature citation score. At the same time, the system extracts the source institution level of the node. The information comes from an internationally recognized cardiovascular disease research center, which is at the highest level in the internal authority ranking database. The literature citation level and source institution level are processed through a standardized mapping function, and finally fused to generate the medical information authority feature value of the node. This value is a normalized numerical value.

[0075] For the system to extract the timeliness feature value, the process relies on the version management metadata built into the node. The system reads the latest revision timestamp of the target node. For example, a node about "clinical application criteria for new antihypertensive drugs" has a timestamp showing the current year. At the same time, the system calculates the historical update cycle of the node, i.e. backtracking all historical versions to calculate the average time interval between adjacent versions. A node that is frequently updated and has a short historical update cycle usually means that its content keeps up with medical progress. The closeness of the latest revision timestamp to the current time and the length of the historical update cycle are combined through a pre-defined timeliness calculation model to generate a timeliness feature value reflecting the freshness of the information.

[0076] For the extraction of audience adaptation degree feature value, the system relies on the user behavior log and content analysis module inside the platform. The system extracts user historical feedback data of semantic matching nodes, including but not limited to the average length of time the node content is read by users, user ratings, sharing times, and subsequent consultation behaviors. A node about "daily life management of hypertension" has high user ratings and high average reading completion, indicating that its content is easy to understand and accept. At the same time, the system will call natural language processing tools to analyze the text content of the node and calculate its reading difficulty coefficient, which takes into account factors such as sentence length, vocabulary professionalism, and term density. User historical feedback data and reading difficulty coefficient are processed through an aggregation algorithm to generate an audience adaptation degree feature value.

[0077] After completing the extraction of three-dimensional feature values, the system enters the generation stage of medical information comprehensive evaluation value. The system assigns a weight coefficient to each feature value. These weight coefficients are not fixed but dynamically generated through an offline machine learning process on massive user interaction data and expert annotation data. Under the current system configuration, the weight coefficient of the medical information authority feature value is set to the highest, followed by the timeliness feature value, and then the audience adaptation degree feature value.

[0078] The medical information comprehensive evaluation value is calculated by a linear weighted summation model, which can be expressed by the following formula:

[0079]

[0080] In this formula, the character represents the final calculated medical information comprehensive evaluation value; the character represents the authority feature value of the medical information extracted and standardized in the previous step; the character represents the timeliness feature value extracted; the character represents the calculated audience adaptation degree feature value; the character represents the dynamic weight coefficient assigned to the authority feature value; the character represents the dynamic weight coefficient assigned to the timeliness feature value; and the character represents the dynamic weight coefficient assigned to the audience adaptation degree feature value.

[0081] In the calculation, the system substitutes the three feature values of each node into this formula. For example, a node with high authority, medium timeliness, and high audience adaptation degree has a higher comprehensive evaluation value, while a node with extremely high authority but extremely low timeliness and extremely high reading difficulty has a lower comprehensive evaluation value, even if the feature value A is multiplied by a high weight , the final comprehensive evaluation value will be significantly affected, thus objectively reflecting the comprehensive quality. The evaluation values of all nodes are finally normalized to a predetermined range to facilitate subsequent comparison and sorting.

[0082] Example 4: Describes an engineering initiated after the system generates a content credibility warning signal, which aims to produce dynamic adjustment factors by obtaining specific data and performing calculations to fine-tune the preliminary evaluation results. The following illustrates its implementation through a specific case.

[0083] ​​​Assume that the system is processing a user query about "Alzheimer's disease drug therapy", and one of the nodes, "clinical efficacy of drug A", has a relatively single information source and lacks recent updates, so its preliminary calculated medical information comprehensive evaluation value is lower than the system's preset threshold lower limit, triggering a content credibility warning signal. The system then starts the process of embodiment 4 for this specific node. The system first needs to obtain the cross-modal association path density value of the node, and the system will comprehensively scan all the data paths connected to the "clinical efficacy of drug A" node in the knowledge graph, and classify and count them according to their modal types. The system will accurately count the number of graphic-text data paths associated with the node, such as the number of linked drug chemical structure diagrams, mechanism of action diagrams, and detailed text explanation documents. At the same time, the system will count the number of video data paths associated with the node, such as the number of links to expert interpretation clips, patient medication guidance animations, or related academic conference videos. In addition, the system will also count the number of clinical case data paths associated with the node, that is, the number of links to case library entries recording the actual clinical application of the drug, patient feedback, and efficacy observation records. The sum of the number of the above three types of paths, after the standardization calculation process set in the system, finally generates the cross-modal association path density value of the node. This value reflects the diversity and richness of the presentation of the node information.

[0084] The system also monitors and obtains the node information update frequency value of the node. The system sets a retrospective time window, for example, the last twenty-four months, and within this time window, the system accesses the version control log of the node, which is a structured file recording all content change history. The system will accurately calculate the number of effective version changes recorded in the log. An effective version change may include updates to treatment data, additions to side effect information, revisions to medication guidelines, or additions to cited literature, etc. Simple text polishing or format adjustment is usually not counted. The total number of changes counted in this monitoring period is processed by the system and directly generated as the node information update frequency value of the node. This value directly reflects the maintenance activity and timeliness of the information content. In order to more clearly show the data acquisition situation of the system when processing multiple nodes, refer to Table 1, which records the example results of the system collecting data for a group of nodes that triggered the warning at a certain time.

[0085] Table 1: Node cross-modal data and update frequency monitoring table.

[0086]

[0087] After obtaining the above raw data, the system enters the coupling calculation stage to generate dynamic adjustment factors. The specific calculation starts from the processing of cross-modal association path density values. The system calls a preset path density benchmark value, which is not a fixed value but a representative value dynamically calculated based on historical data of similar nodes. The calculation method is: path density benchmark value = sum of cross-modal association path density values of similar nodes / number of similar nodes. Among them, "similar nodes" are divided according to the type of medical information (for example, if the current node is a "clinical efficacy of drug A" node, the similar nodes are all drug treatment class efficacy nodes); the cross-modal association path density value is the sum of the number of graphic and text paths, the number of video paths, and the number of clinical case paths associated with the node, that is: cross-modal association path density value = number of graphic and text paths + number of video paths + number of clinical case paths, and the system will recalculate and update the path density benchmark value every morning based on the full data of similar nodes in the past three months to ensure its timeliness and representativeness. The system generates a modality coverage coefficient through ratio calculation: modality coverage coefficient = cross-modal association path density value of current node / path density benchmark value; this coefficient reflects the degree of information diversity of the node relative to the average level of similar nodes. If the coefficient > 1, it means that the information diversity (cross-modal resource coverage) of the current node is better than the average level of similar nodes; if the coefficient = 1, it means that its information diversity is equal to the average level of similar nodes; if the coefficient < 1, it means that its information diversity is lower than the average level of similar nodes.

[0088] After completing the processing of cross-modal association path density values, the system simultaneously calculates the node information update frequency value to generate a timeliness enhancement coefficient. The system compares the update frequency value of the node with another preset update frequency benchmark value, that is: update frequency ratio = node information update frequency value of current node / update frequency benchmark value. The update frequency benchmark value is also dynamically generated by the system, for example, taking the median of all node update frequency values. The system performs logarithmic normalization processing on the ratio of the node information update frequency value to the update frequency benchmark value. This method standardizes the original update frequency value with high dispersion (such as 0-20 times / 24 months) to a preset reasonable interval (0.05-2) through the compression characteristics and linear interval mapping of the logarithmic function, which not only eliminates the excessive influence of extreme values on the evaluation results, but also preserves the relative differences in the update activity of different nodes, finally generating a timeliness enhancement coefficient. The specific calculation is divided into two steps:

[0089] Basic logarithmic conversion: eliminate zero value and extreme value interference based on smoothing coefficient, formula: wherein is the original update frequency, is the update frequency benchmark value, is the intermediate value after logarithmic conversion, is the smoothing coefficient (default 0.5);

[0090] Interval mapping: scaling the logarithmic conversion result to the target interval, the formula is: wherein , are the minimum and maximum values of the logarithmic conversion values of the same type of nodes, respectively, is the time-dependent strengthening coefficient, , are the lower limit 0.05 and the upper limit 2 of the time-dependent strengthening coefficient, respectively.

[0091] For example, when the original update frequency , the reference value , the calculated time-dependent strengthening coefficient is about 0.618; when , , the time-dependent strengthening coefficient is about 1.463, which realizes the smoothing and standardization of the original data. The coefficient reflects the strength of the node update activity relative to the reference level.

[0092] The system multiplies the calculated modal coverage coefficient and the time-dependent strengthening coefficient, and the product of the two coefficients is the dynamic adjustment factor generated for the specific node. This factor is a comprehensive adjustment parameter that considers both the richness of information dimension and the timeliness of information maintenance. If a node has a low initial score, but it is associated with multiple forms of learning materials and is updated frequently, then the dynamic adjustment factor calculated through this process will be a larger positive number, significantly improving its final evaluation value in subsequent correction. Conversely, a node with single and outdated content will have a small or even zero dynamic adjustment factor, and its evaluation value will not be effectively improved. This process introduces a more detailed and dynamic consideration dimension for the evaluation of medical information quality.

[0093] Example 5: The final output stage of the entire retrieval method is described, which is the core of the accurate revision and sorting of the target retrieval set after deep analysis. This process takes the calculation results and converts abstract evaluation values into an ordered retrieval result list that can directly face users. The following is a specific example to illustrate its operation mechanism: Assume that the system is processing a query about "coronary artery disease intervention". After processing all previous steps, the system has constructed a target retrieval set containing dozens of related nodes, and has calculated a medical information comprehensive evaluation value and a dynamic adjustment factor for each node in the set. At this time, the system finds that the node about "long-term efficacy of degradable stents" has a relatively low medical information comprehensive evaluation value due to the fact that the studies it cites are still in the follow-up stage, and the evidence level is relatively preliminary, triggering the early warning mechanism. However, subsequent analysis shows that the node is associated with rich imaging data, surgical simulation videos, and early clinical cases from multiple centers (high cross-modal association path density value), and has three data update records in the past half year (high node information update frequency value), resulting in a large dynamic adjustment factor.

[0094] The system performs revision operations to obtain the final medical information optimization evaluation value. The revision process is arithmetic, and the system directly adds the medical information comprehensive evaluation value of each node to the dynamically generated dynamic adjustment factor. For the "long-term efficacy of degradable stents" node, its originally low comprehensive evaluation value is significantly improved after adding a positive dynamic adjustment factor. The significance of this revision process is that it does not deny the cautious judgment of the information evidence level in the initial evaluation, but at the same time recognizes the value of the information in terms of presentation form diversity and update enthusiasm, thereby obtaining a more comprehensive and balanced final score. Conversely, a node about "old balloon dilation technology" may have authoritative content but is outdated, with few update records, so its dynamic adjustment factor will be very small, and its optimization evaluation value will basically remain the same, objectively reflecting its current limited reference value.

[0095] After completing the optimization evaluation value calculation for all nodes in the target retrieval set, the system enters the priority sorting stage. The sorting algorithm arranges all nodes in descending order according to the calculated medical information optimization evaluation value from high to low. The node with the highest optimization evaluation value represents that its comprehensive quality, timeliness, richness, and user adaptability have all been highly recognized by the system, so it is placed at the forefront of the result list.

[0096] The sorting algorithm also predefines rules for handling parallel cases. When the medical information optimization evaluation values of two or more nodes are exactly the same, the system will not perform random sorting, but will break the tie according to a secondary index: the cross-modal correlation path density value. Nodes with higher cross-modal correlation path density values will obtain a higher ranking. The design philosophy of this rule is that, in the case of equivalent overall information evaluation values, those information resources that can meet the different learning preferences and understanding abilities of users through various forms such as text, video, and cases are considered to have a slight priority for presentation.

[0097] The system outputs a strictly sorted medical information retrieval result list. Taking the "coronary artery disease intervention" query as an example, the node with the highest ranking may be the "latest version of clinical practice guidelines" node, which has a very high optimization evaluation value; the node that may follow it may be the "drug-eluting stent animation demonstration" node, which has a slightly lower evaluation value but a very high cross-modal density; and some nodes of popular science articles with single content and long-term non-updating will be at the end of the list. This ordered list is presented to the user through the user interface, effectively reducing the user's information screening cost and guiding the user to pay attention to and consume high-quality and highly relevant medical information content.

[0098] It should be noted that the relational terms herein, such as first and second, are used only to distinguish one entity or action from another entity or action, and do not necessarily require or imply that these entities or actions have any such actual relationship or order. In addition, the terms "comprise", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or device.

[0099] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.

Claims

1.A medical information retrieval method based on an intelligent medical popular science platform, characterized in that, The method comprises the following steps: collecting a user search request, dividing the user search request into a medical concept entity and an associated description entity; extracting a core concept node and an adjacent expansion node in a preset medical knowledge graph based on the medical concept entity to generate an initial search range; performing semantic filtering on the initial search range based on the associated description entity, screening out a semantic matching node, and constructing a target search set; analyzing the medical information in the target search set to generate a medical information optimization evaluation value; performing priority sorting on the target search set according to the medical information optimization evaluation value, and outputting a sorted medical information search result; the analysis of the medical information in the target search set to generate a medical information optimization evaluation value comprises: performing multi-dimensional feature extraction on the medical information in the target search set to obtain a medical information authority feature value, a timeliness feature value and an audience adaptation degree feature value; performing weighted calculation on the medical information authority feature value, the timeliness feature value and the audience adaptation degree feature value to generate a medical information comprehensive evaluation value; comparing the medical information comprehensive evaluation value with a preset evaluation threshold interval, if the medical information comprehensive evaluation value is lower than the lower limit of the evaluation threshold interval, generating a content credibility early warning signal; based on the content credibility early warning signal, obtaining the cross-modal association path density value and the node information update frequency value of the semantic matching node; coupling calculation of the cross-modal association path density value and the node information update frequency value to generate a dynamic adjustment factor; based on the dynamic adjustment factor, correcting the medical information comprehensive evaluation value to obtain a medical information optimization evaluation value; the specific way of multi-dimensional feature extraction is: extracting the literature citation level and the source agency level of the semantic matching node to generate the medical information authority feature value; extracting the latest revision timestamp and the historical update period of the semantic matching node to generate the timeliness feature value; extracting the user historical feedback data and the reading difficulty coefficient of the semantic matching node to generate the audience adaptation degree feature value; the specific way of obtaining the cross-modal association path density value and the node information update frequency value is: counting the number of graphic data paths, video data paths and clinical case data paths associated with the semantic matching node to calculate the cross-modal association path density value; monitoring the version change frequency of the semantic matching node in a preset time window to generate the node information update frequency value; the specific way of generating a dynamic adjustment factor is: performing ratio calculation on the cross-modal association path density value and the path density reference value to obtain a modal coverage coefficient; performing logarithmic normalization on the node information update frequency value and the update frequency reference value to obtain a timeliness enhancement coefficient; multiplying the modal coverage coefficient and the timeliness enhancement coefficient to generate the dynamic adjustment factor. 2.The medical information retrieval method based on the intelligent medical popular science platform according to claim 1, characterized in that, The specific way of generating an initial search range is: identifying the main node position of the medical concept entity in the medical knowledge graph, extracting the first-level associated nodes directly connected with the main node and their attribute labels; Traverse the interaction path of the primary association node to obtain all secondary association nodes and their relationship types; Merge the main node, primary association node and secondary association node to form the initial search range. 3.The medical information retrieval method based on the intelligent medical popular science platform according to claim 1, characterized in that, The specific way of constructing the target search set is: Parse the syntax structure of the association description entity, extract the keyword weight distribution and logical connection relationship; Calculate the semantic similarity value of each node in the initial search range and the keyword weight distribution; Screen the nodes with a semantic similarity value greater than a preset similarity threshold value and mark them as semantic matching nodes; Aggregate all the semantic matching nodes and their associated edges to form the target search set. 4.The medical information retrieval method based on the intelligent medical popular science platform according to claim 1, characterized in that, The specific way of generating a medical information comprehensive evaluation value is: Set the first weight coefficient of the medical information authority feature value, the second weight coefficient of the timeliness feature value and the third weight coefficient of the audience adaptation degree feature value; Multiply the medical information authority feature value by the first weight coefficient, the timeliness feature value by the second weight coefficient, and the audience adaptation degree feature value by the third weight coefficient, and then add and sum to obtain the medical information comprehensive evaluation value. 5.The medical information retrieval method based on the intelligent medical popular science platform according to claim 1, characterized in that, The specific way of correcting the medical information comprehensive evaluation value to obtain a medical information optimization evaluation value is: Add the medical information comprehensive evaluation value and the dynamic adjustment factor to obtain the medical information optimization evaluation value. 6.The medical information retrieval method based on the intelligent medical popular science platform according to claim 5, characterized in that, The specific way of prioritizing the target search set according to the medical information optimization evaluation value is: Sort the nodes in the target search set from high to low according to the medical information optimization evaluation value; When the medical information optimization evaluation value is the same, perform secondary sorting according to the cross-modal association path density value.

Citation Information

Patent Citations

  • Disease science popularization error correction method and system based on artificial intelligence

    CN120108774A

  • AI medical information accurate retrieval method based on knowledge graph

    CN120561370A