Method and apparatus for searching epidemic information

CN117573702BActive Publication Date: 2026-09-25MAGIC CUBE MEDICAL TECH (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311609700.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2026-09-25
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

[0007]然而,现有的流病分析多集中在大型流行病项目中,此类数据源虽提供了重要的信息,但分析对象单一片面,且分析粒度不够精细,难免存在数据误差和混杂,故无法满足用户对流病信息的高精度检索需求

Benefits of technology

[0023]上述流病信息检索方法和装置,服务器按照预设的疾病体系清洗流行病学数据,以此构建得到流病信息数据库之后,即可通过接收并响应携带有流病信息检索词的流病信息检索请求,在预先构建的流病信息数据库中筛选出与流病信息检索词相匹配的目标流病信息发送至客户端,为客户端用户提供全面的流病信息检索平台及精细的流病信息检索结果,满足用户对流病信息检索需求的同时,还可提高流病信息检索效率,节省信息调研人力物力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117573702B_ABST
    Figure CN117573702B_ABST
Patent Text Reader

Abstract

The application provides an epidemic information retrieval method and device. The method comprises the following steps: firstly, cleaning epidemiological data according to a preset disease system, and constructing an epidemic information database; then, receiving and responding to an epidemic information retrieval request carrying an epidemic information retrieval word; and finally, screening target epidemic information matched with the epidemic information retrieval word from the pre-constructed epidemic information database, and sending the target epidemic information to a client, so as to provide a comprehensive epidemic information retrieval platform and fine epidemic information retrieval result for the client user, meet the user's epidemic information retrieval demand, improve the epidemic information retrieval efficiency, and save manpower and material resources for information research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and apparatus for retrieving epidemic information. Background Technology

[0002] With the rapid development of science and society and people's increasing emphasis on health, health risk assessment has become an indispensable part of society, especially in predicting disease risk factors.

[0003] Epidemiological data analysis plays a crucial role in disease prevention and control. By collecting, integrating, and analyzing large amounts of health data and related information, it reveals the patterns and risk factors of diseases, thereby providing people with better means of preventing and controlling diseases.

[0004] First, epidemiological data analysis can help identify the transmission routes and patterns of diseases. By analyzing the spread of viruses or bacteria in different populations, the transmission paths of diseases can be inferred, allowing for targeted measures to reduce the risk of disease transmission. For example, in some past influenza outbreaks, epidemiological data analysis helped scientists determine the transmission routes of viruses, leading to the development of effective isolation and protective measures, which successfully controlled the spread of the epidemic.

[0005] Secondly, epidemiological data analysis can also reveal disease risk factors. By analyzing the health data and behavioral habits of large populations, it is possible to identify specific factors closely related to the risk of disease. For example, by analyzing the health data of smokers and non-smokers, we can find that smoking is one of the important factors leading to many health problems, thereby guiding the government to take corresponding measures to reduce the smoking population and thus reduce the incidence of related diseases.

[0006] Furthermore, epidemiological data analysis is of great significance to the medical field. By analyzing patient case data, it can provide data support for the development of innovative drugs for related diseases. Researchers can use the results of epidemiological data analysis to identify specific targets of diseases, design suitable drugs, and conduct clinical trials. Simultaneously, epidemiological data analysis can also help assess the economic benefits of drugs, provide a basis for decision-making on drug pricing and usage policies, and provide data support for competitive landscape analysis, market potential prediction, and pharmacoeconomic evaluation of innovative drugs for corresponding diseases.

[0007] However, existing epidemiological analyses are mostly concentrated in large-scale epidemic projects. While such data sources provide important information, the analysis is one-dimensional and lacks fine-grained analysis, inevitably leading to data errors and confounding. Therefore, they cannot meet users' needs for high-precision retrieval of epidemiological information. Furthermore, data providers must spend a significant amount of time manually cleaning and structuring the data, resulting in often untimely data updates. Summary of the Invention

[0008] The purpose of this invention is to provide a method and apparatus for retrieving epidemiological information, which combines artificial intelligence algorithms to mine epidemiological data from massive amounts of information, thereby building an information retrieval platform to achieve structured management of epidemiological information, meet users' needs for retrieving epidemiological information, and improve the retrieval efficiency of epidemiological information.

[0009] In a first aspect, the present invention provides a method for retrieving epidemic information, comprising: Receive requests for information on infectious diseases that carry search terms related to infectious diseases; In response to requests for retrieval of epidemiological information, the system filters out epidemiological information that matches the search terms from a pre-built epidemiological information database and uses it as the target epidemiological information. The epidemiological information database is constructed by cleaning epidemiological data step by step according to a pre-defined disease system. Send the target flow defect information to the client for display.

[0010] In some embodiments of the present invention, in response to an epidemiological information retrieval request, epidemiological information matching the epidemiological information retrieval terms is selected from a pre-constructed epidemiological information database as target epidemiological information. This includes: in response to an epidemiological information retrieval request, determining the retrieval term type of the epidemiological information retrieval terms; and, based on the retrieval term type, selecting epidemiological information of the target population matching the epidemiological information retrieval terms from the pre-constructed epidemiological information database as target epidemiological information. The retrieval term type includes at least one of the following: title retrieval terms, disease retrieval terms, clinical staging retrieval terms, biomarker retrieval terms, baseline feature retrieval terms, pathology retrieval terms, target retrieval terms, region retrieval terms, and dimension retrieval terms. The target population includes the basic population and / or subtype population; the epidemiological information includes historical epidemiological information and / or predicted epidemiological information.

[0011] In some embodiments of the present invention, the number of search terms for epidemic information retrieval is more than one. Based on the search term type, epidemic information of the target population that matches the epidemic information retrieval term is selected from a pre-constructed epidemic information database as target epidemic information. This includes: if the search term types of each epidemic information retrieval term are consistent, then the epidemic information of the target population that matches each epidemic information retrieval term is selected from the pre-constructed epidemic information database, and the union of each epidemic information is taken to obtain the target epidemic information; if the search term types of each epidemic information retrieval term are inconsistent, then the epidemic information of the target population that matches each epidemic information retrieval term is selected from the pre-constructed epidemic information database, and the intersection of each epidemic information is taken to obtain the target epidemic information.

[0012] In some embodiments of the present invention, historical epidemic information includes at least one of new cases, deaths, and total number of patients; predicted epidemic information includes at least one of first predicted epidemic information, second predicted epidemic information, and third predicted epidemic information; the first predicted epidemic information includes at least one of new cases, deaths, and total number of patients; the second predicted epidemic information includes population potential information; and the third predicted epidemic information includes drug potential information.

[0013] In some embodiments of the present invention, before responding to an epidemiological information retrieval request and filtering out epidemiological information matching the epidemiological information retrieval terms from a pre-constructed epidemiological information database as target epidemiological information, the method further includes: acquiring biomedical data; analyzing the biomedical data according to its data type to extract epidemiological data; and analyzing the epidemiological data based on a preset disease system to construct an epidemiological information database.

[0014] In some embodiments of the present invention, the analysis of biomedical data to extract epidemiological data based on the data type of the biomedical data includes: if the data type of the biomedical data is text, then the biomedical data is analyzed using a trained first model to extract epidemiological data; wherein the trained first model is obtained by fine-tuning a large-scale pre-trained language model; if the data type of the biomedical data is chart, and the chart is standard, then the biomedical data is analyzed using a trained second model to extract epidemiological data; wherein the trained second model is obtained by pre-training an initial language model based on entity alignment, a preset entity filling task, and a cloze test task; if the data type of the biomedical data is chart, and the chart is non-standard, then the biomedical data is analyzed using a trained third model to extract epidemiological data; wherein the trained third model is obtained by training a sample document set with pre-annotated chart areas covering title areas.

[0015] In some embodiments of the present invention, epidemiological data is analyzed based on a preset disease system to construct an epidemiological information database, including: performing entity identification on the epidemiological data to obtain a first entity and a first entity category, wherein the first entity category includes at least one of disease, stage, pathology, biomarker, and baseline characteristics; combining each first entity according to the first entity category and the preset disease system to determine the base population and / or subtype population in the epidemiological data; performing data cleaning on the epidemiological data for the base population and / or subtype population to obtain at least one of new cases, deaths, and total number of patients as historical epidemiological information; determining at least one of a first predicted epidemiological information, a second predicted epidemiological information, and a third predicted epidemiological information based on the historical epidemiological information as predicted epidemiological information; and summarizing the predicted epidemiological information and / or historical epidemiological information to obtain epidemiological information to construct an epidemiological information database.

[0016] In some embodiments of the present invention, for the basic population and / or subtype population, epidemiological data is cleaned to obtain at least one of new cases, deaths, and total number of patients as historical epidemiological information. This includes: for the basic population and / or the subtype population, entity recognition is performed on the epidemiological data to obtain a second entity; wherein the second entity is a region-related entity; based on the second entity, at least one of new cases, deaths, and total number of patients in the epidemiological data is cleaned to obtain historical epidemiological information.

[0017] In some embodiments of the present invention, the epidemiological information retrieval method further includes: if the epidemiological data includes the proportion of new cases in a subtype population but does not include new cases in a subtype population, then based on the proportion of new cases and the corresponding new cases in the basic population, the method obtains the new cases in the subtype population as historical epidemiological information; and / or if the epidemiological data includes the proportion of deaths in a subtype population but does not include deaths in a subtype population, then based on the proportion of deaths in a subtype population and the corresponding deaths in the basic population, the method obtains the deaths in a subtype population as historical epidemiological information; and / or if the epidemiological data includes the proportion of the total number of cases in a subtype population but does not include the total number of cases in a subtype population, then based on the proportion of the total number of cases and the corresponding total number of cases in the basic population, the method obtains the total number of cases in a subtype population as historical epidemiological information.

[0018] In some embodiments of the present invention, determining at least one of a first predicted epidemic information, a second predicted epidemic information, and a third predicted epidemic information as predicted epidemic information based on historical epidemic information includes: calculating at least one of new cases, deaths, and total number of patients based on the numerical change rate of historical epidemic information as the first predicted epidemic information; and / or obtaining population potential information based on the first predicted epidemic information or the total number of new cases or patients as the second predicted epidemic information; and / or obtaining drug potential information based on the second predicted epidemic information as the third predicted epidemic information.

[0019] In a second aspect, the present invention provides an epidemiological information retrieval device, comprising: The request receiving module is used to receive requests for retrieving epidemic information carrying keywords for epidemic information retrieval; The information retrieval module is used to respond to requests for retrieval of epidemiological information. It filters out epidemiological information that matches the search terms in a pre-built epidemiological information database and uses it as the target epidemiological information. The epidemiological information database is constructed by cleaning epidemiological data step by step according to a preset disease system. The information feedback module is used to send target flow defect information to the client for display.

[0020] Thirdly, the present invention also provides a computer device, comprising: One or more processors; The memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the above-described method for retrieving disease information.

[0021] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the epidemiological information retrieval method.

[0022] Fifthly, embodiments of the present invention provide a computer program product or computer program, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the method provided in the first aspect described above.

[0023] The aforementioned method and apparatus for retrieving epidemiological information involves a server that cleans epidemiological data according to a pre-defined disease system to construct an epidemiological information database. After this database is built, the server can receive and respond to epidemiological information retrieval requests carrying epidemiological information retrieval terms. The server then filters out target epidemiological information matching the retrieval terms from the pre-built database and sends it to the client. This provides client users with a comprehensive epidemiological information retrieval platform and detailed retrieval results, satisfying users' needs for epidemiological information retrieval while improving retrieval efficiency and saving manpower and resources for information research. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of a scenario for the epidemiological information retrieval method in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the method for retrieving epidemiological information in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of the epidemic information retrieval device in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] It should be noted that in the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0028] Meanwhile, the epidemiological information retrieval method provided in this embodiment of the invention can be applied to, for example... Figure 1The illustrated epidemiological information retrieval system includes a client 102 and a server 104. The client 102 can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication over a bidirectional communication link. Such a device can include cellular or other communication devices with single-line displays, multi-line displays, or no multi-line displays. Specifically, the client 102 can be a desktop terminal or a mobile terminal; it can also be a mobile phone, tablet computer, or laptop computer. The server 104 can be a standalone server or a server network or server cluster, including but not limited to computers, network hosts, single network servers, multiple network server sets, or cloud servers composed of multiple servers. The cloud server consists of a large number of computers or network servers based on cloud computing. Furthermore, the client 102 and server 104 establish a communication connection through a network, which can be any of a wide area network (WAN), local area network (LAN), or metropolitan area network (MAN).

[0029] Furthermore, those skilled in the art will understand that Figure 1 The application environment shown is merely one applicable scenario for the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include more than one. Figure 1 The number of devices shown may be more or less. For example, Figure 1 Only one server is shown. It is understood that this epidemic information retrieval system may also include one or more other devices, which are not specifically limited here. Additionally, the epidemic information retrieval system may include a storage device for storing data, such as epidemic information.

[0030] certainly, Figure 1 The schematic diagram of the epidemiological information retrieval system shown is merely an example. The epidemiological information retrieval system and scenario described in the embodiments of the present invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. As those skilled in the art will know, with the evolution of epidemiological information retrieval systems and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0031] Epidemiology is the science that studies the distribution of diseases and health conditions in a population and their influencing factors, as well as strategies and measures for disease prevention and control and health promotion. Collecting and cleaning epidemiological data can clarify the severity of disease in a population (e.g., incidence, prevalence), identifying high-incidence populations and regions, the proportion of each disease stage, and the proportion of each molecular subtype, thus providing a basis for disease prevention and treatment or uncovering unmet clinical needs. Based on this, this invention proposes to identify high-risk populations (i.e., subdivided indication populations) by clarifying the three-dimensional (human, temporal, spatial) distribution of diseases. Then, it proposes to associate epidemic information with different subdivided indication populations, building an information retrieval platform that allows users to query corresponding epidemic information within subdivided indication populations. This facilitates later assessment of market size and the potential of subdivided indications. Detailed platform application and construction steps are as follows.

[0032] See Figure 2 This is a flowchart illustrating a method for retrieving epidemiological information according to an embodiment of the present invention. This embodiment mainly applies this method to the above-mentioned... Figure 1 Taking server 104 as an example, the method includes steps S201 to S203, as follows:

[0033] S201, Receive an epidemiological information retrieval request carrying epidemiological information retrieval terms.

[0034] Since search terms are usually used to summarize relevant words for the content to be searched, search terms for epidemiological information can be used to summarize relevant words for the epidemiological information to be searched, and can be used on information retrieval platforms to allow users to search for and obtain specific epidemiological information.

[0035] For example, epidemiological information search terms include, but are not limited to: title search terms (various sub-indications, such as: liver cancer, liver cancer / hepatocellular carcinoma, liver cancer / hepatocellular carcinoma / stage III-IV, etc.), disease search terms (various disease names, such as: liver cancer, non-small cell lung cancer, triple-negative breast cancer, etc.), clinical staging search terms (such as: stage I, stage II, stage III, etc.), biomarker search terms (such as: 11q deletion, ABL1 mutation, etc.), baseline characteristic search terms (patient baseline characteristics, such as: recurrence, disseminated metastasis, male, female, etc.), pathology search terms (such as: carcinosarcoma, adenosarcoma, etc.), target search terms (such as: PDE-phosphodiesterase, PBP-penicillin-binding protein), regional search terms (such as: global, China, United States, etc.), and dimension search terms (such as: new cases, deaths, total number of patients).

[0036] In a specific implementation, a user can send a virus information retrieval request to the server 104 through the client 102, or through other devices. The "other devices" mentioned here can be devices without a communication connection to the client 102, or devices with a communication connection to the client 102; specific embodiments of the present invention are not limited to these.

[0037] For example, server 104 renders the search page of the epidemic information retrieval platform to client 102 for display. The user can then determine at least one epidemic information search term through the epidemic information retrieval page displayed on client 102. Client 102 then sends an epidemic information retrieval request carrying one or more epidemic information search terms to server 104, instructing server 104 to respond to the request and obtain the target epidemic information required by the user. The methods for determining the epidemic information search terms mentioned in this embodiment include, but are not limited to: clicking, double-clicking, or long-pressing preset candidate search terms, or entering non-preset search terms.

[0038] S202, in response to the request for retrieval of epidemiological information, select epidemiological information that matches the search terms in the pre-built epidemiological information database as the target epidemiological information; wherein, the epidemiological information database is constructed by cleaning epidemiological data step by step according to the preset disease system.

[0039] In this embodiment of the invention, the disease system can be obtained by combining information from at least one dimension, including disease, stage, pathology, biomarkers, and baseline characteristics. For example, the disease system can be composed of: a first layer of disease, a second layer of disease + stage, a third layer of disease + stage + pathology, a fourth layer of disease + stage + pathology + biomarkers, and a fifth layer of disease + stage + pathology + biomarkers + baseline characteristics.

[0040] However, it should be noted that the above system represents a special case where disease, stage, pathology, biomarkers, and baseline features are all present simultaneously. If one or more fields are absent, the hierarchical order will change. For example, if the "stage" field is not currently detected, the first layer is disease, the second layer is disease + pathology, the third layer is disease + pathology + biomarkers, and the fourth layer is disease + pathology + biomarkers + baseline features. Similarly, if the "stage" and "biomarkers" fields are not currently detected, the first layer is disease, the second layer is disease + pathology, and the third layer is disease + pathology + baseline features.

[0041] The epidemiological information can include historical epidemiological information and / or predicted epidemiological information. Historical epidemiological information refers to basic epidemiological information for a specific historical period (e.g., 2018-2023), while predicted epidemiological information refers to at least one of the following: basic epidemiological information, population potential information, and drug potential information for a specific future period (e.g., 2024-2028). Basic epidemiological information (unit: number of people) includes at least one of new cases, deaths, and total number of patients. New cases represent the number of people experiencing new events within a specific period, deaths represent the number of people experiencing deaths within a specific period, and total number of patients represents the total number of people infected within a specific period. Population potential information (unit: amount) is used to present the sales revenue of a specific indication population, and drug potential information (unit: amount) is used to present the market valuation of the corresponding drug.

[0042] In practice, after receiving an epidemiological information retrieval request, server 102 responds by filtering epidemiological information from the epidemiological information database that matches the search term, using this information as the target epidemiological information. However, it should be noted that the criteria for "matching" mentioned above include, but are not limited to: identical words, word inclusion, and semantic similarity. For example, if the epidemiological information retrieval term is "liver cancer," the matching value in the epidemiological information database could be either "liver cancer" (which matches the word) or "hepatocellular carcinoma" (which contains the word "liver cancer"). Similarly, if the epidemiological information retrieval term is "central nervous system infection," the matching value in the epidemiological information database could be "epidemic encephalitis B" (which has the same semantic meaning).

[0043] In one embodiment, step S202 includes: responding to an epidemiological information retrieval request and determining the search term type of the epidemiological information search terms; based on the search term type, filtering epidemiological information of the target population that matches the epidemiological information search terms from a pre-built epidemiological information database as target epidemiological information; wherein, the search term type includes at least one of the following: title search terms, disease search terms, clinical staging search terms, biomarker search terms, baseline feature search terms, pathology search terms, target search terms, region search terms, and dimension search terms; the target population includes the basic population and / or subtype population; and the epidemiological information includes historical epidemiological information and / or predicted epidemiological information.

[0044] The search terms can be categorized as follows: title terms can be summarized based on specific indications; disease terms can be summarized based on disease names; clinical stage terms can be summarized based on clinical trial progress; biomarker terms can be summarized based on substances or indicators produced in the body; baseline characteristic terms can be summarized based on patient characteristics; pathology terms can be summarized based on the morphology, structure, and function of human tissues or cells; target terms can be summarized based on drug targets; regional terms can be summarized based on regions; and dimensional terms can be summarized based on statistical dimensions. Examples of each search term have been shown above and will not be repeated here.

[0045] Determining whether a specific indication population belongs to the basic population or a subtype population depends not only on the population cleaning results of the epidemiological data to be cleaned, but also on whether the cleaned population has leaf nodes. That is, the basic population is the subtype indication population that serves as the root node in the population cleaning results, while the subtype population is the subtype indication population that serves as the leaf node in the population cleaning results. For example, if the population that can be extracted from the article includes: ① non-small cell lung cancer, ② non-squamous non-small cell lung cancer, and ③ squamous non-small cell lung cancer, then the baseline population is ① non-small cell lung cancer, and ② non-squamous non-small cell lung cancer and ③ squamous non-small cell lung cancer are both subtypes. As another example, if the population extracted from the article includes: ① non-squamous non-small cell lung cancer, ② non-small cell lung cancer / non-squamous / adenocarcinoma, and ③ non-small cell lung cancer / non-squamous / adenocarcinoma / BRAF mutation, then the baseline population is ① non-squamous non-small cell lung cancer, ② non-squamous non-small cell lung cancer / non-squamous / adenocarcinoma is a subtype of ① non-squamous non-small cell lung cancer, and ③ non-squamous non-small cell lung cancer / non-squamous / adenocarcinoma / BRAF mutation is also a subtype of ① non-squamous non-small cell lung cancer.

[0046] The historical and predicted epidemic information has been explained above and will not be repeated here.

[0047] In specific implementation, after receiving and responding to the epidemic information retrieval request, server 102 can detect the number of search terms for the epidemic information retrieval. If the number of search terms is one, it is only necessary to determine the search term type of the epidemic information retrieval term in order to determine the corresponding result matching rule and complete the matching determination of the target epidemic information. If the number of search terms is greater than one, it is necessary not only to determine the search term type of each epidemic information retrieval term, but also to further determine whether the multiple search term types are consistent, and then determine the matching rule based on the consistency judgment result. Here, this embodiment will describe in detail the target epidemic information determination steps when the number of search terms is one; the target epidemic information determination steps when the number of search terms is greater than one will be described in detail in the next embodiment.

[0048] Specifically, when server 102 responds to an epidemiological information retrieval request and receives the epidemiological information retrieval term "liver cancer," it determines that the number of retrieval terms is one. At this point, it only needs to determine the retrieval term type of "liver cancer" to lock in the matching results based on the retrieval term type. For example, if "liver cancer" is a title retrieval term, the target epidemiological information is the matching result that matches the word "liver cancer"; if "liver cancer" is a disease retrieval term, the target epidemiological information includes matching results that match the word "liver cancer," as well as matching results whose words contain "liver cancer." Of course, the above example is for illustrative purposes only. In actual application scenarios, different matching rules can be associated with different retrieval term types according to requirements. Specific embodiments of the present invention are not limited to this.

[0049] In one embodiment, the number of search terms for epidemiological information is more than one. Based on the search term type, epidemiological information matching the search terms is selected from a pre-built epidemiological information database as target epidemiological information. This includes: if the search term types of each epidemiological information search term are consistent, then epidemiological information matching each search term is selected from the pre-built epidemiological information database, and the union of the epidemiological information is taken to obtain the target epidemiological information; if the search term types of each epidemiological information search term are inconsistent, then epidemiological information matching each search term is selected from the pre-built epidemiological information database, and the intersection of the epidemiological information is taken to obtain the target epidemiological information.

[0050] In practice, if the epidemiological information search terms received by server 102 include "liver cancer" which belongs to the disease search terms and "China" which belongs to the regional search terms, it can be determined that the search term types of the two epidemiological information search terms are inconsistent. At this time, there are "19" search results that match "liver cancer". After taking the intersection with "China", only "6" search results remain.

[0051] Furthermore, if the disease information search terms "liver cancer" and "breast cancer" received by server 102 are both classified as disease search terms, it can be determined that the search term types of the two disease information search terms are consistent. At this time, there are "19" search results matching "liver cancer" and "36" search results matching "breast cancer", and the final search results are "55".

[0052] In one embodiment, before step S202, the method further includes: acquiring biomedical data; analyzing the biomedical data according to its data type to extract epidemiological data; and analyzing the epidemiological data based on a preset disease system to construct an epidemiological information database.

[0053] Biomedical data can be data in any format and language related to biomedicine. Examples include biomedical literature, publicly available information from biomedical companies, officially published annual reports (such as the China Cancer Registry Annual Report and the US Cancer Annual Report), and IPO (Initial Public Offering) prospectuses.

[0054] In specific implementation, server 102 can first obtain biomedical data in one of the following ways: 1. In a normal network structure, server 104 can receive biomedical data to be analyzed from client 102 or other cloud devices with network connections; 2. In a pre-built blockchain network, server 104 can synchronously obtain biomedical data to be analyzed from other terminal nodes or server nodes. The blockchain network can be a public chain, a private chain, etc.; 3. In a pre-built tree structure, server 104 can request biomedical data to be analyzed from an upper-level server or poll from a lower-level server.

[0055] Then, server 102 can extract epidemiological data from the biomedical data according to the data type, including text and chart types. The extraction methods for different types of data will be described in detail below. Finally, after analyzing the epidemiological data, server 102 can use a standard entity dictionary to standardize each entity in the epidemiological data. At the same time, it can divide each sub-indication according to the hierarchical relationship in the disease system, thereby connecting a relatively complete system of sub-indications and treating them as target populations. Corresponding values ​​can be extracted from the currently obtained data for correlation, or the epidemiological information of subtype indications can be calculated according to the publicly available proportions in the data, thereby constructing an epidemiological information database.

[0056] In one embodiment, analyzing biomedical data and extracting epidemiological data based on the data type of the biomedical data includes: if the data type of the biomedical data is text, then analyzing the biomedical data using a trained first model to extract epidemiological data; wherein the trained first model is obtained by fine-tuning a large-scale pre-trained language model; if the data type of the biomedical data is chart, and the chart is standard, then analyzing the biomedical data using a trained second model to extract epidemiological data; wherein the trained second model is obtained by pre-training an initial language model based on entity alignment, a preset entity filling task, and a cloze test task; if the data type of the biomedical data is chart, and the chart is non-standard, then analyzing the biomedical data using a trained third model to extract epidemiological data; wherein the trained third model is obtained by training a sample document set with pre-annotated chart areas covering title areas.

[0057] In practical implementation, if the biomedical data is plain text without charts or graphs, machine learning techniques can be used for text classification to identify and extract epidemiological data from the biomedical data. To improve the accuracy and efficiency of the text classification model, this solution proposes using a multilingual BERT large-scale pre-trained model combined with fine-tuning to train the first model, thus obtaining a trained first model for document classification. This model allows us to process text in multiple languages ​​simultaneously without needing to customize models for each language separately.

[0058] Specifically, the crawled biomedical data is first preprocessed, including removing stop words, punctuation marks, and other useless information, and then converted into token sequences for model processing. Then, pre-labeled keywords (e.g., onset, death, patients) from the data annotation stage are used to filter out documents related to epidemiology, thus identifying and extracting epidemiological data from the massive biomedical data.

[0059] Furthermore, if the biomedical data is in the form of charts and is a standard table, a trained second model can be used to extract information from the text content to identify epidemiological data, which includes, but is not limited to: disease name, target information, region, statistical dimensions (death cases, new cases, total number of patients), time, growth rate, etc.

[0060] Specifically, the second model effectively utilizes tables within the text to enhance the structure of information extraction. For text containing standard tables, the second model first extracts the first row and first column information from the table and determines whether it is a header row or header column. If so, it extracts entity pairs from the header content and table content to construct a cloze test template, which is then analyzed in conjunction with the article information to extract epidemiological data from biomedical data. The model structure of the second model is as follows:

[0061] 1) Input layer: Used to receive text sequences and cloze test templates (if there is table information), and encode the text sequences into semantic vector sequences.

[0062] 2) BiLSTM (Bidirectional Long Short-Term Memory) layer: used to take the encoded semantic vector sequence as input to BiLSTM to obtain the semantic information of each word in the context.

[0063] 3) CRF (Conditional Random Fields) layer: used to solve complex situations such as entity intersection or nesting through global optimization.

[0064] 4) Pointer Network layer: used to find the corresponding position in the original text for the missing entities in the cloze test template.

[0065] 5) Output and Correction Layer: This layer fuses the outputs of the CRF layer and the Pointer Network layer. If the Pointer Network layer can find the corresponding entity in the original text, but the CRF layer did not predict it, then that result will also be used as the final output entity.

[0066] Furthermore, if the biomedical data is in the form of charts, but not standard tables (standard definition: the first row and first column are both titles), the epidemiological data can be extracted as follows: First, a trained third-party model is used to extract table information from the biomedical data. This model learns from training data to predict the possible locations of charts in the article and accurately obtains the table titles. Once the location of the charts is known, the text content can be extracted. For plain text charts, the text information can be accurately obtained by directly parsing the content of the source file and combining it with chart location; for image-type charts, OCR (Optical Character Recognition) technology can be used to extract text information from the charts. Finally, after extracting the text information from the charts, the tables can be classified based on this text information to filter out images or tables containing epidemiological information. Here, because the third-party model is trained on a set of sample documents with pre-labeled chart areas covering the title areas, it can predict the possible locations of charts in the article and accurately obtain the table titles.

[0067] In one embodiment, epidemiological data is analyzed based on a preset disease system to construct an epidemiological information database, including: entity identification of the epidemiological data to obtain a first entity and a first entity category, wherein the first entity category includes at least one of disease, stage, pathology, biomarker, and baseline characteristics; combining the first entities according to the first entity category and the preset disease system to determine the base population and / or subtype population in the epidemiological data; performing data cleaning on the epidemiological data for the base population and / or subtype population to obtain at least one of new cases, deaths, and total number of patients as historical epidemiological information; determining at least one of a first predicted epidemiological information, a second predicted epidemiological information, and a third predicted epidemiological information based on the historical epidemiological information as predicted epidemiological information; and summarizing the predicted epidemiological information and / or historical epidemiological information to obtain epidemiological information to construct an epidemiological information database.

[0068] In specific implementation, after server 102 analyzes and extracts epidemiological data from biomedical data, it first performs disease entity identification on the epidemiological data to obtain first entities belonging to any one of the following types: disease, stage, pathology, biomarker, or baseline feature. Then, a standard entity dictionary is used to standardize each first entity to obtain uniquely named first entities. Finally, the entity type (i.e., first entity type) of each standard first entity is analyzed, and the first entities are combined according to a preset disease system to determine the base population and / or subtype population. In this embodiment, the standard entity dictionary is a dictionary composed of standardized entities.

[0069] Furthermore, if the currently identified sub-indication population includes a basic population and a subtype population, server 102 can first perform data cleaning on the basic population, that is, clean out the new cases, deaths, and / or total number of patients in the basic population, as historical epidemiological information for that basic population. Then, data cleaning is performed on the subtype population, that is, clean out the new cases, deaths, and / or total number of patients in the subtype population, as historical epidemiological information for that subtype population. However, if the historical epidemiological information of the subtype population cannot be obtained through data cleaning, server 104 can also obtain it through other methods, such as by proportional conversion. The specific conversion method will be described in detail in subsequent embodiments.

[0070] Furthermore, after obtaining historical epidemiological information for each subdivided indication population, the first predicted epidemiological information for the corresponding population can be calculated using the historical epidemiological information. Then, based on the first predicted epidemiological information, the second predicted epidemiological information for the corresponding population can be calculated, and based on the second predicted epidemiological information, the third predicted epidemiological information for the corresponding population can be calculated. Finally, all the cleaned or calculated epidemiological information is integrated and stored in the database as an epidemiological information database, which is available for the server 104 to call to meet the user's retrieval needs for epidemiological information.

[0071] In one embodiment, for the basic population and / or subtype population, epidemiological data is cleaned to obtain at least one of new cases, deaths, and total number of patients as historical epidemiological information. This includes: for the basic population and / or the said subtype population, performing entity identification on the epidemiological data to obtain a second entity; wherein the second entity is a region-related entity; and based on the second entity, cleaning at least one of new cases, deaths, and total number of patients from the epidemiological data as historical epidemiological information.

[0072] The second entity includes, but is not limited to, regional names such as global, China, the United States, the European Union, and Japan. The purpose of identifying the second entity is to further clean the epidemiological data to obtain the number of new cases, deaths, or total patients belonging to different regions. At the same time, it can be used as a regional search term to meet users' more refined search needs.

[0073] In practice, if both the base population and the subtype population exist, server 104 can first perform entity identification related to the second entity for the base population. Then, for one or more identified second entities, it can cleanse the base population to obtain the number of new cases, deaths, or total patients under the corresponding second entity, using at least one of the cleaned new cases, deaths, or total patients as historical epidemiological information. After obtaining the historical epidemiological information of the base population under the corresponding second entity, server 104 can further determine whether the epidemiological data also records the number of new cases, deaths, or total patients of the subtype population under the same second entity. If it does, it can directly cleanse to obtain the number of new cases, deaths, or total patients corresponding to different years; if it does not, it can be obtained through analysis based on other key information.

[0074] For example, the number of new lung cancer cases in China as the basic population, after data cleaning, can be obtained as follows: 2011 (651,053), 2012 (704,800), 2013 (732,800), 2014 (782,000), 2015 (787,000), and 2016 (828,100).

[0075] For example, after data cleaning, the number of newly diagnosed lung cancer / metastatic cases in China as a subtype population can be obtained as follows: 2011 (371, 100), 2012 (401, 736), 2013 (417, 696), 2014 (445, 740), 2015 (448, 590), and 2016 (472, 017).

[0076] In one embodiment, the epidemiological information retrieval method further includes: if the epidemiological data includes the proportion of new cases in a subtype population but not the proportion of new cases in the subtype population, then obtaining the new cases in the subtype population based on the proportion of new cases and the corresponding new cases in the basic population, as historical epidemiological information; and / or if the epidemiological data includes the proportion of deaths in a subtype population but not the proportion of deaths in the subtype population, then obtaining the deaths in the subtype population based on the proportion of deaths and the corresponding deaths in the basic population, as historical epidemiological information; and / or if the epidemiological data includes the proportion of the total number of cases in a subtype population but not the total number of cases in the subtype population, then obtaining the total number of cases in the subtype population based on the proportion of the total number of cases and the corresponding total number of cases in the basic population, as historical epidemiological information.

[0077] In specific implementation, this embodiment will explain in detail how the server 104 analyzes and obtains the historical epidemiological information of the subtype population based on other key information when the epidemiological data does not directly record the new cases, deaths and total number of cases of the subtype population.

[0078] Specifically, the number of new cases in the subtype population = the percentage of new cases in the subtype population * the number of new cases in the basic population; the number of deaths in the subtype population = the percentage of deaths in the subtype population * the number of deaths in the basic population; and the total number of cases in the subtype population = the percentage of total cases in the subtype population * the total number of cases in the basic population.

[0079] For example, the epidemiological data currently being analyzed includes a subtype of "lung cancer / metastasis" and a base population of "lung cancer" as its parent node. While the data does not explicitly record new cases of "lung cancer / metastasis," it does record new cases of "lung cancer" in China (values ​​for 2011-2016 are shown above), and also records that the proportion of new cases of "lung cancer / metastasis" in China is "0.57". Therefore, after data cleaning, the new cases of lung cancer / metastasis in China as a subtype can be obtained as follows: 2011 (371, 100=651, 053*0.57), 2012 (401, 736=704, 800*0.57), 2013 (417, 696=732, 800*0.57), 2014 (445, 740=782, 000*0.57), 2015 (448, 590 = 787,000 * 0.57, 2016 (472,017 = 828,100 * 0.57).

[0080] In one embodiment, determining at least one of a first predicted epidemic information, a second predicted epidemic information, and a third predicted epidemic information as predicted epidemic information based on historical epidemic information includes: calculating at least one of new cases, deaths, and total number of patients based on the numerical change rate of historical epidemic information as the first predicted epidemic information; and / or obtaining population potential information based on the first predicted epidemic information or the total number of new cases or patients as the second predicted epidemic information; and / or obtaining drug potential information based on the second predicted epidemic information as the third predicted epidemic information.

[0081] In practice, whether it is the basic population or the subtype population, the numerical change rate of historical epidemiological information can be obtained from the original epidemiological data or by analyzing and calculating the historical epidemiological information. However, if it has been recorded in the original text, the numerical change rate recorded in the original text will be used as the basis for subsequent analysis. Conversely, if it has not been recorded in the original text, the historical epidemiological information can be analyzed, and the numerical change rate can be calculated to determine the first predicted epidemiological information.

[0082] For example, the number of new lung cancer cases in China from 2011 to 2016 were: 2011 (651,053), 2012 (704,800), 2013 (732,800), 2014 (782,000), 2015 (787,000), and 2016 (828,100). Since the original text did not record the rate of change for these new cases, server 104 can calculate the 5-year average growth rate of new lung cancer cases in China as "4.93%" based on these values. This rate of change is then used as the rate of change required for subsequent analysis to predict the number of new lung cancer cases in China from 2024 to 2028, specifically as follows: 2024 (1,216,828), 2025 (1,276,799), 2026 (1,339,725), and 2027 (1, 405, 753), 2028 (1, 475, 035).

[0083] Furthermore, there are two methods for analyzing and obtaining population potential information. One is for non-oncology sub-indication populations, where population potential information = total number of patients * treatment cost (treatment cost / person / year); the other is for oncology sub-indication populations, where population potential information = new cases * diagnosis rate * penetration rate * treatment cost. Diagnosis rate and penetration rate can be obtained by cleaning the original data. Treatment cost can be determined by: for the same population, cleaning the original data to aggregate all available drug data, then determining the price of each drug and / or other prices, and summing them as the treatment cost. Drug potential information is determined by population potential information and market share. Market share determination needs to consider: the target drug's order of market launch in its respective market segment (e.g., large molecule, small molecule, etc.), its position in guidelines, the company that developed it, and its therapeutic position (the order in which doctors recommend medication), etc.

[0084] S203, send the target flow defect information to the client for display.

[0085] In practice, after the server 104 obtains the target disease information required by the user, it can send the target disease information to the client 102 so that the client 102 can display the target disease information to the user through its display screen, thereby satisfying the user's disease information retrieval needs.

[0086] It is understood that the display formats of target infectious disease information on client 102 include, but are not limited to: lists, grid views, summary cards, map views, charts, or graphs. A list arranges search results into a simple list according to sequence or importance; each item may include a title, summary, and link. A grid view displays search results in a grid format, with each item having a box containing a title, image, and brief description. Summary cards display each search result as a small card containing a title, summary, and link; cards are typically arranged vertically or horizontally, but can also be arranged diagonally or in a stacked manner, as not limited in specific embodiments of the present invention. A map view presents the results in a map format; for example, target infectious disease information can be displayed based on regions divided on a map. Charts or graphs can display target infectious disease information in chart, graph, or visualization formats.

[0087] Therefore, this invention proposes to segment and clean epidemiological information from the perspective of subdivided indications. This not only improves data accuracy (i.e., for subdivided indication granularity, data can be cleaned and organized more precisely, because different diseases or symptoms may have different definitions and classification standards, as well as corresponding data fields. The cleaning process can classify data according to specific diseases or symptoms, which can reduce data errors and confounding, and improve data accuracy), but also optimizes data analysis research (after subdividing data to the indication granularity, it is easier to conduct targeted data analysis and research, helping researchers better understand the differences between different subgroups and discover relevant factors or trends under a specific indication, which helps improve the effectiveness and reliability of research results), and supports drug development and evaluation (to help researchers better understand the clinical characteristics, patient needs, and market potential of target diseases, thereby guiding research and development strategies and decisions).

[0088] In the above-described method for retrieving epidemiological information, the server cleans epidemiological data step by step according to a preset disease system to construct an epidemiological information database. Then, it can receive and respond to epidemiological information retrieval requests carrying epidemiological information retrieval terms, filter out epidemiological information matching the epidemiological information retrieval terms from the pre-constructed epidemiological information database, and obtain target epidemiological information that can be sent to the client for display, so as to meet the user's retrieval needs for epidemiological information and improve the retrieval efficiency of epidemiological information.

[0089] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0090] To better implement the epidemiological information retrieval method provided in the embodiments of the present invention, based on the epidemiological information retrieval method proposed in the embodiments of the present invention, the embodiments of the present invention also provide an epidemiological information retrieval device, such as... Figure 3 As shown, the epidemic information retrieval device 300 includes: The request receiving module 310 is used to receive a request for retrieving epidemic information carrying epidemic information retrieval terms; The information retrieval module 320 is used to respond to the request for retrieval of epidemic disease information, and to filter out the epidemic disease information that matches the search terms in the pre-built epidemic disease information database as the target epidemic disease information; wherein, the epidemic disease information database is constructed by cleaning epidemiological data step by step according to the preset disease system; The information feedback module 330 is used to send target flow disease information to the client for display.

[0091] In one embodiment, the information retrieval module 320 is further configured to respond to an epidemiological information retrieval request, determine the retrieval term type of the epidemiological information retrieval term, and, based on the retrieval term type, filter out epidemiological information of the target population that matches the epidemiological information retrieval term from a pre-constructed epidemiological information database as target epidemiological information; wherein, the retrieval term type includes at least one of the following: title retrieval term, disease retrieval term, target retrieval term, region retrieval term, dimension retrieval term; the target population includes the basic population and / or subtype population; the epidemiological information includes historical epidemiological information and / or predicted epidemiological information.

[0092] In one embodiment, if the number of search terms for epidemic information is more than one, the information retrieval module 320 is further configured to: if the search term types of each epidemic information search term are consistent, then in the pre-built epidemic information database, filter out the epidemic information of the target population that matches each epidemic information search term, and take the union of each epidemic information to obtain the target epidemic information; if the search term types of each epidemic information search term are inconsistent, then in the pre-built epidemic information database, filter out the epidemic information of the target population that matches each epidemic information search term, and take the intersection of each epidemic information to obtain the target epidemic information.

[0093] In one embodiment, historical epidemic information includes at least one of new cases, deaths, and total number of patients; predicted epidemic information includes at least one of a first predicted epidemic information, a second predicted epidemic information, and a third predicted epidemic information; the first predicted epidemic information includes at least one of new cases, deaths, and total number of patients; the second predicted epidemic information includes population potential information; and the third predicted epidemic information includes drug potential information.

[0094] In one embodiment, the epidemic information retrieval device 300 further includes a database construction module for acquiring biomedical data; analyzing the biomedical data according to its data type to extract epidemiological data; and analyzing the epidemiological data based on a preset disease system to construct an epidemic information database.

[0095] In one embodiment, the database construction module is further configured to: if the biomedical data is text-based, analyze the biomedical data using a trained first model to extract epidemiological data; wherein the trained first model is obtained by fine-tuning a large-scale pre-trained language model; if the biomedical data is chart-based and the chart is standard, analyze the biomedical data using a trained second model to extract epidemiological data; wherein the trained second model is obtained by pre-training an initial language model based on entity alignment, a preset entity filling task, and a cloze test task; if the biomedical data is chart-based and the chart is non-standard, analyze the biomedical data using a trained third model to extract epidemiological data; wherein the trained third model is obtained by training a set of sample documents with pre-annotated chart areas covering title areas.

[0096] In one embodiment, the database construction module is further configured to perform entity identification on epidemiological data to obtain a first entity and a first entity category, wherein the first entity category includes at least one of disease, stage, pathology, biomarker, and baseline characteristics; combine the first entities according to the first entity category and a preset disease system to determine the base population and / or subtype population in the epidemiological data; perform data cleaning on the epidemiological data for the base population and / or subtype population to obtain at least one of new cases, deaths, and total number of patients as historical epidemiological information; determine at least one of first predicted epidemiological information, second predicted epidemiological information, and third predicted epidemiological information as predicted epidemiological information based on the historical epidemiological information; and summarize the predicted epidemiological information and / or historical epidemiological information to obtain epidemiological information to construct an epidemiological information database.

[0097] In one embodiment, the database construction module is further configured to perform entity identification on epidemiological data for the basic population and / or the subtype population to obtain a second entity; wherein the second entity is a region-related entity; based on the second entity, at least one of the new cases, deaths, and total number of patients in the epidemiological data is cleaned out as historical epidemiological information.

[0098] In one embodiment, the database construction module is further configured to: if the epidemiological data includes the proportion of new cases in the subtype population but does not include new cases in the subtype population, then obtain the new cases in the subtype population based on the proportion of new cases and the corresponding new cases in the basic population, as historical epidemiological information; and / or if the epidemiological data includes the proportion of deaths in the subtype population but does not include deaths in the subtype population, then obtain the deaths in the subtype population based on the proportion of deaths and the corresponding deaths in the basic population, as historical epidemiological information; and / or if the epidemiological data includes the proportion of the total number of cases in the subtype population but does not include the total number of cases in the subtype population, then obtain the total number of cases in the subtype population based on the proportion of the total number of cases and the corresponding total number of cases in the basic population, as historical epidemiological information.

[0099] In one embodiment, the database construction module is further configured to calculate at least one of new cases, deaths, and total number of patients based on the numerical change rate of historical epidemic information, as the first predicted epidemic information; and / or obtain population potential information based on the first predicted epidemic information or the total number of new cases or patients, as the second predicted epidemic information; and / or obtain drug potential information based on the second predicted epidemic information, as the third predicted epidemic information.

[0100] In the above embodiments, after the server cleans the epidemiological data step by step according to the preset disease system to build an epidemic disease information database, it can receive and respond to epidemic disease information retrieval requests carrying epidemic disease information retrieval terms. It can then filter out epidemic disease information that matches the epidemic disease information retrieval terms from the pre-built epidemic disease information database, thereby obtaining target epidemic disease information that can be sent to the client for display, so as to meet the user's retrieval needs for epidemic disease information and improve the retrieval efficiency of epidemic disease information.

[0101] It should be noted that the specific limitations regarding the epidemiological information retrieval device can be found in the limitations of the epidemiological information retrieval method described above, and will not be repeated here. Each module in the aforementioned epidemiological information retrieval device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independently of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.

[0102] In some embodiments of this application, the epidemic information retrieval device 300 can be implemented as a computer program, and the computer program can be implemented in, for example... Figure 4 The computer device shown operates on this device. The computer device's memory can store the various program modules that make up the disease information retrieval device 300, for example, Figure 3 The request receiving module 310, information retrieval module 320, and information feedback module 330 shown; the computer program composed of each program module causes the processor to execute the steps in the disease information retrieval method of the various embodiments of this application described in this specification. For example, Figure 4 The computer device shown can be used as follows Figure 3 The request receiving module 310 in the illustrated epidemiological information retrieval device 300 executes step S201. The computer device can execute step S202 via the information retrieval module 320. The computer device can execute step S203 via the information feedback module 330. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external computer devices via a network connection. When the computer program is executed by the processor, it implements an epidemiological information retrieval method.

[0103] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0104] In some embodiments of this application, a computer device is provided, including one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor as described in the resource deployment information retrieval method. The steps of the epidemiological information retrieval method described here may be steps from the epidemiological information retrieval methods of the various embodiments described above.

[0105] In some embodiments of this application, a computer-readable storage medium is provided, storing a computer program. The computer program is loaded by a processor, causing the processor to execute the steps of the above-described epidemic information retrieval method. The steps of the epidemic information retrieval method here can be the steps in the epidemic information retrieval methods of the various embodiments described above.

[0106] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0107] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0108] The above provides a detailed description of the method and apparatus for retrieving epidemiological information according to embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for retrieving epidemiological information, characterized in that, include: Receive requests for information on infectious diseases that carry search terms related to infectious diseases; Acquiring biomedical data; Based on the data type of the biomedical data, analyze the biomedical data to extract epidemiological data; The epidemiological data is analyzed based on a pre-defined disease system to construct an epidemiological information database, including: Entity identification is performed on the epidemiological data to obtain a first entity and a first entity category, wherein the first entity category includes at least one of disease, stage, pathology, biomarker, and baseline feature; Based on the first entity category and the preset disease system, the first entities are combined to determine the base population and / or subtype population in the epidemiological data; For the basic population and / or the subtype population, the epidemiological data is cleaned to obtain at least one of the following: new cases, deaths, and total number of patients, which is used as historical epidemiological information; Based on the historical epidemic information, at least one of the first predicted epidemic information, the second predicted epidemic information, and the third predicted epidemic information is determined as the predicted epidemic information; The first predicted epidemic information includes at least one of new cases, deaths, and total number of patients; the second predicted epidemic information includes population potential information; and the third predicted epidemic information includes drug potential information. By summarizing the predicted infectious disease information and / or the historical infectious disease information, the infectious disease information is obtained to construct the infectious disease information database; The disease system is obtained by combining information from at least one dimension, including disease, stage, pathology, biomarkers, and baseline characteristics. In response to the request for retrieval of epidemic information, the system filters out epidemic information that matches the search terms from a pre-built epidemic information database and uses it as the target epidemic information. The epidemic information database is constructed by cleaning epidemiological data step by step according to a preset disease system. The target flow defect information is sent to the client for display.

2. The method as described in claim 1, characterized in that, The step of responding to the epidemic information retrieval request by filtering out epidemic information matching the search terms from a pre-built epidemic information database as target epidemic information includes: In response to the epidemic information retrieval request, determine the retrieval term type of the epidemic information retrieval term; Based on the search term type, the target population's epidemic information that matches the epidemic information search term is selected from the pre-constructed epidemic information database and used as the target epidemic information; The search term types include at least one of the following: title search terms, disease search terms, clinical staging search terms, biomarker search terms, baseline feature search terms, pathology search terms, target search terms, region search terms, and dimension search terms; The target population includes the basic population and / or subtype population; the epidemic information includes historical epidemic information and / or predicted epidemic information.

3. The method as described in claim 1, characterized in that, The step of analyzing the biomedical data according to its data type and extracting epidemiological data includes: If the biomedical data is of text type, the epidemiological data is extracted by analyzing the biomedical data through a trained first model; wherein, the trained first model is obtained by training a large-scale pre-trained language model using a fine-tuning method; If the biomedical data is of chart type and the chart is of standard type, then the biomedical data is analyzed by the trained second model to extract the epidemiological data; wherein, the trained second model is obtained by pre-training the initial language model based on entity alignment, preset entity filling tasks and cloze test tasks; If the biomedical data is of chart type and the chart is non-standard, the biomedical data is analyzed by a trained third model to extract the epidemiological data; wherein the trained third model is trained based on a set of sample documents with pre-labeled chart areas covering the title area.

4. The method as described in claim 1, characterized in that, The epidemiological data is cleaned for the base population and / or the subtype population to obtain at least one of the following: new cases, deaths, and total number of patients, which is used as historical epidemiological information: For the base population and / or the subtype population, entity identification is performed on the epidemiological data to obtain a second entity; wherein, the second entity is a region-related entity; Based on the second entity, at least one of the following is cleaned from the epidemiological data: new cases, deaths, and total number of patients, which is used as the historical epidemiological information.

5. The method as described in claim 4, characterized in that, The method further includes: If the epidemiological data includes the proportion of new cases in the subtype population but does not include new cases in the subtype population, then based on the proportion of new cases and the corresponding new cases in the baseline population, the new cases in the subtype population are obtained as the historical epidemiological information; and / or If the epidemiological data includes the percentage of deaths in the subtype population but does not include deaths in the subtype population, then the deaths in the subtype population are obtained based on the percentage of deaths and the corresponding deaths in the baseline population, and are used as the historical epidemiological information; and / or If the epidemiological data includes the percentage of the total number of cases in the subtype population, but does not include the total number of cases in the subtype population, then the total number of cases in the subtype population is obtained based on the percentage of the total number of cases and the total number of cases in the corresponding base population, and is used as the historical epidemiological information.

6. The method as described in claim 1, characterized in that, The step of determining at least one of the first predicted infectious disease information, the second predicted infectious disease information, and the third predicted infectious disease information as predicted infectious disease information based on the historical infectious disease information includes: Based on the rate of change of the historical epidemiological information, at least one of the following—new cases, deaths, and total number of patients—is calculated and used as the first predicted epidemiological information in the predicted epidemiological information; and / or Based on the total number of new cases or patients in the first predicted epidemic information, population potential information is obtained and used as the second predicted epidemic information in the predicted epidemic information; and / or Based on the second predicted epidemic information, drug potential information is obtained as the third predicted epidemic information in the predicted epidemic information.

7. An epidemiological information retrieval device, characterized in that, include: The request receiving module is used to receive requests for retrieving epidemic information carrying keywords for epidemic information retrieval; Database building module: used to acquire biomedical data; Based on the data type of the biomedical data, analyze the biomedical data to extract epidemiological data; The epidemiological data is analyzed based on a pre-defined disease system to construct an epidemiological information database, including: Entity identification is performed on the epidemiological data to obtain a first entity and a first entity category, wherein the first entity category includes at least one of disease, stage, pathology, biomarker, and baseline feature; Based on the first entity category and the preset disease system, the first entities are combined to determine the base population and / or subtype population in the epidemiological data; For the basic population and / or the subtype population, the epidemiological data is cleaned to obtain at least one of the following: new cases, deaths, and total number of patients, which is used as historical epidemiological information; Based on the historical epidemic information, at least one of the first predicted epidemic information, the second predicted epidemic information, and the third predicted epidemic information is determined as the predicted epidemic information; The first predicted epidemic information includes at least one of new cases, deaths, and total number of patients; the second predicted epidemic information includes population potential information; and the third predicted epidemic information includes drug potential information. By summarizing the predicted infectious disease information and / or the historical infectious disease information, the infectious disease information is obtained to construct the infectious disease information database; The disease system is obtained by combining information from at least one dimension, including disease, stage, pathology, biomarkers, and baseline characteristics. The information retrieval module is used to respond to the epidemic information retrieval request and filter out the epidemic information that matches the epidemic information retrieval terms from the pre-constructed epidemic information database as the target epidemic information; wherein, the epidemic information database is constructed by cleaning epidemiological data step by step according to a preset disease system; The information feedback module is used to send the target flow defect information to the client for display.

Citation Information

Patent Citations

  • Rare disease epidemiological database construction method and system based on case registration and search engine

    CN114334171A

  • Disease information mining and retrieval method and device, electronic equipment and storage medium

    CN114400099A