Drug data analysis and retrieval method, device, electronic device and storage medium
Through automated aggregation and analysis of drug information, the problem of inefficient research on drug research and development competition data has been solved, and an efficient and accurate collection of drug information has been achieved, and a decision to establish drug research and development projects has been supported.
Patent Information
- Application Number
- CN202111413586.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The research on competitive data on traditional Chinese medicine research and development relies on manual sorting, which is inefficient, and the reliability and accuracy of the analysis results are poor, which cannot effectively support drug research and development project establishment decisions.
By obtaining drug information in various regions, including marketing information, registration information and clinical trial information, the drug identification information and enterprise information are aggregated to generate a collection of drug information, and an automated data analysis method is provided.
It improves the efficiency and accuracy of drug data analysis, reduces costs, provides comprehensive and reliable data support, and provides convenient data support for drug research and development competition analysis.
Smart Images

Figure CN114218269B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a drug data analysis and retrieval method, device, electronic device and storage medium. Background Art
[0002] For pharmaceutical companies, how to develop and launch suitable drug pipelines in a fiercely competitive environment is an eternal business focus.
[0003] For project establishment or strategic development departments, daily work involves researching drug information and making decisions about project establishment or business partnerships based on the results. Among this information, the competitive landscape of drug R&D is one of the most critical pieces of data.
[0004] Currently, there are no commercial databases supporting drug R&D competition data, and research on this data is largely conducted manually. This data is fraught with immense amounts of information, and the sources of drug information are fragmented, with information from these sources often overlapping and intersecting, yet not fully encompassed. Manual analysis of drug R&D competition within this vast amount of information is extremely time-consuming and inefficient, and often limited by data integrity and personal understanding, resulting in poor reliability and accuracy. Summary of the Invention
[0005] The present invention provides a drug data analysis and retrieval method, device, electronic device and storage medium to address the defects of the prior art in manually analyzing the competitive situation of drug research and development, which is extremely time-consuming and labor-intensive, extremely inefficient, and has poor reliability and accuracy of the analysis results.
[0006] The present invention provides a drug data analysis method, comprising:
[0007] Obtaining drug information in each region, wherein the drug information includes at least one of marketing information, registration information, and clinical trial information;
[0008] Based on the identification information and enterprise information of each drug in the drug information of each region, the drug information of each region is aggregated to obtain a drug information set.
[0009] According to a drug data analysis method provided by the present invention, the marketing information includes at least one of identification information and company information of each marketed drug;
[0010] The registration information includes at least one of the identification information, company information, acceptance number, review items and review conclusions of each registered drug;
[0011] The clinical trial information includes at least one of identification information, company information, trial stage and trial status of each trial drug.
[0012] According to a drug data analysis method provided by the present invention, obtaining drug information of each region includes:
[0013] Performing enterprise matching on the enterprise names of listed drugs in the listed drug data of any listed drug in any region to obtain the enterprise names in the listed drug data;
[0014] Based on the company name, determine the company information of any marketed drug in any region, wherein the company information includes the license holder information and the manufacturer information, or includes the company group information;
[0015] The enterprise information of any of the listed drugs is stored in the drug listing information of any region.
[0016] According to a drug data analysis method provided by the present invention, obtaining drug information of each region includes:
[0017] Determine the acceptance number of any drug based on the registration application data of any drug in any region, and determine the review items of any drug based on the acceptance number of any drug;
[0018] and / or, based on the registration application data of any of the drugs, determine the review conclusion of any of the drugs;
[0019] The review items and / or review conclusions of any of the drugs are stored in the drug registration information of any of the regions.
[0020] According to a drug data analysis method provided by the present invention, the drug information of each region is aggregated based on the drug identification information and company information of each drug in the drug information of each region to obtain a drug information set, including:
[0021] Based on the fact that the identification information of each drug in each region is the same and the enterprise information is partially the same, the drug information of each region is aggregated to obtain a drug information set.
[0022] According to a drug data analysis method provided by the present invention, based on the fact that the identification information of each drug in each region is the same and the company information is partially the same, the drug information of each region is aggregated to obtain a drug information set, including:
[0023] Based on the fact that the identification information of each registered drug and part of the enterprise information in the registration information of each region are the same, the registration information of each region is aggregated to obtain the registration association information table of each region;
[0024] Generate an initial information set based on the registration association information table of each area;
[0025] Based on the fact that the identification information and enterprise information of each drug in each region are the same, the remaining drug information of each region and the initial information set are aggregated to obtain the drug information set. The remaining drug information is the information in the drug information of the corresponding region excluding the drug registration information.
[0026] According to a drug data analysis method provided by the present invention, the drug information set includes the research and development progress of each drug, and the research and development progress is determined based on the following steps:
[0027] Determine the identification information of the target drug;
[0028] Searching the marketing information for data related to the identification information of the target drug; if so, determining the research and development progress of the target drug based on the marketing information of the target drug; otherwise, searching the registration information for data related to the identification information of the target drug;
[0029] If the registration information contains data related to the identification information of the target drug, the research and development progress of the target drug shall be determined based on the review items and / or review conclusions in the registration information of the target drug; otherwise, the research and development progress of the target drug shall be determined based on the trial phase and / or trial status in the clinical trial information of the target drug.
[0030] The present invention also provides a drug data retrieval method, comprising:
[0031] Get the target search term entered by the user;
[0032] The drug information and / or R&D progress corresponding to the target search term is obtained by screening from a drug information set, wherein the drug information set is determined based on any of the drug data analysis methods described above.
[0033] The present invention also provides a drug data analysis device, comprising:
[0034] a drug information acquisition unit, configured to acquire drug information in each region, wherein the drug information includes at least one of marketing information, registration information, and clinical trial information;
[0035] The information set construction unit is used to aggregate the drug information of each region based on the identification information and enterprise information of each drug in the drug information of each region to obtain a drug information set.
[0036] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any of the above-described drug data analysis methods or drug data retrieval methods are implemented.
[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described drug data analysis methods or drug data retrieval methods.
[0038] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned drug data analysis methods or drug data retrieval methods.
[0039] The drug data analysis and retrieval method, apparatus, electronic device, and storage medium provided by the present invention obtain at least one of the following: drug marketing information, registration information, and clinical trial information for each region. This information is then aggregated to form a drug information set. This method enables comprehensive and reliable regional drug information to be obtained, effectively improving the efficiency of regional drug information collection and reducing the cost of regional drug data analysis. Furthermore, the drug information set can provide data support for analyzing competitive conditions in drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 Schematic diagram of the process of the drug data analysis method provided by the present invention;
[0042] Figure 2 This is one of the flow charts of step 110 in the drug data analysis method provided by the present invention;
[0043] Figure 3 This is the second flow chart of step 110 in the drug data analysis method provided by the present invention;
[0044] Figure 4 It is a flowchart of the method for constructing a drug information set provided by the present invention;
[0045] Figure 5It is a flow chart of the information aggregation method provided by the present invention;
[0046] Figure 6 It is a flow chart of the method for determining the R&D progress provided by the present invention;
[0047] Figure 7 It is a flowchart of the drug data retrieval method provided by the present invention;
[0048] Figure 8 It is a structural schematic diagram of the drug data analysis device provided by the present invention;
[0049] Figure 9 It is a structural diagram of the drug data retrieval device provided by the present invention;
[0050] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0052] The basic information on drugs or targets pharmaceutical companies need to research is scattered across websites and databases in various countries. A single country can easily have over 10,000 pharmaceutical and R&D companies, making homogeneous competition a frequent occurrence in product development. Whether developing generic drugs or innovative drugs, there's a clear trend of overcrowding. A new target with a high probability of drug development can spark R&D efforts by hundreds of companies in just a few years. A single generic drug with proven efficacy often sees hundreds to thousands of companies launch its own copies. This overcrowding creates a massive amount of information to contend with when researching drugs or targets. Sorting through this vast amount of information to rank the progress of each company's pipeline is challenging and time-consuming. This is true even for the Chinese pharmaceutical industry, and the challenge of acquiring and compiling competitive intelligence data is even greater when viewed globally.
[0053] Thirdly, due to complex circumstances such as company name changes, joint R&D registrations, transfers of clinical or manufacturing approvals, and changes in marketing authorization, the company entity or entity name for the same drug in different data sources may change. Manually sorting out this process and understanding the context of corporate changes is a complex and arduous process.
[0054] Finally, the databases of drug regulatory authorities in various countries are often designed for management processes, not for the needs of industry users. Data obtained by industry users often needs to be converted into drug and target data for research. For example, the National Medical Products Administration (NMPA) manages marketed drugs in China by approval number. The same drug from the same company may have multiple, even dozens, of approval numbers, each with different information such as approval date and strength. For example, drug registration data in Country A is managed by the Center for Drug Evaluation (CDE) of the National Medical Products Administration (NMPA) by acceptance number. The same drug from the same company may have several, even dozens, of acceptance numbers. To fully understand a drug's R&D and registration progress, milestone history, and milestone dates, users must review all relevant acceptance numbers to draw a basic conclusion, which is extremely time-consuming and inefficient.
[0055] In summary, a large number of pharmaceutical company employees engage in this kind of primary drug information research daily. Even with access to commercial databases to export source data, the process remains purely manual, making it extremely inefficient. Simply researching basic information on a single drug may require reviewing and organizing thousands of data points from multiple data sources. This repetitive work requires a high level of experience and commitment, while also severely limiting employee development opportunities.
[0056] Therefore, the market urgently needs more efficient and automated drug R&D competition data analysis methods that can improve the work efficiency of pharmaceutical company staff.
[0057] In response to the above problems, an embodiment of the present invention provides a drug data analysis method. Figure 1 FIG. 1 is a flow chart of the drug data analysis method provided by the present invention, such as Figure 1 As shown, the method includes:
[0058] Step 110: Obtain drug information in each region, where the drug information includes at least one of marketing information, registration information, and clinical trial information.
[0059] Specifically, regional drug information refers to drug information for each country or region. For example, it may be drug information for Country A, Country B, or Country C, or even for Region D. This is not specifically limited in the present embodiment. The types of information included in drug information for different regions may be the same or different. For example, drug information for Country A may include marketing information, registration information, and clinical trial information, while drug information for Country B may only include marketing information.
[0060] Drug information for each region can be obtained from databases of national or regional drug regulatory authorities or commercial databases. For example, Country A could be China, Country B could be the United States, Country C could be Japan, and Region D could be the European Union. Drug information sources for China include, but are not limited to, the National Medical Products Administration (NMPA), the Center for Drug Evaluation (CDE) of the National Medical Products Administration, and the China Clinical Trial Registry (ChiCTR); drug information sources for the United States include, but are not limited to, Drugs@FDA, the Orange Book, the Purple Book, and the Center for Biologics Evaluation and Research (CBER); drug information sources for Japan include, but are not limited to, the official platform of the Pharmaceuticals and Medical Devices Agency (PDMA); and drug information sources for Europe include, but are not limited to, the Heads of Medicines Agencies (HMA), the European Medicines Agency (EMA), and the EU clinical trial database, EudraCT.
[0061] Data analysis of drug information in each region can be conducted from at least one of the following three aspects: market launch information, registration information, and clinical trial information:
[0062] Market information is used to identify marketed drugs. Marketed drugs are those that have been reviewed and approved by the national drug regulatory authorities and issued a drug production (or trial production) approval number or an import drug registration certificate. Specifically, this information may include the drug's name, strength, approval number, manufacturer, or marketing authorization holder.
[0063] Registration information represents information about registered drugs. Registered drugs are those for which registration applications have been submitted in accordance with legal procedures and relevant requirements, and for which the National Drug Administration has reviewed and issued an administrative licensing decision. Registration information may include, but is not limited to, the drug name, registration application category, and applicant.
[0064] Clinical trial information is used to represent information about drugs that are undergoing or have completed clinical trials. Clinical trial information may include, but is not limited to, the drug name, company information, trial phase, and trial status.
[0065] Step 120 , based on the identification information and company information of each drug in the drug information of each region, the drug information of each region is aggregated to obtain a drug information set.
[0066] Specifically, the drug information of each region originates from the official website or database platform of each country or region, and the information of these data sources is mutually intertwined and partially overlapped, but not completely contained. Taking China as an example, the marketing information, registration information and clinical trial information are respectively scattered in the National Medical Products Administration (NMPA), the Center for Drug Evaluation (CDE) of the National Medical Products Administration and the China Clinical Trial Registry (ChiCTR). At the same time, NMPA contains both the marketing information and the registration information of some drugs; CDE contains both the registration information and the clinical trial information of some drugs. Therefore, it is necessary to aggregate the drug information of each region and obtain a drug information set after aggregation. Here, the drug information set contains at least one of the drug marketing information, registration information and clinical trial information of each region.
[0067] Here, the identification information of each drug refers to information that can identify the drug. The identification information can specifically include the drug name, generic name, and dosage form of each drug, or the drug name, dosage form, and strength of each drug, which is not specifically limited in the embodiment of the present invention.
[0068] The enterprise information of each drug refers to information related to the enterprise involved in the drug, which may include manufacturer information, license holder information, historical enterprise information, etc., and may also include group information, etc.
[0069] It should be noted that the identification information and company information of each drug are included in at least one of the marketing information, registration information and clinical trial information of the drugs in each region. By obtaining at least one of the marketing information, registration information and clinical trial information of the drugs in each region, the identification information and company information of each drug can be obtained.
[0070] When aggregating drug information of each region, the identification information and company information of each drug in the drug information of each region can be used as a basis. Specifically, drugs with the same identification information and the same company information can be merged to obtain a drug information set.
[0071] Taking into account the complex situations such as company name change, cooperative R&D registration, transfer of clinical approval or production approval, change of marketing owner, etc., the company information of the same drug in different data sources will change. When aggregating the drug information of each region, it is also possible to merge drugs with the same identification information and partially the same company information to obtain a drug information set.
[0072] The resulting drug information collection includes at least one of: drug launch information, drug registration information, and drug clinical trial information for each region. The aggregation process also filters out overlapping data, resulting in complete, accurate, and non-duplicated basic drug data. Through query and search, users can quickly and easily understand the competitive status of a drug's R&D efforts in each region, significantly improving the efficiency of pharmaceutical company staff.
[0073] The drug data analysis method provided by an embodiment of the present invention obtains at least one of the following: drug marketing information, registration information, and clinical trial information for each region. This information is then aggregated to form a drug information set. This method enables comprehensive and reliable regional drug information to be obtained, effectively improving the efficiency of regional drug information collection and reducing the cost of regional drug data analysis. Furthermore, the drug information set can provide data support for analyzing competitive conditions in drug research and development.
[0074] Based on the above embodiment, the marketing information includes at least one of identification information and company information of each marketed drug;
[0075] Registration information includes at least one of the following: identification information, company information, acceptance number, review items, and review conclusions for each registered drug;
[0076] The clinical trial information includes at least one of the identification information of each trial drug, company information, trial stage and trial status.
[0077] Specifically, the listing information may include at least one of the listing drug's identification information and company information. Specifically, the listing drug's identification information may include the name, generic name, and dosage form of the listing drug. The corresponding drug name can be obtained based on the approval number and matched against a pre-established drug dictionary to obtain the standard drug name, generic name, and dosage form information. Drugs with identical identification information are considered to be the same drug.
[0078] The corporate information of marketed drugs may include information on license holders, manufacturers and historical companies.
[0079] Registration information may include at least one of the following: identification information, company information, acceptance number, review items, and review conclusions. The identification information of a registered drug can be used to obtain the corresponding drug name based on the acceptance number, and then matched against a pre-established drug dictionary to obtain the standard drug name, its generic name, and dosage form information.
[0080] The enterprise information of registered drugs can be obtained from CDE and NMPA based on the same acceptance number, and the enterprise information of the two platforms can be merged and cleaned, such as removing special characters and deduplication; the merged enterprise information can be matched with the constructed enterprise dictionary to obtain standard enterprise information.
[0081] One acceptance number corresponds to one piece of registration information, but the same drug from the same company may have multiple acceptance numbers.
[0082] Review items can be obtained based on the acceptance number. For example, the review items may be clinical application, production application or one-time import, etc.
[0083] The review conclusion information includes but is not limited to: approval for production, approval for supplementation, approval for re-registration, approval for one-time import, approval for technology transfer, approval for subpackaging or one-time import, etc.
[0084] Clinical trial information includes at least one of identification information for each trial drug, company information, trial phase, and trial status. Specifically, the identification information for each trial drug may include the name, generic name, and dosage form of the trial drug. The drug name can be obtained based on the clinical registration number and matched against a constructed drug dictionary to obtain the standardized drug name, generic name, and dosage form information.
[0085] If drug name information is not directly available, the drug name can be extracted based on the registered study name or the trial text related to the drug name through entity recognition, rule matching, etc. The obtained drug name is then matched against the constructed drug dictionary to obtain the standard drug name, its generic name, and dosage form information.
[0086] The company information of each test drug can be obtained by matching the acquired company name with a pre-built company dictionary to obtain a standard company name.
[0087] The trial phases of each investigational drug include: clinical phase I (e.g., Phase I, Phase Ib, Phase Ia, etc.), clinical phase I / II (e.g., Phase I / II, Phase Ib / II, etc.), clinical phase II (e.g., Phase II, Phase IIa, etc.), clinical phase II / III (Phase II / III, etc.), clinical phase III (Phase III, Phase IIIa, etc.), clinical phase IV (Phase IV, etc.), BE trial, and other.
[0088] Specifically, the "trial title" and "trial stage" of the clinical registration can be obtained from the original website. If the "trial stage" cannot be obtained directly, it can be extracted from the "trial title" and the extracted trial stage can be cleaned into a standard trial stage according to certain rules.
[0089] The trial status of each trial drug includes: ongoing (not yet recruited), ongoing (recruiting), ongoing (recruiting completed), completed, voluntarily suspended or terminated, stopped, etc. The trial status can be standardized based on the captured trial status to obtain the standard trial status. For example:
[0090] 1) If the collected test status begins with "active pause" or "active termination", then return to "active pause or termination";
[0091] 2) If the collected test status begins with "ordered to suspend" or "ordered to terminate", then return to "stopped", etc.
[0092] The method provided in the embodiment of the present invention obtains standardized drug information for each region by performing data cleaning on at least one of marketing information, registration information, and clinical trial information.
[0093] Based on any of the above embodiments, Figure 2 This is one of the flow charts of step 110 in the drug data analysis method provided by the present invention, such as Figure 2 As shown, step 110 specifically includes:
[0094] Step 111a, performing enterprise matching on the enterprise names of the listed drugs in the listed drug data of any listed drug in any region to obtain the enterprise names in the listed drug data;
[0095] Step 112a, based on the company name, determine the company information of any marketed drug in any region, the company information including the license holder information and the manufacturer information, or including the company group information;
[0096] Step 113a: Store the company information of any marketed drug into the drug market information of any region.
[0097] Specifically, the name of the listed pharmaceutical company of any listed drug in any region can be obtained based on the approval number, and matched in a pre-built enterprise dictionary to obtain a standard enterprise name; then, based on the standard enterprise name, the enterprise information of the listed drug in the region can be determined.
[0098] The manufacturer is the company that actually produces the drug, and the licensee is the holder of the drug marketing authorization, that is, the company that bears primary responsibility for the quality of the target drug throughout its life cycle. If the licensee's company name is empty, the manufacturer's information can be automatically filled in with the licensee's information.
[0099] A corporate group is a group consisting of multiple subsidiaries. For example, when a multinational pharmaceutical company launches the same drug in different countries, it is often marketed under the names of its subsidiaries in different countries. In this case, corporate information can include corporate group information.
[0100] Based on any of the above embodiments, step 112a may further include:
[0101] If a change in the licensee information of the approval number corresponding to any drug is detected, the licensee names before and after the change will be determined based on the licensee change information, and the licensee information will be updated to the changed licensee name, and the licensee name before the change will be stored in the historical enterprise information in the enterprise information.
[0102] Specifically, based on the same approval number, there may be changes in the license holder information. If a change in the license holder information of any drug corresponding to the approval number is detected, the changed license holder name will be used as the new license holder, and the license holder before the change will be stored in the historical enterprise.
[0103] The method provided by the embodiment of the present invention ensures the integrity and accuracy of enterprise information by determining the information of license holders, production enterprises, group enterprises and historical enterprises, thereby being able to provide complete and accurate listing information.
[0104] Based on any of the above embodiments, Figure 3 This is the second flow chart of step 110 in the drug data analysis method provided by the present invention. Figure 3 As shown, step 110 specifically includes:
[0105] Step 111b, based on the registration application data of any registered drug in any region, determining the acceptance number of the registered drug, and based on the acceptance number of the registered drug, determining the review items of the registered drug;
[0106] and / or, based on the registration application data of any registered drug, determine the review conclusion of the registered drug;
[0107] Step 112b: storing the review items and / or review conclusions of the registered drug into the drug registration information of the region.
[0108] Specifically, an acceptance number corresponds to a piece of application data for any registered drug in any region. Based on the acceptance number, the review items for the registered drug can be determined and information can be entered. For example, when the acceptance number begins with JT, the application item is JT, indicating "one-time import"; when the acceptance number begins with CQZ, JQZ, CSZ, or JSZ, the application item is S, indicating "application for production"; for other applications, the fourth character of the acceptance number is used as the application item value, such as L, indicating "application for clinical trials", etc.
[0109] Based on the registration application data of any registered drug, the review conclusion of the registered drug can be calculated in real time. The steps for real-time calculation of the review conclusion are as follows:
[0110] 1) Initial review conclusion information is: None;
[0111] 2) First, based on the registration application data of any registered drug obtained, determine the corresponding review conclusion (such as review conclusion A or review conclusion B), and then compare it with the stored review conclusion to determine whether there is a change. If there is a change, record the corresponding review conclusion and store it.
[0112] The review conclusion information includes but is not limited to: approval for production, approval for supplementation, approval for re-registration, approval for one-time import, approval for technology transfer, approval for subpackaging, approval for one-time import, etc.
[0113] After obtaining the review items and / or review conclusions of any registered drug, the review items and / or review conclusions of the registered drug shall be stored in the drug registration information of the region.
[0114] Based on any of the above embodiments, the review conclusion determination rule may be: according to the acceptance number of any registered drug, query the information related to the registered drug, and determine the review conclusion based on the relevant information obtained from the query. The specific rules may be as follows:
[0115] If the clinical trial notification issuance catalog information is obtained and the stored review conclusion information is not available, the review conclusion is determined to be clinical approval;
[0116] If the information of the marketed drug (including the technical review report and instructions) is obtained, and the stored review conclusion information is not available, the review conclusion is determined to be approved for production;
[0117] If the information of the old certificate for a specific drug being replaced with a new one is obtained, and the stored review conclusion information is "Not Available", when the header of the test acceptance number is JYHB, JYSB, JYZB, JYBB or JYFB field, the review conclusion will be determined as "Approved for Supplementation";
[0118] If the corresponding title contains "Drug Approval Opinion Notice", the review conclusion will be determined as "Not Approved". If the title also contains "Drug Approval Document", then when the acceptance number is longer than 4, if the acceptance number starts with JT, the review conclusion will be determined as "Approved for One-Time Import"; if the fourth digit of the acceptance number is "T", the review conclusion will be determined as "Approved for Technology Transfer".
[0119] The method provided in the embodiment of the present invention obtains the review items and / or review conclusions of any registered drug by performing data analysis on the registration application data of the registered drug, thereby obtaining complete registration information.
[0120] Based on any of the above embodiments, step 120 specifically includes:
[0121] Step 121 , based on the fact that the identification information of each drug in each region is the same and the company information is partially the same, the drug information of each region is aggregated to obtain a drug information set.
[0122] Specifically, due to the complex historical evolution of drugs, including joint R&D, transfers of clinical or manufacturing approvals, and the marketed goods holder system, company information in different data sources can change. This means that drugs with the same identification information in different regions may correspond to multiple company information due to changes in company information.
[0123] Therefore, the identification information of each drug in each region can be exactly the same, and the drug information of each region can be aggregated as long as there is one same enterprise information to obtain a drug information set.
[0124] Based on any of the above embodiments, Figure 4 This is a flow chart of the method for constructing a drug information set provided by the present invention. Figure 4 As shown, step 121 specifically includes:
[0125] Step 1211: Based on the fact that the identification information of each registered drug and the company information are identical in the registration information of each region, the registration information of each region is aggregated to obtain a registration association information table of each region;
[0126] Step 1212: Generate an initial information set based on the registration association information table of each area;
[0127] In step 1213, based on the fact that the identification information of each drug in each region is the same and the enterprise information is the same, the remaining drug information of each region and the initial information set are aggregated to obtain a drug information set. The remaining drug information is the information in the drug information of the corresponding region excluding the drug registration information.
[0128] Specifically, due to the complex historical evolution of drugs, which involve joint R&D, transfers of clinical or manufacturing approvals, and the marketed goods holder system, company information in different data sources can change. This means that a registered drug with the same identifying information may have multiple acceptance numbers associated with it due to changes in company information. Therefore, this registration information can be aggregated to create a registration association information table for each region.
[0129] The registration information for each region is based on the same identification information for each registered drug and partially the same company information. This means that if the identification information for each registered drug is exactly the same and there is only one piece of the same company information, the registration information for multiple acceptance numbers can be aggregated to obtain a registration association information table for each region. The identification information for each registered drug includes the drug name, generic name, and dosage form.
[0130] After obtaining the registration association information tables for each region, information can be aggregated to generate an initial information set. This initial information set includes the drug registration information for each region. Subsequent aggregation of drug information for each region can be performed based on the initial information set. Because the initial information set integrates the information of all relevant companies in each region during the drug registration and submission phase, it contains the most and most complete enterprise information, facilitating data aggregation when other information is stored in the drug information set. Furthermore, prioritizing information aggregation for each region's registration association information tables to generate the initial information set avoids the system burden of subsequent simultaneous calculations of larger amounts of data and improves computational efficiency.
[0131] After obtaining the initial information set, the remaining drug information for each region is aggregated with the initial information set to obtain the drug information set. The remaining drug information is the information in the corresponding region's drug information excluding drug registration information. For example, if the drug information for Country A can include Country A's marketing information, Country A's registration information, and Country A's clinical trial information, the remaining drug information will be the Country A's marketing information and Country A's clinical trial information. For another example, if the drug information for Country B can include Country B's marketing information and Country B's registration information, the remaining drug information will be the Country B's marketing information.
[0132] Specifically, when aggregating information, the identification information of each drug in each region is the same and the company information is partially the same. This means that all drugs in each region with the same name, generic name, dosage form and company information as long as there is one drug in common are aggregated into one piece of drug data, and this piece of drug data is stored in the drug information collection.
[0133] The drug data analysis method provided by the embodiments of the present invention first aggregates registration information from each region to obtain a registration association information table for each region; then, based on the registration association information table for each region, an initial information set is obtained; and finally, the initial information set is aggregated with the remaining drug information to obtain a drug information set. This method prioritizes aggregation of registration information from each region, avoiding the system burden of concurrently computing larger amounts of data and improving computational efficiency.
[0134] Based on any of the above embodiments, Figure 5 It is a flow chart of the information aggregation method provided by the present invention, such as Figure 5 As shown, the method specifically includes:
[0135] Step 510: Determine the target identification information in the information to be aggregated, and search the database for a lock status associated with the target identification information.
[0136] Step 520: If yes, wait for the lock state to be released;
[0137] Step 530: Otherwise, set the target identification information to a locked state, perform information aggregation of the information to be aggregated, and release the locked state of the target identification information after the information aggregation is completed.
[0138] Specifically, the information to be aggregated refers to the information that needs to be aggregated, and the target identification information refers to the generic name and dosage form corresponding to the target drug name. To prevent accidental data deletion and ensure data integrity and accuracy, before aggregation is performed, the target identification information in the information to be aggregated is confirmed to be locked before aggregation is performed. If the identification information is not locked, the target identification information is locked before aggregation is performed. The target identification information is released from the locked state after aggregation is completed.
[0139] When aggregating data in a distributed environment, multiple threads may process data simultaneously. This can lead to multiple deletions and merges of the same generic name and dosage form. During the first merge, a second deletion occurs simultaneously, causing a data merge error. To avoid this error, distributed locks can be used to ensure that the generic names and dosage forms processed by each thread at the same time do not overlap. The specific processing steps are as follows:
[0140] (1) When a thread processes new generic names, dosage forms, and companies that need to be updated, the generic names and dosage forms are stored as locks in Redis;
[0141] (2) When another thread receives a new generic name, dosage form, and company, it checks whether the lock in redis exists. If not, it directly uses this group of generic names and dosage forms as the lock and performs the following aggregation processing; if it exists, it determines whether this group of generic names and dosage forms has any intersection with the generic name and dosage form in the lock. If not, it adds another lock to the current generic name and dosage form; if so, it waits for the lock to be released and re-requests the lock processing at regular intervals. When the number of requests exceeds a certain number, the generic name, dosage form, and company are put back into the queue to be processed, releasing the current resources.
[0142] The method provided by the embodiment of the present invention avoids accidental deletion of data and ensures the integrity and accuracy of data by setting the locking state of target identification information and then performing information aggregation of the information to be aggregated.
[0143] Based on any of the above embodiments, the drug information set includes the research and development progress of each drug. Figure 6 It is a flow chart of the method for determining the R&D progress provided by the present invention, such as Figure 6 As shown, the method includes:
[0144] Determine the identification information of the target drug;
[0145] Searching the marketing information for data related to the identification information of the target drug; if so, determining the research and development progress of the target drug based on the marketing information; otherwise, searching the registration information for data related to the identification information of the target drug;
[0146] If the registration information contains data related to the identification information of the target drug, the research and development progress of the target drug shall be determined based on the review items and / or review conclusions in the registration information of the target drug; otherwise, the research and development progress of the target drug shall be determined based on the trial phase and / or trial status in the clinical trial information of the target drug.
[0147] Specifically, the R&D progress may be the target drug's highest R&D progress, for example, "marketed," "marketing application," "clinical approval," or "clinical application." The target drug refers to the drug for which the highest R&D progress needs to be determined. The target drug's identification information may include its generic name and dosage form.
[0148] First, check whether there is a target drug in the marketed information. If so, the target drug’s highest R&D progress can be determined as “marketed”; if the target drug is not found, check whether there is a target drug in the registration information. If so, the target drug’s highest R&D progress can be determined based on the target drug’s review items and review conclusions; if the target drug is not found in the registration information, check whether there is a target drug in the clinical trial information. If so, the target drug’s highest R&D progress can be determined based on the target drug’s clinical trial phase and trial status. If the target drug is not found in the clinical trial information, the target drug’s highest R&D progress can be determined as “not declared”.
[0149] The method provided in the embodiment of the present invention determines the maximum R&D progress of the target drug by sequentially searching for the identification information of the target drug in the marketing information, registration information and clinical trial information, thereby further improving the efficiency of drug data analysis and reducing manpower and time costs.
[0150] Based on any of the above embodiments, taking Country A as an example, the R&D progress of the target drug in Country A can be determined based on the following steps:
[0151] 1. First, judge based on drug listing information:
[0152] (1) Search the drug marketing information based on the generic name + dosage form information. If data can be found, the highest R&D progress is "marketed";
[0153] (2) If the corresponding generic name + dosage form information is not found, further search for drug registration information;
[0154] 2. Judgment based on drug registration information:
[0155] (1) Search the drug registration information based on the generic name + dosage form information. If data can be found, further judgment will be made based on the "Review Matters" and "Review Conclusions". If no data is found, directly search for drug clinical trial information;
[0156] (2) If the information in the “Application Matters” or “Review Conclusions” contains matters related to drug marketing, for example, if the “Application Matters” contains T (Technology Transfer), or the “Review Conclusions” contains “Approved for Production” or “Approved for Import”, the highest R&D progress is “Marketed”;
[0157] (3) If none of the above information is included, determine whether the information in "Application Matters" or "Review Conclusions" contains matters related to the drug's marketing application. For example, if "Application Matters" includes S (application for production) and there is no review conclusion, the highest R&D progress is "application for marketing";
[0158] (4) If none of the above information is included, determine whether the information in the “Application Matters” or “Review Conclusions” contains matters related to the drug’s approval for clinical trials. For example, if the “Review Conclusions” includes “Approved for Clinical Trials”, the highest R&D progress is “Approved for Clinical Trials”;
[0159] (5) If none of the above information is included, determine whether the information in "Application Matters" or "Review Conclusions" contains matters related to the drug's clinical application. For example, if "Application Matters" includes L (clinical application) and there is no review conclusion, the highest R&D progress is "clinical application";
[0160] 3. Judgment based on drug clinical trial information:
[0161] (1) Obtain the highest progress in Country A according to the following priorities:
[0162] I. If the "Experimental Stage" includes Clinical Phase IV, and the "Trial Status" does not include active suspension, termination, or halt, the highest R&D progress is "Marketed"; if the "Trial Status" includes the above information, the highest R&D progress is "Marketed (Inactive)";
[0163] II. If "Experimental Stage" is not "other" and "Experimental Status" does not include active suspension, termination, or suspension, the highest R&D progress is the R&D stage represented by the current "Experimental Stage." If "Experimental Status" includes the above information, the highest R&D progress is the Inactive status of the R&D stage represented by the current "Experimental Stage."
[0164] III. If the trial phase is "other" and the "trial status" does not include voluntary suspension, termination, or suspension, the highest R&D progress is "clinical research"; if the "trial status" includes the above information, the highest R&D progress is "clinical research (Inactive)";
[0165] (2) If the corresponding generic name + dosage form information is not found in the drug clinical information, it means that the target drug has not yet been declared, and the highest R&D progress is "no declaration".
[0166] The present invention also provides a drug data retrieval method. Figure 7 FIG. 1 is a flow chart of the drug data retrieval method provided by the present invention, such as Figure 7 As shown, the method includes:
[0167] Step 710, obtaining the target search term input by the user;
[0168] Step 720 , screening the drug information set to obtain drug information and / or R&D progress for each region corresponding to the target search term, wherein the drug information set is determined based on a drug data analysis method.
[0169] Specifically, after obtaining a drug information set according to the drug data analysis method of the above embodiment, a one-click drug search platform can be established. The target search term here can be a drug name and / or a target name. It should be noted that the target name can be obtained based on a pre-established drug-target association relationship to obtain the target name corresponding to the target drug. The drug information set can include both drug names and target names.
[0170] After obtaining the drug name and / or target name input by the user, the drug information and / or research and development progress corresponding to the drug name and / or target name can be filtered from the drug information set.
[0171] The drug data retrieval method provided by the embodiment of the present invention can quickly obtain the R&D competition status of target drugs in various regions based on any dimension or combination of dimensions of drugs and targets, thereby improving the efficiency of data retrieval.
[0172] Based on any of the above embodiments, in order to ensure the real-time and validity of the data, the data in the drug information set may be updated in real time based on changes in the acquired data.
[0173] If there is a change in the original acquired data (such as drug name, generic name, dosage form, company, etc.) in the drug information of each region, it will trigger the change of the drug information data of each region, and further update the drug information collection.
[0174] Specifically, the generic name, dosage form, and company of the data before the change are traversed in the drug information set, and all the queried data are deleted; then, based on the data after the change, they are merged in the drug information set according to the above-mentioned information aggregation method.
[0175] Accordingly, a change in the original data obtained based on at least one of the clinical trial information, registration information, and market launch information triggers the real-time calculation of the maximum R&D progress of the target drug.
[0176] Based on any of the above embodiments, the steps for generating a drug information set are as follows:
[0177] (1) Obtain drug marketing information, drug registration information, and clinical trial information of Country A;
[0178] Country A can be China, and obtain Chinese drug marketing information from NMPA, drug registration information from CDE and NMPA, and drug clinical trial information from ChiCTR and CDE.
[0179] (2) Obtain drug marketing information in Country B;
[0180] Country B can be the United States. Data on drugs marketed in the United States are registered on multiple platforms, such as Drugs@FDA, OrangeBook, Purple Book, and CBER (Center for Biologics Evaluation and Research). The drug information registered in these data sources intersects and partially overlaps with each other, but is not completely included.
[0181] Obtain US marketed drug data from Drugs@FDA, Orange Book, Purple Book, and CBER, including but not limited to: application number, generic name, dosage form, and company information;
[0182] Specifically, based on the information obtained from the above four different data sources, at a preset fixed time, the generic name information and dosage form information of the drug are collected from the four different data sources respectively, and matched with the preset generic name dictionary and dosage form dictionary respectively to obtain the standard generic name information and dosage form information;
[0183] Based on the acquired company information, standard company names and group names are matched against pre-built company dictionaries. For example, AstraZeneca, a leading global multinational pharmaceutical company headquartered in London, UK, has subsidiaries in different countries around the world responsible for local operations. The same drug from a multinational pharmaceutical company is often marketed under the names of subsidiaries in different countries. To avoid the situation where multiple companies may market the same drug during the subsequent merging of drug information tables due to differences in the company names of drug R&D applications in various countries, a group dimension is established to ensure the accuracy and completeness of the drug information collection data.
[0184] The data from four data sources with the same generic name, dosage form, and company were merged to obtain the U.S. drug launch information.
[0185] (3) Obtain drug marketing information in Region C;
[0186] Region C can be Europe. European drug marketing data is registered on multiple platforms, such as HMA (The Heads of Medicines Agencies) and EMA (European Medicines Agency). Drug information registered with HMA refers to drugs that have obtained marketing authorization in any EU member state, while drug information registered with EMA refers to drugs approved through the centralized approval process.
[0187] Obtain European marketed drug data from HMA and EMA, including but not limited to: product number, generic name, dosage form, and company information;
[0188] Specifically, based on the information obtained from the above two different data sources, at a preset fixed time, the generic name information, dosage form information, and company information of the drug are collected from the two different data sources respectively, and matched with the preset generic name dictionary, dosage form dictionary, and company dictionary respectively to obtain the standard generic name information, dosage form information, company information, and group information;
[0189] The data from the two data sources were merged according to the same generic name, same dosage form, and same company to obtain the European drug launch information.
[0190] (4) Obtain drug marketing information in Country D;
[0191] Country D can be Japan. Drug data for Japanese marketed products is registered on the PDMA (Pharmaceuticals and Medical Devices Agency) platform. Drugs marketed under the same generic name but in different dosage forms by the same manufacturer are registered in the same entry. Furthermore, if a reexamination or supplemental application occurs, an additional entry will be added for the same generic name and dosage form by the same manufacturer. Therefore, to ensure data accuracy and facilitate subsequent data calculations, the following are required:
[0192] I. Split data for the same generic name under the same enterprise according to different dosage forms;
[0193] II. Aggregate data based on the same generic name, same dosage form, and some of the same companies; ultimately generate information on drugs marketed in Japan;
[0194] Specifically, obtain data on Japanese marketed drugs from PMDA, including but not limited to: generic name information, dosage form information, and company information;
[0195] At a preset fixed time, collect the generic name information, dosage form information, and company information of the drug, and match them with the preset generic name dictionary, dosage form dictionary, and company dictionary respectively to obtain the standard generic name information, dosage form information, company information, and group information;
[0196] I. If there is more than one dosage form information in the acquired data, split the data according to different dosage forms;
[0197] II. Pre-aggregate the data based on the same generic name, same dosage form, and partially identical companies to obtain information on drugs marketed in Japan.
[0198] (5) Based on the fact that the identification information of each registered drug in the registration information of each region is the same and the company information is partially the same, the registration information of each region is aggregated to obtain the registration association information table of each region.
[0199] In drug registration information, the data can first be differentiated based on whether the "review item" begins with JT (one-time import); then the data can be aggregated based on the fact that the identification information of each registered drug is the same and the company information is partially the same. The reason is that one-time imported drugs may be for reference research or only for sales processing, and will not be included in the Chinese drug approval process system, that is, they will not be registered and there will be no subsequent application matters.
[0200] The registration association information table for each region can be displayed in the form shown in Table 1. For example, registered drugs with the drug name "nevirapine sustained-release tablets", the generic name "nevirapine", the dosage form "sustained-release tablets", and some of the same company information can be aggregated to obtain the registration association information table shown in Table 1.
[0201] Table 1
[0202]
[0203] (6) Generate an initial information set based on the registration association information table of each area.
[0204] Here, only the registration association information table exists in country A. Therefore, the initial information set is the registration association information table of country A. If registration association information tables exist in other regions, the registration association information tables of each region are aggregated to obtain the initial information set.
[0205] (7) Generate a drug information set based on the initial information set and the remaining drug information.
[0206] Here, the remaining drug information is the marketing information in Country A, clinical trial information in Country A, marketing information in Country B, marketing information in Region C, and marketing information in Country D.
[0207] The drug information set includes but is not limited to: drug name, generic name, dosage form, target, licensee, partner, manufacturer and historical enterprise. The drug information set for each region can be displayed in the form shown in Table 2, where:
[0208] ① Target: obtain the target name corresponding to the target drug based on the pre-established drug-target relationship;
[0209] ② The license holder is the license holder information or enterprise information in the enterprise information;
[0210] ③Manufacturer and historical enterprise are respectively the production enterprise information and historical enterprise information in the enterprise information;
[0211] ④Cooperative enterprises. Since the enterprise information of the licensed dealer will be covered by the enterprise name in the information form of each listed drug, cooperative enterprises are set up to ensure the integrity and traceability of enterprise data.
[0212] Table 2
[0213]
[0214] The specific information aggregation process is as follows:
[0215] I. The data in the initial information set is stored as the basis in the drug information set (the initial information set integrates the information of all relevant companies in the drug registration and declaration stage in each region. It involves the most and most complete company information, which is convenient for data merging when other information is stored in the drug information set). The "drug approval" mark is added and the company information is stored in the partner company (when the data is subsequently merged, the license holder information will be overwritten by the company information. To ensure data traceability, the company information can be stored separately in the partner company).
[0216] II. Furthermore, based on the same generic name, same dosage form, and same enterprise part (all enterprise names in each sub-table and all enterprise names in the drug information master table have the same enterprise name), the data will be merged into the drug information set and the corresponding source tag will be added; if the enterprise information includes enterprise group information, the enterprise group information will be aggregated.
[0217] If the data in all drug marketing information tables can be merged with the data already stored in the drug information set, the company information in each drug marketing information table will overwrite the original licensee information. In this embodiment of the present invention, the licensee information in Country A's drug marketing information is preferably used as the final licensee information. Therefore, if the target drug has been marked as "Marketed in Country A", the company information in other drug marketing information tables will no longer overwrite the licensee information and will be stored in the partner company.
[0218] The company information in the drug clinical information table of country A and the drug marketing information table of country D may contain multiple company names. To avoid overwriting of company information after merging, the information will be stored separately in the cooperative company after merging.
[0219] When merging the remaining information data with the drug information set data, during the data merging process of Country A's drug marketing information, Country A's drug clinical information, and Country D's drug marketing information, since there are multiple company names in the company information, the data can be sorted during the merging process and then merged to ensure the accuracy of the final merged data; wherein, the sorting method is: aggregate the data based on the same generic name and the same dosage form, ① arrange them from most to least according to the number of aggregated companies, ② if the number of companies is the same and ≥2, further search in the synthesized data based on the generic name, dosage form, and company, and arrange them from most to least according to the number of data items found.
[0220] III. If the merger is not possible, store the data directly in the drug information collection and add the corresponding source tag.
[0221] The drug data analysis device provided by the present invention is described below. The drug data analysis device described below and the drug data analysis method described above can be referenced to each other.
[0222] Figure 8 This is a schematic diagram of the structure of the drug data analysis device provided by the present invention. Figure 8 As shown, the device includes:
[0223] a drug information acquisition unit 810 for acquiring drug information in each region, wherein the drug information includes at least one of marketing information, registration information, and clinical trial information;
[0224] The information set construction unit 820 is used to aggregate the drug information of each region based on the identification information and company information of each drug in the drug information of each region to obtain a drug information set.
[0225] The apparatus provided in an embodiment of the present invention obtains at least one of the following: drug marketing information, registration information, and clinical trial information for each region. This information is then aggregated to form a drug information set. This method enables comprehensive and reliable regional drug information to be obtained, effectively improving the efficiency of regional drug information collection and reducing the cost of regional drug data analysis. Furthermore, the drug information set can provide data support for analyzing competitive conditions in drug research and development.
[0226] Based on any of the above embodiments, the marketing information includes at least one of identification information and company information of each marketed drug;
[0227] The registration information includes at least one of the identification information, company information, acceptance number, review items and review conclusions of each registered drug;
[0228] The clinical trial information includes at least one of identification information, company information, trial stage and trial status of each trial drug.
[0229] Based on any of the above embodiments, the drug information acquisition unit 810 is used to:
[0230] Perform enterprise matching on the marketed drug data of any marketed drug in any region to obtain the company name in the marketed drug data;
[0231] Based on the company name, determine the company information of any marketed drug in any region, wherein the company information includes the license holder information and the manufacturer information, or includes the company group information;
[0232] The enterprise information of any of the listed drugs is stored in the drug listing information of any region.
[0233] Based on any of the above embodiments, the drug information acquisition unit 810 is used to:
[0234] Determine the acceptance number of any registered drug based on the registration application data of any registered drug in any region, and determine the review items of any registered drug based on the acceptance number of any registered drug;
[0235] and / or, determining the review conclusion of any registered drug based on the registration application data of the any registered drug;
[0236] The review items and / or review conclusions of any registered drug are stored in the drug registration information of any region.
[0237] Based on any of the above embodiments, the information set construction unit 820 is configured to:
[0238] Based on the fact that the identification information of each drug in each region is the same and the enterprise information is partially the same, the drug information of each region is aggregated to obtain a drug information set.
[0239] Based on any of the above embodiments, the information set construction unit 820 is configured to:
[0240] Based on the fact that the identification information of each registered drug and part of the enterprise information in the registration information of each region are the same, the registration information of each region is aggregated to obtain the registration association information table of each region;
[0241] Generate an initial information set based on the registration association information table of each area;
[0242] Based on the fact that the identification information and enterprise information of each drug in each region are the same, the remaining drug information of each region and the initial information set are aggregated to obtain the drug information set. The remaining drug information is the information in the drug information of the corresponding region excluding the drug registration information.
[0243] Based on any of the above embodiments, the device further includes a research and development progress determination unit, which is configured to:
[0244] Determine the identification information of the target drug;
[0245] Searching the marketing information for data related to the identification information of the target drug; if so, determining the research and development progress of the target drug based on the marketing information; otherwise, searching the registration information for data related to the identification information of the target drug;
[0246] If the registration information contains data related to the identification information of the target drug, the research and development progress of the target drug shall be determined based on the review items and / or review conclusions in the registration information of the target drug; otherwise, the research and development progress of the target drug shall be determined based on the trial phase and / or trial status in the clinical trial information of the target drug.
[0247] Figure 9 This is a schematic diagram of the structure of the drug data retrieval device provided by the present invention. Figure 9 As shown, the device includes:
[0248] The search term acquisition unit 910 is used to acquire a target search term input by a user;
[0249] The information screening unit 920 screens the drug information and / or R&D progress corresponding to the target search term from the drug information set, where the drug information set is determined based on a drug data analysis method.
[0250] The drug data retrieval device provided by the present invention can quickly obtain the R&D competition status of target drugs in various regions based on any dimension or combination of dimensions of drugs and targets, thereby improving the efficiency of data retrieval.
[0251] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 may call logic instructions in the memory 1030 to execute a drug data analysis method or a drug data retrieval method, wherein the drug data analysis method includes: obtaining drug information in each region, wherein the drug information includes at least one of marketing information, registration information, and clinical trial information; and aggregating the drug information in each region based on the identification information and company information of each drug in the drug information in each region to obtain a drug information set.
[0252] The drug data retrieval method includes: obtaining a target search term input by a user; screening drug information and / or R&D progress corresponding to the target search term from a drug information set, wherein the drug information set is determined based on a drug data analysis method.
[0253] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0254] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the drug data analysis method or drug data retrieval method provided by the above methods, wherein the drug data analysis method includes: obtaining drug information in each region, the drug information including at least one of marketing information, registration information and clinical trial information; based on the identification information and enterprise information of each drug in the drug information of each region, aggregating the drug information of each region to obtain a drug information set.
[0255] The drug data retrieval method includes: obtaining a target search term input by a user; screening drug information and / or R&D progress corresponding to the target search term from a drug information set, wherein the drug information set is determined based on a drug data analysis method.
[0256] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the drug data analysis method or drug data retrieval method provided by the above-mentioned methods, wherein the drug data analysis method includes: obtaining drug information in each region, wherein the drug information includes at least one of marketing information, registration information and clinical trial information; based on the identification information and enterprise information of each drug in the drug information of each region, aggregating the drug information of each region to obtain a drug information set.
[0257] The drug data retrieval method includes: obtaining a target search term input by a user; screening drug information and / or R&D progress corresponding to the target search term from a drug information set, wherein the drug information set is determined based on a drug data analysis method.
[0258] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0259] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0260] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A drug data retrieval method, characterized in that: include: Get the target search term entered by the user; The target search term is a drug name and / or a target name; The drug information and / or R&D progress corresponding to the target search term is obtained by screening from a drug information set, wherein the drug information set is determined based on the following steps: Obtaining drug information in each region, including marketing information, registration information, and clinical trial information; Based on the identification information and company information of each drug in each region, the drug information of each region is aggregated to obtain a drug information set; The aggregation of drug information in each region includes: Determine the target identification information in the information to be aggregated, and search the database for a lock state associated with the target identification information; if so, wait for the lock state to be released; otherwise, set the lock state of the target identification information, perform information aggregation of the information to be aggregated, and release the lock state of the target identification information after the information aggregation is completed; the target identification information refers to the generic name and dosage form corresponding to the target drug name; The marketing information includes the identification information and company information of each marketed drug; the registration information includes the identification information, company information, acceptance number, review items and review conclusions of each registered drug; the clinical trial information includes the identification information, company information, trial stage and trial status of each trial drug; The obtaining of drug information in each region includes: Performing enterprise matching on the names of listed drug companies in the listed drug data of any listed drug in any region to obtain the company names in the listed drug data; determining the company information of any listed drug in any region based on the company names, the company information including the licensee information and the manufacturing company information, or including the company group information; storing the company information of any listed drug in the drug listing information of any region; if a change in the licensee information corresponding to the approval number of any drug is detected, determining the licensee names before and after the change based on the licensee change information, updating the licensee information to the changed licensee name, and storing the licensee name before the change in the historical company information in the company information; Determining, based on the registration application data of any registered drug in any region, the acceptance number of the registered drug, and determining, based on the acceptance number of the registered drug, the review items of the registered drug; and / or, based on the registration application data of the registered drug, determining the review conclusion of the registered drug; and storing the review items and / or review conclusion of the registered drug in the drug registration information of the registered drug in the region; The drug information set includes the research and development progress of each drug, which is determined based on the following steps: Determine the identification information of the target drug; search the marketing information for data related to the identification information of the target drug; if the marketing information contains data related to the identification information of the target drug, determine the research and development progress of the target drug based on the marketing information of the target drug; otherwise, search the registration information for data related to the identification information of the target drug; if the registration information contains data related to the identification information of the target drug, determine the research and development progress of the target drug based on the review items and / or review conclusions in the registration information of the target drug; otherwise, determine the research and development progress of the target drug based on the trial phase and / or trial status in the clinical trial information of the target drug.
2. The drug data retrieval method according to claim 1, characterized in that: The drug information of each region is aggregated based on the drug identification information and company information of each drug in the drug information of each region to obtain a drug information set, including: Based on the fact that the identification information of each drug in each region is the same and the enterprise information is partially the same, the drug information of each region is aggregated to obtain a drug information set.
3. The drug data retrieval method according to claim 2, characterized in that: Based on the fact that the identification information of each drug in each region is the same and the company information is partially the same, the drug information of each region is aggregated to obtain a drug information set, including: Based on the fact that the identification information of each registered drug and part of the enterprise information in the registration information of each region are the same, the registration information of each region is aggregated to obtain the registration association information table of each region; Generate an initial information set based on the registration association information table of each area; Based on the fact that the identification information of each drug in each region is the same and the enterprise information is partially the same, the remaining drug information of each region and the initial information set are aggregated to obtain the drug information set. The remaining drug information is the information in the drug information of the corresponding region except the drug registration information.
4. A drug data retrieval device, characterized in that: include: A search term acquisition unit, used to acquire a target search term input by a user; The target search term is a drug name and / or a target name; An information screening unit, configured to screen the drug information set to obtain drug information and / or R&D progress corresponding to the target search term; A drug information acquisition unit, configured to acquire drug information in each region, including marketing information, registration information, and clinical trial information; An information set building unit is used to aggregate the drug information of each region based on the identification information and company information of each drug in the drug information of each region to obtain a drug information set; The aggregating of drug information in each region includes: determining target identification information in the information to be aggregated, and searching in a database whether there is a lock state associated with the target identification information; if so, waiting for the lock state to be released; otherwise, setting the lock state of the target identification information, performing information aggregation of the information to be aggregated, and releasing the lock state of the target identification information after the information aggregation is completed; the target identification information refers to the generic name and dosage form corresponding to the target drug name; The marketing information includes the identification information and company information of each marketed drug; the registration information includes the identification information, company information, acceptance number, review items and review conclusions of each registered drug; the clinical trial information includes the identification information, company information, trial stage and trial status of each trial drug; The drug information acquisition unit is used to: Performing enterprise matching on the marketed drug data of any marketed drug in any region to obtain the enterprise name in the marketed drug data; determining the enterprise information of any marketed drug in any region based on the enterprise name, wherein the enterprise information includes the licensee information and the manufacturing enterprise information, or includes the enterprise group information; storing the enterprise information of any marketed drug in the drug marketed information of any region; if a change in the licensee information corresponding to the approval number of any drug is detected, determining the licensee names before and after the change based on the licensee change information, updating the licensee information to the changed licensee name, and storing the licensee name before the change in the historical enterprise information in the enterprise information; Determining, based on the registration application data of any registered drug in any region, the acceptance number of the registered drug, and determining, based on the acceptance number of the registered drug, the review items of the registered drug; and / or, based on the registration application data of the registered drug, determining the review conclusion of the registered drug; and storing the review items and / or review conclusion of the registered drug in the drug registration information of the registered drug in the region; The drug data retrieval device also includes a research and development progress determination unit for: Determine the identification information of the target drug; search the marketing information for data related to the identification information of the target drug; if data related to the identification information of the target drug exists in the marketing information, determine the research and development progress of the target drug based on the marketing information of the target drug; otherwise, search the registration information for data related to the identification information of the target drug; if data related to the identification information of the target drug exists in the registration information, determine the research and development progress of the target drug based on the review items and / or review conclusions in the registration information of the target drug; otherwise, determine the research and development progress of the target drug based on the trial phase and / or trial status in the clinical trial information of the target drug.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the drug data retrieval method according to any one of claims 1 to 3 are implemented.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the drug data retrieval method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Medicine information library building method based on network crawler
CN106777165A
Intelligent big data analysis method and system for drug group purchase
CN113034104A