Method, device, equipment and medium for analyzing pharmaceutical patent information based on large models

Through the large-scale model-based pharmaceutical patent information analysis method, the key information in pharmaceutical patents is automatically parsed, which solves the problem of time-consuming and labor-intensive manual analysis and achieves efficient and low-cost information extraction and R&D direction guidance.

CN119476258BActive Publication Date: 2025-09-19SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411669913.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-09-19
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing technologies for extracting and parsing key information in pharmaceutical patents rely on manual reading, which is time-consuming and labor-intensive, and requires high professional qualities of professionals, resulting in waste of human resources and low efficiency in information analysis.

Method used

A large-scale model-based pharmaceutical patent information analysis method is used to automatically analyze pharmaceutical patents by fine-tuning the pre-trained model to obtain key information such as targets, indications, and applicant companies. The target R&D value is determined based on the domestic and international R&D status and company attention.

Benefits of technology

It realizes the automation and rapid analysis of pharmaceutical patent information, saves labor costs and time, reduces the requirements for professional quality, and provides timely guidance on R&D direction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476258B_ABST
    Figure CN119476258B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, equipment, and medium for analyzing pharmaceutical patent information based on a large model, which relates to the field of artificial intelligence technology, including: inputting the pharmaceutical patent uploaded by the current user end into the fine-tuned large model to parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; determining the disease field of the current pharmaceutical patent based on the indication in the target key information, and obtaining the highest domestic and foreign R&D status of the current pharmaceutical patent from relevant websites based on the disease field and the target in the target key information; determining the level of the target based on the highest domestic and foreign R&D status, and judging whether the applicant company of the current pharmaceutical patent is a company of key concern to obtain a judgment result; judging whether the target has R&D value based on the judgment result and the target level. This application greatly improves the efficiency of pharmaceutical patent information analysis, saves the manpower and time costs of pharmaceutical patent analysis, and avoids waste of human resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for analyzing pharmaceutical patent information based on a large model. Background Art

[0002] Currently, extracting and parsing key information (such as targets and indications) from pharmaceutical patents to analyze the current situation and provide guidance to companies on R&D direction relies heavily on the manual reading and statistical analysis skills of specialized professionals. On average, each person reviews three to five patents per week. This process not only relies on the technical expertise of professionals but is also time-consuming and labor-intensive. Furthermore, because pharmaceutical patents are highly specialized and written in multiple languages, they require a high level of professionalism. Manual analysis consumes significant human resources, increasing the labor cost of patent information analysis and hindering timely guidance on R&D directions. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a large-scale model-based pharmaceutical patent information analysis method, device, equipment, and medium that can automatically extract and organize key information from pharmaceutical patents and analyze their current status, significantly improving the efficiency of patent information analysis, saving the manpower and time costs of patent analysis, and avoiding the waste of human resources. At the same time, it reduces the professional quality requirements for professionals, thereby providing timely and accurate guidance for enterprises to determine their R&D directions. The specific plan is as follows:

[0004] In a first aspect, the present application discloses a method for analyzing pharmaceutical patent information based on a large model, comprising:

[0005] Get the pharmaceutical patent uploaded by the current user and get the current pharmaceutical patent;

[0006] Input the current pharmaceutical patent into the fine-tuned large model, and use preset prompt information to parse the key information in the current pharmaceutical patent to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset; the target key information includes the target, indication, and the applicant company of the current pharmaceutical patent;

[0007] Determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain the target disease field, and obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information;

[0008] Determine the level of the target based on the highest domestic and international R&D status, obtain the target level, and judge whether the applicant company of the current pharmaceutical patent is a company of key concern, and obtain a judgment result;

[0009] Determine whether the target in the target disease field has research and development value based on the judgment result and the target level.

[0010] Optionally, obtaining the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information includes:

[0011] Utilize a web crawler and crawl the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information.

[0012] Optionally, the large model-based pharmaceutical patent information analysis method further includes:

[0013] Collect different types of historical pharmaceutical patents to obtain a pharmaceutical patent dataset;

[0014] Convert all historical pharmaceutical patents in the pharmaceutical patent dataset into a unified data format to obtain a converted dataset;

[0015] Text recognition is performed on each pharmaceutical patent in the converted data set to obtain historical patent text information, and the historical patent text information is input into a pre-trained model for model fine-tuning to obtain the fine-tuned large model; the pre-trained model is located in the inference server.

[0016] Optionally, performing text recognition on each pharmaceutical patent in the converted data set to obtain historical patent text information includes:

[0017] Optical character recognition tools are used to identify the text information of each pharmaceutical patent in the converted data set to obtain historical patent text information.

[0018] Optionally, the current pharmaceutical patent is input into the fine-tuned large model to analyze the key information in the current pharmaceutical patent using preset prompt information to obtain target key information, including:

[0019] Convert the current pharmaceutical patent into an image to obtain a target image, and cut the target image into multiple sub-images;

[0020] Text recognition is performed on each of the sub-images to obtain multiple current patent text information, and the current patent text information is input into the fine-tuned large model in turn to analyze the key information in the current medical patent using preset prompt information to obtain the target key information.

[0021] Optionally, determining whether the target in the target disease field has research and development value based on the judgment result and the target level includes:

[0022] Obtaining the locations where the target and the indication appear in the current pharmaceutical patent to obtain first location information, and counting the number of times the target and the indication appear in the current pharmaceutical patent to obtain number information;

[0023] Obtain the position where the applicant company appears in the current pharmaceutical patent to obtain second position information, and determine whether the target in the target disease field has research and development value based on the judgment result, the target level, the first position information, the number information and the second position information.

[0024] Optionally, the large model-based pharmaceutical patent information analysis method further includes:

[0025] The target key information, the target disease field, the highest domestic and international R&D status, the target level, the judgment result and the judgment result of whether it has R&D value are sent to the user terminal for structured display on the human-computer interaction interface of the user terminal.

[0026] In a second aspect, the present application discloses a large-scale model-based pharmaceutical patent information analysis device, comprising:

[0027] The patent acquisition module is used to obtain the pharmaceutical patent uploaded by the current user and obtain the current pharmaceutical patent;

[0028] A parsing module is configured to input the current pharmaceutical patent into a fine-tuned large model, and to parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning a pre-trained model using a pharmaceutical patent dataset; the target key information includes the target, indication, and the applicant company of the current pharmaceutical patent;

[0029] A field determination module is used to determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain a target disease field;

[0030] A status acquisition module is used to obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information;

[0031] A level determination module, configured to determine the level of the target based on the highest domestic and international R&D status, thereby obtaining the target level;

[0032] A judgment module, configured to judge whether the applicant company of the current pharmaceutical patent is a company of key concern, and obtain a judgment result;

[0033] The R&D value determination module is used to determine whether the target in the target disease field has R&D value based on the determination result and the target level.

[0034] In a third aspect, the present application discloses an electronic device comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the aforementioned large model-based pharmaceutical patent information analysis method is implemented.

[0035] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned large-model-based pharmaceutical patent information analysis method.

[0036] It can be seen that this application first obtains the pharmaceutical patent uploaded by the current user terminal to obtain the current pharmaceutical patent, and then inputs the current pharmaceutical patent into the fine-tuned large model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, so as to use the preset prompt information to parse the key information in the current pharmaceutical patent to obtain the target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, wherein the target key information includes the target, indication and the applicant company of the current pharmaceutical patent. Then, according to the indication in the target key information, the disease field of the current pharmaceutical patent is determined to obtain the target disease field, and according to the target disease field and the target in the target key information, the highest domestic and foreign R&D status of the current pharmaceutical patent is obtained from the relevant website, and then based on the highest domestic and foreign R&D status, the level of the target is determined to obtain the target level, and it is judged whether the applicant company of the current pharmaceutical patent is a key focus company to obtain a judgment result, and finally, according to the judgment result and the target level, it is determined whether the target in the target disease field has R&D value. This application first analyzes the current pharmaceutical patent input by the user through a fine-tuned large model to obtain the target key information including targets, indications and the applicant company of the current pharmaceutical patent, and then determines the disease field of the current pharmaceutical patent, the highest domestic and foreign R&D status, the target level and whether it belongs to a key company based on the target key information. Finally, based on the determined information, it determines whether the current target has R&D value. This application uses large model technology to deeply understand and analyze pharmaceutical patents, and realizes the extraction, collation and current status analysis of key information in automated pharmaceutical patents. Compared with traditional manual patent information analysis methods, it not only greatly improves the efficiency of patent information analysis, but also saves the manpower and time costs of patent analysis, avoids the waste of human resources, and at the same time, reduces the professional quality requirements for professionals, and greatly shortens the time and confirmation cycle of patent key information statistics, thereby providing timely and accurate guidance for enterprises to determine the R&D direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0038] Figure 1 This is a flow chart of a method for analyzing pharmaceutical patent information based on a large model disclosed in this application;

[0039] Figure 2This is a flowchart of a specific method for analyzing pharmaceutical patent information based on a large model disclosed in this application;

[0040] Figure 3 This is a schematic diagram of the structure of a large-scale model-based pharmaceutical patent information analysis device disclosed in this application;

[0041] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0043] The present application embodiment discloses a method for analyzing pharmaceutical patent information based on a large model, see Figure 1 As shown, the method includes:

[0044] Step S11: Obtain the pharmaceutical patent uploaded by the current user terminal to obtain the current pharmaceutical patent.

[0045] In this embodiment, the pharmaceutical patent currently uploaded by the user through the patent upload portal in the human-computer interaction interface on the user terminal is first retrieved to obtain the current pharmaceutical patent. It should be noted that the pharmaceutical patents uploaded by the user terminal can be either a single patent or multiple patents uploaded in batches. In other words, this solution supports batch analysis of multiple pharmaceutical patent information. Furthermore, the uploaded pharmaceutical patents can be backed up for subsequent access and manual analysis.

[0046] Step S12: Input the current pharmaceutical patent into the fine-tuned large model to parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset; the target key information includes the target, indication and the applicant company of the current pharmaceutical patent.

[0047] In this embodiment, after obtaining the pharmaceutical patent uploaded by the current user, the current pharmaceutical patent is further input into a large model (i.e., the fine-tuned large model) obtained by fine-tuning a pre-trained model using a pharmaceutical patent dataset. The large model is then questioned using semantic information provided by preset prompts in the model to parse key information in the current pharmaceutical patent and, in accordance with preset requirements, provide key information including the target, indication, and the applicant company of the current pharmaceutical patent. The preset prompts include, but are not limited to, semantic information regarding the target, indication, and patent-owning company. Targets include, but are not limited to, single targets, simultaneous targets, and multiple targets. Indications can be specifically categorized into 18 major disease areas, such as cardiovascular diseases, infectious diseases, tumors, and endocrine and metabolic diseases. Indications can be further subdivided into nearly 2,000 categories, including B-cell leukemia, leukocyte adhesion deficiency, and interleukin-1 receptor antagonist deficiency. Questioning the large model can be performed using pre-screened content.

[0048] In a specific embodiment, the Prompt may be:

[0049] i. Target Inquiry: The inventor title is #XXX#, and the inventor content is #XXX#. Please extract the target information (in English). The extracted target information should strictly follow the "Target:......" format. If there are multiple targets, separate them with ";". Do not explain them, and do not include redundant information in your answer. Please answer in English.

[0050] ii. Ask the applicant: "The inventor applicant is #XXX#." Please extract the applicant company information (in English). The extracted applicant company should strictly follow the format of 'Applicant:......'. If there are multiple applicant companies, separate them with ';'. Do not explain or include any unnecessary information in your answer. Please answer in English. [Example 1]: 'the inventor applicant is # REVOLUTION MEDICINES, INC. [US / US];700 Saginaw Drive,Redwood City,California 94063(US).#', the extracted information is 'Applicant:REVOLUTION MEDICINES, INC' [Example 2]: 'the inventor applicant is # MEMORIALSLOAN KETTERING CANCER CENTER[US.US];Office Of Technology Development,1275York Aenue,New York,NY 10065(US).MEMORIAL"HOSPITAL FOR CANCER AND ALLIEDDISESES [US / US]; Office Of Technology Development,1275 York Aenue,New York,NY10065(US).#', the extracted information is 'Applicant:MEMORIAL SLOAN KETTERING CANCER CENTER;MEMORIAL HOSPITAL FOR CANCER AND ALLIED DISESES'.

[0051] In this embodiment, the process of obtaining the fine-tuned large model may specifically include: collecting different types of historical medical patents to obtain a medical patent data set; converting all historical medical patents in the medical patent data set into a unified data format to obtain a converted data set; performing text recognition on each medical patent in the converted data set to obtain historical patent text information, and inputting the historical patent text information into a pre-trained model to perform model fine-tuning to obtain the fine-tuned large model; the pre-trained model is located in the inference server. In this embodiment, historical medical patents from different fields may be collected to obtain a medical patent dataset. Furthermore, considering that the historical medical patents may be downloaded through different channels during the collection process, the multiple historical medical patents in the historical medical patent dataset may exist in different text formats, such as PDF (Portable Document Format) and Word (a text file format). To facilitate subsequent unified model fine-tuning training, all historical medical patents in the historical medical patent dataset may be converted into data formats. For example, historical medical patents in PDF, Word, and other formats may be uniformly converted into jpg (Joint Photographic Experts Group, an image format using a hybrid compression method) image format. Then, text recognition is performed on each medical patent in the converted dataset to obtain historical patent text information including the title, abstract, target, indication, etc. of each patent. This historical patent text information is then input into a pre-trained model for model fine-tuning, thereby obtaining the fine-tuned large model. The pre-trained model may be located in an inference server.

[0052] It is understood that large language models (LLMs) possess the complex capabilities and characteristics of comprehensive analysis and problem-solving at a deeper level. They can demonstrate human-like thinking and intelligence, capture complex semantic relationships within language, and engage in human-level language interaction. In one specific embodiment, the pre-trained model can be an L0 Hairuo large model with 60 bytes of parameters and a maximum supported token count of 4096. The model has capabilities for entity extraction, semantic matching, intelligent summarization, keyword extraction, knowledge classification, and text error correction. It employs a framework of BERT, Bilstm, and CRF, and annotates the training set using BIO (a commonly used sequence annotation method, Begin Inside Outside).

[0053] It's important to note that when fine-tuning a pre-trained model, you can choose an appropriate fine-tuning strategy based on task requirements and available resources. Specifically, consider whether to perform full or partial fine-tuning, as well as the level and scope of fine-tuning. Since low-level features near the input layer are more general, freezing the weights of these layers preserves the general feature extraction capabilities learned during pre-training. High-level features near the output layer, on the other hand, require task-specific fine-tuning to adapt to new task requirements.

[0054] Furthermore, before fine-tuning the model, it's necessary to determine hyperparameters for the fine-tuning process, such as the learning rate, batch size, and number of training rounds. For fine-tuning layers, a smaller learning rate can be used to avoid excessive damage to pre-trained weights; for frozen layers, the learning rate can be zero. The choice of these hyperparameters has a significant impact on the performance and convergence speed of fine-tuning. Considering that weights determine the model's emphasis on different input features, with larger weights corresponding to input features having a greater impact on the output, while smaller weights have a smaller impact, setting weights for the pre-trained model can improve its performance. For partial fine-tuning, the parameters of the higher layers near the output layer are randomly initialized; for full fine-tuning, all model parameters are randomly initialized. This solution allows for the use of partial fine-tuning for model training.

[0055] After determining the fine-tuning strategy, the strategy can be used to fine-tune the pre-trained model based on the pharmaceutical patent dataset. During the fine-tuning process, the model parameters can be gradually adjusted according to the set hyperparameters and optimization algorithm to minimize the loss function; wherein, the loss function is used to calculate the difference between the model's prediction and the true label. Specifically, the binary cross entropy can be maximized to correctly determine the probability of the target in the text. The calculation formula is:

[0056] ;

[0057] Where, represents the true label (value is 0 or 1), represents the predicted value of the model, and n is the total number of samples.

[0058] Furthermore, the Hinge Loss function is used to calculate the probability distribution difference between the target extracted from the input medical patent text and the correct target. The calculation formula is:

[0059] ;

[0060] Where, represents the true label (value is +1 or -1), Represents the predicted value of the model.

[0061] Furthermore, during the process of model fine-tuning, the validation set can be used to regularly evaluate the model and adjust the hyperparameters or fine-tuning strategy based on the evaluation results. In addition, after fine-tuning is completed, the test set can be used to evaluate the final fine-tuned model to obtain the final performance indicators.

[0062] In this embodiment, performing text recognition on each pharmaceutical patent in the converted dataset to obtain historical patent text information may include: using an optical character recognition tool to recognize the text information of each pharmaceutical patent in the converted dataset to obtain the historical patent text information. In other words, an optical character recognition (OCR) tool may be used to recognize the text information of each pharmaceutical patent in the converted dataset.

[0063] Step S13: Determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain the target disease field, and obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information.

[0064] In this embodiment, after parsing the key information in the current pharmaceutical patent using the preset prompt information in the fine-tuned large model to obtain the target key information, the disease field to which the current pharmaceutical patent belongs is determined according to the indications in the above target key information to obtain the target disease field, and then the highest foreign R&D status and the highest domestic R&D status of the current pharmaceutical patent are obtained from relevant websites based on the target disease field and the target in the above target key information; wherein, the highest R&D status refers to the latest R&D stage that the target has reached in the corresponding disease field, including no application, preclinical, non-active clinical application, application in clinical, application for clinical, approved clinical, in clinical, clinical phase I, clinical phase I / II, clinical phase II, clinical phase II / III, clinical phase III, non-active application for marketing, application in marketing, application for marketing, and approved marketing, a total of 16 stages.

[0065] It should be pointed out that in order to improve the accuracy of pharmaceutical patent analysis, after obtaining the target key information, it can be displayed on the manual interaction interface so that users can modify it in the manual interaction interface, such as deletion, addition, modification, etc.

[0066] In a specific embodiment, obtaining the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information may specifically include: utilizing a web crawler to crawl the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information. In other words, crawling the highest domestic and international R&D status from relevant websites based on the results returned by the fine-tuned large model.

[0067] Step S14: Determine the level of the target based on the highest R&D status at home and abroad to obtain the target level, and judge whether the applicant company of the current pharmaceutical patent is a company of key concern to obtain a judgment result.

[0068] In this embodiment, the level of the target in the aforementioned key target information can be determined based on the highest domestic and international R&D status. That is, the patent is rated based on the target's highest domestic R&D status and the highest foreign R&D status. Then, a determination is made as to whether the applicant company for the aforementioned current pharmaceutical patent is a company of key focus, thereby obtaining a corresponding determination result. It is understood that, based on whether the company to which the patent belongs is a major competitor and an object of focus, it is possible to determine whether the R&D of the target is urgent or needs to be abandoned. This determination of a company of key focus can be made using a preset database containing competitors and objects of focus.

[0069] It should be pointed out that this application has pre-classified various types of targets into different levels based on the highest R&D status at home and abroad. For example, multiple types of targets are divided into three levels: A, B, and C. When both the highest R&D status at home and abroad has not reached "clinical development", the target's level is A; when the highest R&D status at home has not reached "clinical development", while the highest R&D status abroad has reached "clinical development" but has not reached "Clinical Phase III", the target's level is B; when the highest R&D status at home has reached "clinical development" but has not reached "Clinical Phase II", while the highest R&D status abroad has reached "Clinical Phase III" but has not reached "application for marketing", the target's level is C. In addition, other situations are not included in the rating and are not worth paying attention to. Among them, level A has the highest R&D value, and level C has the lowest R&D value.

[0070] Step S15: Determine whether the target in the target disease field has research and development value based on the judgment result and the target level.

[0071] In this embodiment, after obtaining the above judgment results (i.e., whether it is a company of key focus) and the above target level, a comprehensive judgment can be made based on these two parts of information as to whether the target in the above target disease field has research and development value.

[0072] In a specific embodiment, determining whether the target in the target disease field has research and development value based on the judgment result and the target level can specifically include: obtaining the appearance position of the target and the indication in the current pharmaceutical patent to obtain first position information, and counting the number of times the target and the indication appear in the current pharmaceutical patent to obtain number information; obtaining the appearance position of the applicant company in the current pharmaceutical patent to obtain second position information, and determining whether the target in the target disease field has research and development value based on the judgment result, the target level, the first position information, the number information and the second position information.

[0073] In this embodiment, the appearance positions of the target and indication in the current pharmaceutical patent can be obtained respectively to obtain the first position information, and then the number of times the target and indication appear in the current pharmaceutical patent can be counted respectively, and then the appearance position of the applicant company in the current pharmaceutical patent can be obtained to obtain the second position information. Finally, based on the judgment result of whether it belongs to a key company, the target level, the first position information, the number information, and the second position information, it can be comprehensively judged whether the target in the disease field to which the current pharmaceutical patent belongs has research and development value.

[0074] Furthermore, after determining whether the target in the target disease field has R&D value based on the judgment result and the target level, it may also include: sending the target key information, the target disease field, the highest R&D status at home and abroad, the target level, the judgment result and the judgment result of whether it has R&D value to the user terminal, so as to perform structured display on the human-computer interaction interface of the user terminal. In this embodiment, in order to facilitate user viewing, the extracted target key information, the disease field of the current pharmaceutical patent, the highest R&D status at home and abroad, the target level, the judgment result and the judgment result of whether it has R&D value can be sent to the human-computer interaction interface of the user terminal for structured display. For example, the extracted key information is marked at the corresponding position of the original text and structured display is performed. In addition, the human-computer interaction interface can also provide modification permissions to facilitate users to manually modify and maintain the displayed information. For example, the user can modify and supplement the information such as targets and indications in the human-computer interaction interface, and can also modify the marked position of the key information.

[0075] It can be seen that the embodiment of the present application first obtains the pharmaceutical patent uploaded by the current user terminal to obtain the current pharmaceutical patent, and then inputs the current pharmaceutical patent into the fine-tuned large model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, so as to use the preset prompt information to parse the key information in the current pharmaceutical patent to obtain the target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, wherein the target key information includes the target, indication and the applicant company of the current pharmaceutical patent. Then, according to the indication in the target key information, the disease field of the current pharmaceutical patent is determined to obtain the target disease field, and according to the target disease field and the target in the target key information, the highest domestic and foreign R&D status of the current pharmaceutical patent is obtained from the relevant website, and then the level of the target is determined based on the highest domestic and foreign R&D status to obtain the target level, and it is judged whether the applicant company of the current pharmaceutical patent is a key focus company to obtain a judgment result, and finally, according to the judgment result and the target level, it is determined whether the target in the target disease field has R&D value. The embodiment of the present application first analyzes the current pharmaceutical patent input by the user through the fine-tuned large model to obtain the target key information including the target, indication and the applicant company of the current pharmaceutical patent, and then determines the disease field of the current pharmaceutical patent, the highest domestic and foreign R&D status, the target level and whether it belongs to a key company based on the target key information. Finally, based on the determined information, it determines whether the current target has R&D value. The embodiment of the present application uses large model technology to deeply understand and analyze pharmaceutical patents, and realizes the extraction, collation and current status analysis of key information in automated pharmaceutical patents. Compared with the traditional manual patent information analysis method, it not only greatly improves the efficiency of patent information analysis, but also saves the manpower and time costs of patent analysis, avoids the waste of human resources, and at the same time, reduces the professional quality requirements for professionals, and greatly shortens the time and confirmation cycle of patent key information statistics, thereby providing timely and accurate guidance for enterprises to determine the R&D direction.

[0076] The present application embodiment discloses a specific method for analyzing pharmaceutical patent information based on a large model, see Figure 2 As shown, the method includes:

[0077] Step S21: Acquire the pharmaceutical patent uploaded by the current user terminal to obtain the current pharmaceutical patent.

[0078] Step S22: convert the current medical patent into a picture to obtain a target picture, and cut the target picture to obtain multiple sub-pictures.

[0079] In this embodiment, in order to improve the efficiency of patent information analysis, after obtaining the pharmaceutical patent uploaded by the user, the above-mentioned current pharmaceutical patent can be converted into an image format, and then the target image can be cut into multiple sub-images for subsequent parallel processing.

[0080] Step S23: Perform text recognition on each of the sub-images to obtain multiple current patent text information, and input the current patent text information into the fine-tuned large model in turn, so as to parse the key information in the current pharmaceutical patent using the preset prompt information to obtain the target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset; the target key information includes the target, indication and the applicant company of the current pharmaceutical patent.

[0081] In this embodiment, after the target image is cut into multiple sub-images, the OCR tool can be used to perform segmented text recognition on each sub-image to obtain multiple current patent text information. The current patent text information is then sequentially input into the large model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, so as to use the preset prompt information (prompt) in different models to ask questions about the above-mentioned current pharmaceutical patent to parse the key information in the current pharmaceutical patent and answer the target key information including the target, indication and the applicant company of the current pharmaceutical patent according to the preset requirements. Among them, the segmented text recognition process supports the recognition of multiple small languages ​​such as Chinese, English, Japanese, Korean, French and German, and can extract information such as patent title and patent number based on text features. In addition, the large model supports the analysis of pharmaceutical patents containing Chinese, English and multiple other languages, and parses out information such as target, indication and patent-owning company in the corresponding text.

[0082] It should be noted that the fixed format refers to the format specified in the prompt, for example:

[0083] i. Question: The inventor title is #XXX#, the inventor content is #XXX#. Please extract the target information (in English). The extracted targets should strictly follow the answer format of 'Target:......'. If there are multiple targets, separate them with ';'. Do not explain them, and do not include any redundant information in your answer. Please answer in English.

[0084] ii.Answer: Target: XXX;XXX.

[0085] Step S24: Determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain the target disease field, and obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information.

[0086] Step S25: Determine the level of the target based on the highest R&D status at home and abroad, obtain the target level, and judge whether the applicant company of the current pharmaceutical patent is a company of key concern, and obtain a judgment result.

[0087] Step S26: Determine whether the target in the target disease field has research and development value based on the judgment result and the target level.

[0088] For more specific processing procedures of the above steps S21, S24 to S26, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0089] It can be seen that the embodiment of the present application first converts the current pharmaceutical patent into a picture to obtain a target picture, and cuts the target picture to obtain multiple sub-pictures, and then performs text recognition on each of the sub-pictures to obtain multiple current patent text information, and inputs the current patent text information into the fine-tuned large model in turn, so as to use the preset prompt information to parse the key information in the current pharmaceutical patent to obtain the target key information, and then determine the disease field of the current pharmaceutical patent according to the indications in the target key information to obtain the target disease field, and obtain the highest domestic and foreign R&D status of the current pharmaceutical patent from relevant websites according to the target disease field and the target in the target key information, and finally determine the level of the target based on the highest domestic and foreign R&D status to obtain the target level, and judge whether the applicant company of the current pharmaceutical patent is a key company to obtain a judgment result, so as to determine whether the target in the target disease field has R&D value according to the judgment result and the target level. The embodiment of this application integrates technologies such as text recognition and language large models, and uses pre-trained large models to deeply understand and analyze pharmaceutical patents, thereby realizing the extraction, collation and current status analysis of key information of pharmaceutical patents. Moreover, this solution supports batch processing of pharmaceutical patents, taking an average of 2 minutes per 100 patents. Compared with the traditional manual analysis method of 3-5 patents per person per week, there is a significant improvement in the speed of information processing. At the same time, it liberates human resources in professional fields, reduces the professional requirements of human resources, greatly shortens the time and confirmation cycle of patent key information statistics, and reduces labor costs and time costs. In addition, this solution also promotes the intelligent and automated development of patent management, and provides guidance on the timely adjustment of R&D directions and the search for scarce R&D directions, laying a solid foundation for the digital transformation and sustainable development of enterprises.

[0090] Correspondingly, the present application also discloses a pharmaceutical patent information analysis device based on a large model, see Figure 3 As shown, the device includes:

[0091] The patent acquisition module 11 is used to acquire the pharmaceutical patent uploaded by the current user terminal and obtain the current pharmaceutical patent;

[0092] The parsing module 12 is configured to input the current pharmaceutical patent into the fine-tuned large model, and parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset; the target key information includes the target, indication, and the applicant company of the current pharmaceutical patent;

[0093] A field determination module 13 is configured to determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain a target disease field;

[0094] A status acquisition module 14 is configured to acquire the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information;

[0095] A level determination module 15 is used to determine the level of the target based on the highest domestic and international R&D status to obtain the target level;

[0096] A judgment module 16 is used to judge whether the company applying for the current pharmaceutical patent is a company of key concern, and obtain a judgment result;

[0097] The R&D value determination module 17 is used to determine whether the target in the target disease field has R&D value based on the determination result and the target level.

[0098] Among them, the specific work processes of the above modules can refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0099] It can be seen that in the embodiment of the present application, the pharmaceutical patent uploaded by the current user terminal is first obtained to obtain the current pharmaceutical patent, and then the current pharmaceutical patent is input into the fine-tuned large model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, so as to use the preset prompt information to parse the key information in the current pharmaceutical patent to obtain the target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset, wherein the target key information includes the target, indication and the applicant company of the current pharmaceutical patent. Then, according to the indication in the target key information, the disease field of the current pharmaceutical patent is determined to obtain the target disease field, and according to the target disease field and the target in the target key information, the highest domestic and foreign R&D status of the current pharmaceutical patent is obtained from the relevant website, and then based on the highest domestic and foreign R&D status, the level of the target is determined to obtain the target level, and it is judged whether the applicant company of the current pharmaceutical patent is a key company to obtain a judgment result. Finally, according to the judgment result and the target level, it is determined whether the target in the target disease field has R&D value. The embodiment of the present application first analyzes the current pharmaceutical patent input by the user through the fine-tuned large model to obtain the target key information including the target, indication and the applicant company of the current pharmaceutical patent, and then determines the disease field of the current pharmaceutical patent, the highest domestic and foreign R&D status, the target level and whether it belongs to a key company based on the target key information. Finally, based on the determined information, it determines whether the current target has R&D value. The embodiment of the present application uses large model technology to deeply understand and analyze pharmaceutical patents, and realizes the extraction, collation and current status analysis of key information in automated pharmaceutical patents. Compared with the traditional manual patent information analysis method, it not only greatly improves the efficiency of patent information analysis, but also saves the manpower and time costs of patent analysis, avoids the waste of human resources, and at the same time, reduces the professional quality requirements for professionals, and greatly shortens the time and confirmation cycle of patent key information statistics, thereby providing timely and accurate guidance for enterprises to determine the R&D direction.

[0100] In some specific embodiments, the state acquisition module 14 may specifically include:

[0101] The crawling unit is used to use a web crawler to crawl the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information.

[0102] In some specific embodiments, the large model-based pharmaceutical patent information analysis device may further include:

[0103] Patent collection unit, used to collect different types of historical pharmaceutical patents to obtain pharmaceutical patent datasets;

[0104] a format conversion unit, configured to convert all historical pharmaceutical patents in the pharmaceutical patent dataset into a unified data format to obtain a converted dataset;

[0105] A first text recognition unit is used to perform text recognition on each pharmaceutical patent in the converted data set to obtain historical patent text information;

[0106] An information input unit is used to input the historical patent text information into a pre-trained model to perform model fine-tuning to obtain the fine-tuned large model; the pre-trained model is located in the inference server.

[0107] In some specific embodiments, the first text recognition unit may specifically include:

[0108] The second text recognition unit is used to use an optical character recognition tool to recognize the text information of each pharmaceutical patent in the converted data set to obtain historical patent text information.

[0109] In some specific embodiments, the parsing module 12 may specifically include:

[0110] An image conversion unit, configured to convert the current pharmaceutical patent into an image to obtain a target image;

[0111] A picture cutting unit, used for cutting the target picture into multiple sub-pictures;

[0112] A third text recognition unit is used to perform text recognition on each of the sub-images to obtain multiple current patent text information;

[0113] The parsing unit is used to input the current patent text information into the fine-tuned large model in sequence, so as to parse the key information in the current medical patent using preset prompt information to obtain target key information.

[0114] In some specific embodiments, the R&D value determination module 17 may specifically include:

[0115] A first position information acquisition unit is configured to acquire the locations where the target and the indication appear in the current pharmaceutical patent to obtain first position information;

[0116] A frequency statistics unit, used to count the number of times the target and the indication appear in the current pharmaceutical patent to obtain frequency information;

[0117] A second position information obtaining unit is configured to obtain a position where the applicant company appears in the current pharmaceutical patent to obtain second position information;

[0118] A research and development value determination unit is used to determine whether the target in the target disease field has research and development value based on the judgment result, the target level, the first position information, the number information and the second position information.

[0119] In some specific embodiments, the large model-based pharmaceutical patent information analysis device may further include:

[0120] An information display unit is used to send the target key information, the target disease field, the highest domestic and foreign R&D status, the target level, the judgment result and the judgment result of whether it has R&D value to the user terminal for structured display on the human-computer interaction interface of the user terminal.

[0121] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.

[0122] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the large model-based pharmaceutical patent information analysis method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0123] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0124] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0125] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, NetWare, Unix, Linux, etc. In addition to including computer programs capable of implementing the large model-based pharmaceutical patent information analysis method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0126] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned large-model-based pharmaceutical patent information analysis method. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be further described here.

[0127] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the large model-based pharmaceutical patent information analysis method disclosed above.

[0128] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0129] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0130] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0131] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0132] The above is a detailed introduction to the large-scale model-based pharmaceutical patent information analysis method, device, equipment and medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present application.

Claims

1. A method for analyzing pharmaceutical patent information based on a large model, characterized in that: include: Get the pharmaceutical patent uploaded by the current user and get the current pharmaceutical patent; Input the current pharmaceutical patent into the fine-tuned large model, and use preset prompt information to parse the key information in the current pharmaceutical patent to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning the pre-trained model using the pharmaceutical patent dataset; the target key information includes the target, indication, and the applicant company of the current pharmaceutical patent; Determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain the target disease field, and obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information; Determine the level of the target based on the highest domestic and international R&D status, obtain the target level, and judge whether the applicant company of the current pharmaceutical patent is a company of key concern, and obtain a judgment result; Determining whether the target in the target disease field has research and development value based on the judgment result and the target level; The inputting the current pharmaceutical patent into the fine-tuned large model to parse key information in the current pharmaceutical patent using preset prompt information to obtain target key information includes: converting the current pharmaceutical patent into an image to obtain a target image, and cutting the target image to obtain multiple sub-images; performing text recognition on each of the sub-images to obtain multiple current patent text information, and sequentially inputting the current patent text information into the fine-tuned large model to parse key information in the current pharmaceutical patent using preset prompt information to obtain target key information; The determining whether the target in the target disease field has research and development value based on the judgment result and the target level includes: obtaining the appearance position of the target and the indication in the current pharmaceutical patent to obtain first position information, and counting the number of times the target and the indication appear in the current pharmaceutical patent to obtain number information; obtaining the appearance position of the applicant company in the current pharmaceutical patent to obtain second position information, and determining whether the target in the target disease field has research and development value based on the judgment result, the target level, the first position information, the number information and the second position information.

2. The method for analyzing pharmaceutical patent information based on a large model according to claim 1, characterized in that: The highest domestic and international R&D status of the current pharmaceutical patent is obtained from relevant websites based on the target disease field and the target in the target key information, including: Utilize a web crawler and crawl the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information.

3. The method for analyzing pharmaceutical patent information based on a large model according to claim 1, characterized in that: Also includes: Collect different types of historical pharmaceutical patents to obtain a pharmaceutical patent dataset; Convert all historical pharmaceutical patents in the pharmaceutical patent dataset into a unified data format to obtain a converted dataset; Text recognition is performed on each pharmaceutical patent in the converted data set to obtain historical patent text information, and the historical patent text information is input into a pre-trained model for model fine-tuning to obtain the fine-tuned large model; the pre-trained model is located in the inference server.

4. The method for analyzing pharmaceutical patent information based on a large model according to claim 3, characterized in that: The text recognition is performed on each pharmaceutical patent in the converted data set to obtain historical patent text information, including: Optical character recognition tools are used to identify the text information of each pharmaceutical patent in the converted data set to obtain historical patent text information.

5. The method for analyzing pharmaceutical patent information based on a large model according to any one of claims 1 to 4, characterized in that: Also includes: The target key information, the target disease field, the highest domestic and international R&D status, the target level, the judgment result and the judgment result of whether it has R&D value are sent to the user terminal for structured display on the human-computer interaction interface of the user terminal.

6. A pharmaceutical patent information analysis device based on a large model, characterized in that: include: The patent acquisition module is used to obtain the pharmaceutical patent uploaded by the current user and obtain the current pharmaceutical patent; A parsing module is configured to input the current pharmaceutical patent into a fine-tuned large model, and to parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; the fine-tuned large model is a model obtained by fine-tuning a pre-trained model using a pharmaceutical patent dataset; the target key information includes the target, indication, and the applicant company of the current pharmaceutical patent; A field determination module is used to determine the disease field of the current pharmaceutical patent based on the indication in the target key information to obtain a target disease field; A status acquisition module is used to obtain the highest domestic and international R&D status of the current pharmaceutical patent from relevant websites based on the target disease field and the target in the target key information; A level determination module, configured to determine the level of the target based on the highest domestic and international R&D status, thereby obtaining the target level; A judgment module, configured to judge whether the applicant company of the current pharmaceutical patent is a company of key concern, and obtain a judgment result; A research and development value determination module, configured to determine whether the target in the target disease field has research and development value based on the determination result and the target level; The parsing module is specifically configured to convert the current pharmaceutical patent into an image to obtain a target image, and to segment the target image into multiple sub-images; perform text recognition on each of the sub-images to obtain multiple pieces of current patent text information, and sequentially input the current patent text information into the fine-tuned large model to parse the key information in the current pharmaceutical patent using preset prompt information to obtain target key information; The R&D value judgment module is specifically used to obtain the appearance positions of the target and the indication in the current pharmaceutical patent to obtain first position information, and count the number of times the target and the indication appear in the current pharmaceutical patent to obtain number information; obtain the appearance position of the applicant company in the current pharmaceutical patent to obtain second position information, and determine whether the target in the target disease field has R&D value based on the judgment result, the target level, the first position information, the number information and the second position information.

7. An electronic device, characterized in that: It comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the large model-based pharmaceutical patent information analysis method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that Used to store computer programs; wherein, when the computer program is executed by a processor, it implements the large model-based pharmaceutical patent information analysis method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Drug target interaction prediction method based on multilayer network and graph coding

    CN113571125A

  • Target spot information mining and retrieval method and device, electronic equipment and storage medium

    CN114255877A