Pathological gene detection report generation system based on natural language processing
By designing a pathological gene detection report generation system based on natural language processing, the problems of low generation efficiency and insufficient analysis depth in the prior art are solved, and high-quality reports are generated quickly and accurately, and the detection work efficiency is improved.
Patent Information
- Application Number
- CN202510068095.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The generation efficiency of existing pathological genetic test reports is inefficient, and it takes two days for experts and assistants to complete, and there are some shortcomings in the depth of the analysis of gene test results and the comprehensiveness of drug-related analysis.
Design a pathological gene detection report generation system based on natural language processing, including the detection result import module, the rule matching algorithm module, the drug toxic side effect analysis module, the gene summary module of unknown significance, the hot gene interpretation module, the report automatic generation module, the report review and adjustment module, the electronic signature and PDF generation module and the report sending module.
Through this system, high-quality pathological genetic testing reports can be generated quickly and accurately, greatly improving the efficiency of report generation, reducing labor costs and human errors, and providing strong support for pathological genetic testing.
Smart Images

Figure CN119993368A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of report generation, and specifically to a pathological gene detection report generation system based on natural language processing. Background Art
[0002] The existing pathological gene test report is obtained from the results of instrument testing by slicing specimens and testing with instruments. The status of various susceptible genes, such as mutant gene information, MMR\MSI\TMB\DDR\Fusion, etc., is judged by experts based on their own experience to summarize the gene mutation and drug tolerance, and then give the risk of gene pathological changes and drug tolerance reports. The preparation of pathological gene test reports is very time-consuming. Generally, one test requires experts and assistants to work for 2 days to complete.
[0003] Chinese patent number CN202411087320.5 discloses a method and system for automated interpretation and management of tumor gene detection reports, but the invention is lacking in the depth of analysis of gene detection results and the comprehensiveness of drug-related analysis.
[0004] In summary, a new technical solution for generating pathological gene test reports based on natural language processing is urgently needed to solve the above technical problems. Summary of the invention
[0005] The purpose of this application is to provide a pathological gene detection report generation system based on natural language processing to solve the technical problems raised in the above-mentioned background technology.
[0006] To achieve the above-mentioned purpose, the present application discloses the following technical solutions: a pathological gene detection report generation system based on natural language processing, the system comprising a detection result import module, a rule matching algorithm module, a drug toxicity and side effect analysis module, an unknown gene summary module, a hotspot gene interpretation module, a report automatic generation module, a report review and adjustment module, an electronic signature and PDF generation module and a report sending module, which are sequentially connected in communication;
[0007] The test result import module is configured to receive the result file of instrument test uploaded by the user;
[0008] The rule matching algorithm module is configured to generate the test results of drug sensitivity / resistance-related genes based on the system's built-in pathology, gene, and medication rules, and the test results at least include the specific results, specific values, clinical guideline recommendations, cancer treatment drug recommendations, and other potential beneficial drug information of single-target targeted drugs, multi-target targeted drugs, immunotherapy drugs, and chemotherapy drug-related genes that are positive or negative;
[0009] The drug toxicity and side effect analysis module is configured to: generate a drug toxicity and side effect result table based on positive genes and drug rules;
[0010] The module for summarizing genes of unknown significance is configured to: summarize all positive genes that are not related to drug performance and generate a list of genes of unknown significance;
[0011] The hotspot gene interpretation module is configured to: generate detailed detection instructions for each gene based on the detection results of each gene based on the pathological type and hotspot gene configuration rules;
[0012] The automatic report generation module is configured to: automatically generate a test report based on the template technology and the algorithm generation results of the above-mentioned stages;
[0013] The report review and adjustment module is configured to allow experts to review and adjust the automatically generated test report and confirm the final report;
[0014] The electronic signature and PDF generation module is configured as follows: the system stamps the expert with an electronic signature and generates a final PDF report;
[0015] The report sending module is configured to send the report to the patient or agent via SMS or email.
[0016] Preferably, the basis for determining the gene detection result in the rule matching algorithm module at least includes:
[0017] In the mutation / SNP detection type, negative is supplemented when there is no data, and positive is determined based on the site situation when there is data. Specifically, the site is a mutation and <5%, wild-type homozygote, [5%-95%) mutant heterozygote, ≥95% mutant homozygote correspond to different positive + judgment conditions. If there is no header, the gene is hidden and not processed;
[0018] In the MMR test type, pMMR is negative, and its judgment condition is that no pathogenic mutation is detected, there is a table header but no data row, or there is data in the Clinvar column but no Pathogenic appears. dMMR is positive +, and its judgment condition is that there is data in the Clinvar column and Pathogenic appears;
[0019] In the MSI test type, MSS is negative, hidden genes are not processed when there is no data, and MSI-L and MSI-H are positive +;
[0020] In the TMB test type, 0 is negative, hidden genes are not processed when there is no data, and (0-6), [6-10), [10-19], and >19 correspond to low, medium, and high positive + judgment conditions, respectively;
[0021] In the DDR test type, no data means negative, there is a table header without data row or more sensitive mutation, general sensitive mutation is determined as positive + based on the specific gene mutation results, and hidden genes are not processed when there is no data;
[0022] In the tumor susceptibility test type, no data means negative, and data means positive + and benign mutation;
[0023] In the fusion detection type, if there is no data, it is negative, and if there is data, it is positive + and it is a gene fusion;
[0024] In the amplification test type, no data is used to compensate for negative, data is available and <2 is positive + and copy number is lost, >2 is positive + and copy number is increased, and normal copy number is negative;
[0025] In the expression detection type, <25% is negative and low expression, [25%-75%) is positive+ and medium expression, and ≥75% is positive+ and high expression;
[0026] In the immunohistochemical detection type, no data is negative, [1-50%) is positive + and expression, and ≥50% is positive + and high expression.
[0027] Preferably, the drug toxicity and side effect analysis module performs analysis based on a preset drug toxicity and side effect knowledge base, and the drug toxicity and side effect knowledge base at least includes drug name, generic name, preparation and specifications, indications, medication method, common adverse reactions, serious adverse reactions, contraindications and precautions for combined medication.
[0028] Preferably, the hotspot gene interpretation module determines whether to display the hotspot interpretation gene based on whether a preset hotspot gene knowledge base is enabled. The hotspot gene knowledge base stores detailed interpretations of hotspot gene detection results. When the hotspot interpretation gene is displayed, the corresponding interpretation content is obtained from the hotspot gene knowledge base. The interpretation content at least includes a gene introduction, guideline-related information and detection supplementary instructions, and records the reference source.
[0029] Preferably, the rules for the gene of unknown significance summary module to generate the gene of unknown significance list include at least:
[0030] The database searches for sites. For positive + sites, based on the drug efficacy analysis, whether the medication is configured, and the selection of hotspot genes, determine whether to display them in the list of genes of unknown significance. The drug analysis display location includes tables for single-target / multi-target / immune / chemotherapy / toxicity and side effect drug analysis.
[0031] Preferably, when generating reference information for the application of targeted / immunotherapy drugs, the automatic report generation module uses preset drug efficacy evaluation rules for judgment based on specific drug information entered by the system for different cancer types and different susceptibility gene types. The drug efficacy evaluation rules are specifically as follows:
[0032] When multiple genes or sites correspond to exactly the same drugs, and X of the N genes have weaker therapeutic effects and Y have better therapeutic effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the therapeutic effect is judged to be good; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the therapeutic effect is judged to be weak; results that do not meet the above conditions are judged to be fair therapeutic effects;
[0033] For the toxicity and side effect assessment, if X of the N genes have severe toxicity and side effects and Y have minor toxicity and side effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the toxicity and side effect are judged to be severe; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the toxicity and side effect are judged to be minor; the results that do not meet the above conditions are judged to be moderate toxicity and side effects.
[0034] Preferably, the system further comprises a pathology management module, which is used to manage pathology information, wherein the pathology information at least comprises newly added pathologies, recorded pathology names, gene numbers, guideline versions and ClinicalTrail update time.
[0035] Preferably, the system also includes a gene management module, which is used to manage gene information, which includes at least gene ID, mutation type, and label, and is used to export / import gene information and medication information according to pathology.
[0036] Preferably, the system further comprises a drug management module, wherein the drug management module is used to manage drug information, wherein the drug information at least comprises drug name, application display, and sorting position.
[0037] Preferably, the automatic report generation module automatically generates a test report based on template technology and the algorithm generation results of the above-mentioned stages in combination with natural language technology.
[0038] Beneficial effects: The pathological gene detection report generation system based on natural language processing of the present application starts from receiving the instrument detection result file, generates comprehensive drug sensitivity / resistance gene detection results through precise rule matching, deeply analyzes drug toxicity and side effects, summarizes genes of unknown significance, and interprets hot genes. Finally, with the help of automatic report generation, review and adjustment, electronic signature and sending processes, it can quickly and accurately generate high-quality pathological gene detection reports, greatly improve the report generation efficiency, reduce labor costs and human errors, and provide strong support for pathological gene detection work. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 A structural block diagram of a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0041] Figure 2 A schematic diagram of a result file of a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0042] Figure 3 A gene detection index configuration table for a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of drug toxicity and side effect knowledge base matching for a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0044] Figure 5 A schematic diagram of a hotspot gene knowledge base of a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0045] Figure 6 A schematic diagram of the rules for the list of genes of unknown significance in the pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0046] Figure 7 A schematic diagram of a medical record management module of a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0047] Figure 8 A schematic diagram of a gene management module of a pathological gene detection report generation system based on natural language processing provided in an embodiment of the present application;
[0048] Fig. 9 Schematic diagram of the drug management module of the pathological gene detection report generation system based on natural language processing provided in the embodiment of the present application Figure 1 ;
[0049] Fig.10 Schematic diagram of the drug management module of the pathological gene detection report generation system based on natural language processing provided in the embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0050] The technical solutions in the embodiments of the present application are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.
[0051] In this article, the term "comprising" is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprising..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0052] In order to solve the problem that pathological gene test reports require experts to manually judge gene mutations and drug tolerance based on instrument test results and personal experience, which is inefficient and prone to errors, this system solution automatically generates corresponding gene test results, drug efficacy evaluation, and hot gene interpretation information by specifying gene result matching rules and algorithms, and automatically generates corresponding test reports according to templates through document template technology.
[0053] The first aspect of this embodiment discloses Figure 1 A pathological gene detection report generation system based on natural language processing is shown, which includes a detection result import module, a rule matching algorithm module, a drug toxicity and side effect analysis module, an unknown gene summary module, a hotspot gene interpretation module, a report automatic generation module, a report review and adjustment module, an electronic signature and PDF generation module and a report sending module that are sequentially connected in communication;
[0054] The test result import module is configured to receive the result files of instrument tests uploaded by users; in a simple example, Figure 2 As shown, the result file is an Excel table.
[0055] The rule matching algorithm module is configured as follows: based on the system's built-in pathology, gene, and medication rules, it generates the test results of drug sensitivity / resistance-related genes, which at least include the specific results, specific values, clinical guideline recommendations, cancer treatment drug recommendations, and other potential beneficial drug information of single-target targeted drugs, multi-target targeted drugs, immunotherapy drugs, and chemotherapy drug-related genes;
[0056] The drug toxicity and side effect analysis module is configured as follows: based on positive genes and drug rules, a drug toxicity and side effect result table is generated;
[0057] The module for summarizing genes of unknown significance is configured as follows: summarizing all positive genes that are not related to drug performance and generating a list of genes of unknown significance;
[0058] The hotspot gene interpretation module is configured as follows: based on the pathological type and hotspot gene configuration rules, a detailed detection description of each gene is generated for the detection results of each gene;
[0059] The report automatic generation module is configured to: automatically generate a test report based on the template technology and the algorithm generation results of the above-mentioned stages;
[0060] The report review and adjustment module is configured to allow experts to review and adjust the automatically generated test report and confirm the final report;
[0061] The electronic signature and PDF generation module is configured as follows: the system stamps the expert’s electronic signature and generates a final PDF report;
[0062] The report sending module is configured to send reports to patients or agents via SMS or email.
[0063] Based on the above, this embodiment realizes that starting from receiving the instrument test result file, comprehensive drug sensitivity / resistance gene test results are generated through precise rule matching, in-depth analysis of drug toxicity and side effects, summary of genes of unknown significance, interpretation of hot genes, and finally with the help of automatic report generation, review and adjustment, electronic signature and sending processes, high-quality pathological gene test reports are generated quickly and accurately, which greatly improves the report generation efficiency, reduces labor costs and human errors, and provides strong support for pathological gene testing.
[0064] Specifically, Figure 3 As shown, the basis for determining the gene detection result in the rule matching algorithm module at least includes:
[0065] In the mutation / SNP detection type, negative is supplemented when there is no data, and positive is determined based on the site situation when there is data. Specifically, the site is a mutation and <5%, wild-type homozygote, [5%-95%) mutant heterozygote, ≥95% mutant homozygote correspond to different positive + judgment conditions. If there is no header, the gene is hidden and not processed;
[0066] In the MMR test type, pMMR is negative, and its judgment condition is that no pathogenic mutation is detected, there is a table header but no data row, or there is data in the Clinvar column but no Pathogenic appears. dMMR is positive +, and its judgment condition is that there is data in the Clinvar column and Pathogenic appears;
[0067] In the MSI test type, MSS is negative, hidden genes are not processed when there is no data, and MSI-L and MSI-H are positive +;
[0068] In the TMB test type, 0 is negative, hidden genes are not processed when there is no data, and (0-6), [6-10), [10-19], and >19 correspond to low, medium, and high positive + judgment conditions, respectively;
[0069] In the DDR test type, no data means negative, there is a table header without data row or more sensitive mutation, general sensitive mutation is determined as positive + based on the specific gene mutation results, and hidden genes are not processed when there is no data;
[0070] In the tumor susceptibility test type, no data means negative, and data means positive + and benign mutation;
[0071] In the fusion detection type, if there is no data, it is negative, and if there is data, it is positive + and it is a gene fusion;
[0072] In the amplification test type, no data is used to compensate for negative, data is available and <2 is positive + and copy number is lost, >2 is positive + and copy number is increased, and normal copy number is negative;
[0073] In the expression detection type, <25% is negative and low expression, [25%-75%) is positive+ and medium expression, and ≥75% is positive+ and high expression;
[0074] In the immunohistochemical detection type, no data is negative, [1-50%) is positive + and expression, and ≥50% is positive + and high expression.
[0075] Based on the above, this embodiment uses the detailed and scientific basis for determining the results of gene testing in the rule matching algorithm module to achieve accurate determination of different types of gene testing results. Whether it is mutation / SNP, MMR, MSI or other types of gene testing, it can accurately distinguish between positive and negative results based on specific conditions, providing a reliable data basis for subsequent drug recommendations, toxicity and side effect analysis, etc., improving the accuracy and scientificity of the entire pathological gene testing report, making the report more clinically valuable, and thus assisting medical personnel in making more reasonable diagnosis and treatment decisions.
[0076] Specifically, Figure 4 As shown, the drug toxicity and side effect analysis module performs analysis based on a preset drug toxicity and side effect knowledge base, which at least includes the drug name, generic name, preparation and specifications, indications, method of use, common adverse reactions, serious adverse reactions, contraindications and precautions for combined use.
[0077] Based on the above, this embodiment uses the comprehensive drug toxicity and side effect knowledge base preset in the drug toxicity and side effect analysis module to achieve a systematic analysis of drug toxicity and side effects. With rich drug information, including name, specification, indication and adverse reactions, etc., after detecting positive genes, it can quickly match and generate a detailed toxicity and side effect result table, so that medical personnel and patients can fully understand the potential risks of drugs, ensure drug safety, and enhance the practicality and reliability of pathological gene detection in clinical drug guidance.
[0078] Specifically, Figure 5 As shown, the hotspot gene interpretation module determines whether to display the hotspot interpretation gene based on whether the preset hotspot gene knowledge base is enabled. The hotspot gene knowledge base stores detailed interpretations of hotspot gene detection results. When the hotspot interpretation gene is displayed, the corresponding interpretation content is obtained from the hotspot gene knowledge base. The interpretation content at least includes a gene introduction, guideline-related information and detection supplementary instructions, and records the reference source.
[0079] Based on the above, this embodiment uses the preset hotspot gene knowledge base and the activation judgment mechanism in the hotspot gene interpretation module to achieve targeted interpretation of hotspot genes. By efficiently judging whether to display hotspot genes and accurately obtaining their interpretation content, it provides medical personnel with in-depth information on key genes, helps to deeply understand the relationship between genes and diseases, enhances the guiding role of pathological gene detection in disease research and diagnosis and treatment, and promotes the development of precision medicine.
[0080] Specifically, Figure 6 As shown, the rules for the gene of unknown significance summary module to generate a list of genes of unknown significance include at least:
[0081] The database searches for sites. For positive + sites, based on the drug efficacy analysis, whether the medication is configured, and the selection of hotspot genes, determine whether to display them in the list of genes of unknown significance. The drug analysis display location includes tables for single-target / multi-target / immune / chemotherapy / toxicity and side effect drug analysis.
[0082] Based on the above, this embodiment uses the specific rules of the gene of unknown significance summary module to achieve effective management of genes of unknown significance. Based on database search and multi-factor comprehensive judgment, it is accurately determined whether to display genes in the list of unknown significance, avoiding information confusion, allowing medical personnel to focus on key gene information, assisting them in quickly screening valuable information in complex gene test results, improving diagnosis and treatment efficiency, and optimizing the information processing process of pathological gene detection.
[0083] Specifically, when the report automatic generation module generates reference information for the application of targeted / immunotherapy drugs, it uses the preset drug efficacy evaluation rules to make judgments based on the specific drug information entered by the system for different cancer types and different susceptibility gene types. The specific drug efficacy evaluation rules are:
[0084] When multiple genes or sites correspond to exactly the same drugs, and X of the N genes have weaker therapeutic effects and Y have better therapeutic effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the therapeutic effect is judged to be good; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the therapeutic effect is judged to be weak; results that do not meet the above conditions are judged to be fair therapeutic effects;
[0085] For the toxicity and side effect assessment, if X of the N genes have severe toxicity and side effects and Y have minor toxicity and side effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the toxicity and side effect are judged to be severe; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the toxicity and side effect are judged to be minor; the results that do not meet the above conditions are judged to be moderate toxicity and side effects.
[0086] Based on the above, this embodiment uses the preset efficacy evaluation rules of the report automatic generation module to achieve the reasonable generation of reference information for the application of targeted / immunotherapy drugs. Based on the drug information of different cancer and gene types and scientific evaluation logic, it accurately judges the efficacy and toxicity of drugs, provides accurate reference for clinical medication, helps medical personnel develop personalized treatment plans, improves the application efficiency of pathological gene detection in the field of tumor treatment, and promotes the implementation of precision medication.
[0087] Specifically, Figure 7 As shown, the system also includes a pathology management module, which is used to manage pathology information. The pathology information at least includes newly added pathologies, recorded pathology names, gene numbers, guideline versions, and ClinicalTrail update time.
[0088] Based on the above, this embodiment uses the system management function of the pathology management module for pathology information to achieve orderly integration and convenient query of pathology data. By recording key information such as newly added pathology, name, number of genes, etc., it is convenient for medical personnel to quickly obtain pathology-related data, provide basic support for gene testing and report generation, improve work efficiency, ensure that pathology gene testing work is carried out under an accurate pathology background, and enhance the practicality and professionalism of the system.
[0089] Specifically, Figure 8 As shown, the system also includes a gene management module, which is used to manage gene information. The gene information includes at least gene ID, mutation type, and label, and is used to export / import gene information and medication information according to pathology.
[0090] Based on the above, this embodiment uses the gene management module's effective management capabilities for gene information to achieve efficient processing and secure transmission of gene data. With the management of information such as gene ID and variant type and the import and export functions according to pathology, the accurate flow of gene data within the system is guaranteed, which facilitates medical personnel to flexibly use gene data according to pathological needs, improves the pertinence and accuracy of pathological gene detection work, and promotes the application of gene detection technology in clinical practice.
[0091] Specifically, Fig. 9 and Fig.10 As shown, the system also includes a drug management module, which is used to manage drug information. The drug information at least includes drug name, application display, and sorting position.
[0092] Based on the above, this embodiment uses the comprehensive management method of drug information by the drug management module to achieve standardized management and rapid retrieval of drug data. By recording information such as drug names and display status, it is convenient for medical personnel to find and select suitable drugs, ensure the accurate application of drug information in the process of generating pathological gene detection reports, improve the efficiency of the system in drug recommendation and management, and assist in the formulation of clinical drug use decisions.
[0093] Specifically, the report automatic generation module automatically generates a test report based on template technology and the algorithm generation results of the above-mentioned stages, combined with natural language technology.
[0094] Based on the above, this embodiment uses the advantages of the automatic report generation module combined with natural language technology to achieve the natural and fluent expression of the test report. When generating the report, the complex genetic test and analysis results are presented in clear and accurate natural language, which makes it easier for patients and non-professional medical personnel to understand the report content, promotes doctor-patient communication, and improves the quality and accessibility of pathological gene testing services.
[0095] In summary, the pathological gene detection report generation system based on natural language processing in this embodiment starts from receiving the instrument detection result file, generates comprehensive drug sensitivity / resistance gene detection results through precise rule matching, deeply analyzes drug toxicity and side effects, summarizes genes of unknown significance, and interprets hot genes. Finally, with the help of automatic report generation, review and adjustment, electronic signature and sending processes, it can quickly and accurately generate high-quality pathological gene detection reports, greatly improve the report generation efficiency, reduce labor costs and human errors, and provide strong support for pathological gene detection work.
[0096] In the embodiments provided in the present application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code or any appropriate combination thereof. For hardware implementation, the processor can be implemented in one or more of the following units: application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, other electronic units designed to implement the functions described herein or their combination. For software implementation, part or all of the flow of the embodiment can be completed by instructing the relevant hardware through a computer program. When implemented, the above program can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication medium includes any medium that is convenient for transmitting a computer program from one place to another. The storage medium can be any available medium that a computer can access. The computer-readable storage medium can include but is not limited to RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer.
[0097] Finally, it should be noted that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A pathological gene detection report generation system based on natural language processing, characterized in that: The system includes a test result import module, a rule matching algorithm module, a drug toxicity and side effect analysis module, a gene summary module of unknown significance, a hotspot gene interpretation module, a report automatic generation module, a report review and adjustment module, an electronic signature and PDF generation module and a report sending module, which are connected in sequence by communication; The test result import module is configured to receive the result file of instrument test uploaded by the user; The rule matching algorithm module is configured to generate the test results of drug sensitivity / resistance-related genes based on the system's built-in pathology, gene, and medication rules, and the test results at least include the specific results, specific values, clinical guideline recommendations, cancer treatment drug recommendations, and other potential beneficial drug information of single-target targeted drugs, multi-target targeted drugs, immunotherapy drugs, and chemotherapy drug-related genes that are positive or negative; The drug toxicity and side effect analysis module is configured to: generate a drug toxicity and side effect result table based on positive genes and drug rules; The module for summarizing genes of unknown significance is configured to: summarize all positive genes that are not related to drug performance and generate a list of genes of unknown significance; The hotspot gene interpretation module is configured to: generate detailed detection instructions for each gene based on the detection results of each gene based on the pathological type and hotspot gene configuration rules; The automatic report generation module is configured to: automatically generate a test report based on the template technology and the algorithm generation results of the above-mentioned stages; The report review and adjustment module is configured to allow experts to review and adjust the automatically generated test report and confirm the final report; The electronic signature and PDF generation module is configured as follows: the system stamps the expert with an electronic signature and generates a final PDF report; The report sending module is configured to send the report to the patient or agent via SMS or email.
2. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The basis for determining the gene detection result in the rule matching algorithm module at least includes: In the mutation / SNP detection type, negative is supplemented when there is no data, and positive is determined based on the site situation when there is data. Specifically, the site is a mutation and <5%, wild-type homozygote, [5%-95%) mutant heterozygote, ≥95% mutant homozygote correspond to different positive + judgment conditions. If there is no header, the gene is hidden and not processed; In the MMR test type, pMMR is negative, and its judgment condition is that no pathogenic mutation is detected, there is a table header but no data row, or there is data in the Clinvar column but no Pathogenic appears. dMMR is positive +, and its judgment condition is that there is data in the Clinvar column and Pathogenic appears; In the MSI test type, MSS is negative, hidden genes are not processed when there is no data, and MSI-L and MSI-H are positive +; In the TMB test type, 0 is negative, hidden genes are not processed when there is no data, and (0-6), [6-10), [10-19], and >19 correspond to low, medium, and high positive + judgment conditions, respectively; In the DDR test type, no data means negative, there is a table header without data row or more sensitive mutation, general sensitive mutation is determined as positive + based on the specific gene mutation results, and hidden genes are not processed when there is no data; In the tumor susceptibility test type, no data means negative, and data means positive + and benign mutation; In the fusion detection type, if there is no data, it is negative, and if there is data, it is positive + and it is a gene fusion; In the amplification test type, no data is used to compensate for negative, data is available and <2 is positive + and copy number is lost, >2 is positive + and copy number is increased, and normal copy number is negative; In the expression detection type, <25% is negative and low expression, [25%-75%) is positive+ and medium expression, and ≥75% is positive+ and high expression; In the immunohistochemical detection type, no data is negative, [1-50%) is positive + and expression, and ≥50% is positive + and high expression.
3. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The drug toxicity and side effect analysis module performs analysis based on a preset drug toxicity and side effect knowledge base, and the drug toxicity and side effect knowledge base at least includes drug name, generic name, preparation and specification, indication, method of administration, common adverse reactions, serious adverse reactions, contraindications and precautions for combined medication.
4. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The hotspot gene interpretation module determines whether to display the hotspot interpretation gene based on whether the preset hotspot gene knowledge base is enabled. The hotspot gene knowledge base stores detailed interpretations of hotspot gene detection results. When displaying the hotspot interpretation gene, the corresponding interpretation content is obtained from the hotspot gene knowledge base. The interpretation content at least includes a gene introduction, guideline-related information and detection supplementary instructions, and records the reference source.
5. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The rules for the gene of unknown significance summary module to generate the gene of unknown significance list include at least: The database searches for sites. For positive + sites, based on the drug efficacy analysis, whether the medication is configured, and the selection of hotspot genes, determine whether to display them in the list of genes of unknown significance. The drug analysis display location includes tables for single-target / multi-target / immune / chemotherapy / toxicity and side effect drug analysis.
6. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: When generating reference information for the application of targeted / immunotherapy drugs, the automatic report generation module uses preset drug efficacy evaluation rules to make judgments based on the specific drug information entered by the system for different cancer types and different susceptibility gene types. The specific drug efficacy evaluation rules are: When multiple genes or sites correspond to exactly the same drugs, and X of the N genes have weaker therapeutic effects and Y have better therapeutic effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the therapeutic effect is judged to be good; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the therapeutic effect is judged to be weak; results that do not meet the above conditions are judged to be fair therapeutic effects; For the toxicity and side effect assessment, if X of the N genes have severe toxicity and side effects and Y have minor toxicity and side effects, when the value of X / N is ≥2 / 3 and the value of Y / N is <1 / 3, the toxicity and side effect are judged to be severe; when the value of Y / N is ≥2 / 3 and the value of X / N is <1 / 3, the toxicity and side effect are judged to be minor; the results that do not meet the above conditions are judged to be moderate toxicity and side effects.
7. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The system also includes a pathology management module, which is used to manage pathology information, and the pathology information at least includes newly added pathologies, recorded pathology names, gene numbers, guideline versions, and ClinicalTrail update time.
8. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The system also includes a gene management module, which is used to manage gene information, which includes at least gene ID, variation type, and label, and is used to export / import gene information and medication information according to pathology.
9. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The system also includes a drug management module, which is used to manage drug information. The drug information at least includes drug name, application display, and sorting position.
10. The pathological gene detection report generation system based on natural language processing according to claim 1 is characterized in that: The automatic report generation module automatically generates a test report based on template technology and the algorithm generation results of the above-mentioned stages in combination with natural language technology.
Citation Information
Patent Citations
A method and system for automatically interpreting and managing tumor gene detection reports
CN118629575B
ATP7B gene mutation next-generation sequencing automatic analysis and interpretation method and report system
CN112233725A
System for rapidly identifying drug identification sites and application thereof
CN117577182A
Method and system for automatically interpreting and managing tumor gene detection report
CN118629575A
Methods and systems for interpretation and reporting of sequence-based genetic tests
EP3161698A1