Gene detection report automatic generation method and device
By combining a pre-trained report generation model and a large language model with an authoritative knowledge base, the problem of existing systems being unable to perform comprehensive reasoning and dynamic adaptation is solved, achieving efficient and consistent quality automatic generation of gene testing reports and supporting personalized output.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU FEIRAN MICROBIOTECHNOLOGY CO LTD
- Filing Date
- 2025-11-28
- Publication Date
- 2026-05-12
AI Technical Summary
Existing microbial gene testing report generation systems cannot achieve comprehensive reasoning, making it difficult to generate reports with clinical guidance significance. Furthermore, they cannot dynamically adapt to different application scenarios and user needs, resulting in lagging knowledge updates and high maintenance costs.
Employing a pre-trained report generation model, based on gene testing results and target metadata, the system generates reports through a large language model. Combining authoritative knowledge bases and expert-reviewed knowledge fragments, it achieves end-to-end automatic report generation, supports personalized report output for different user roles and scenarios, and updates in real time.
It achieves efficient and consistent report generation, eliminates subjective differences in expert interpretation, supports large-scale concurrent processing, and can generate reports of varying depths according to different user needs, ensuring the professionalism and readability of the reports.
Smart Images

Figure CN122024988A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene technology, and in particular to a method and device for automatically generating gene testing reports. Background Technology
[0002] With the continuous maturation of high-throughput sequencing technology and its decreasing cost, microbial gene detection has gradually moved from basic scientific research to various practical applications such as clinical diagnosis, environmental monitoring, food safety, and public health. High-throughput sequencing technology can perform unbiased, high-coverage detection of the genetic material of all microorganisms in a sample, thereby comprehensively and accurately analyzing the microbiome composition of the target region, including the types, abundance, and functional potential of various microorganisms such as bacteria, viruses, fungi, and archaea. This capability has demonstrated its immense value in areas such as tracing the source of infectious disease outbreaks, hospital infection control, environmental and ecological assessment, food contamination tracing, and research on the relationship between the microbiome and host health.
[0003] A typical microbial gene sequencing workflow usually includes key steps such as sample collection, nucleic acid (DNA / RNA) extraction, library construction, high-throughput sequencing, bioinformatics analysis, and final report generation. The bioinformatics analysis stage involves multiple complex steps, including raw sequencing data quality control, sequence alignment or taxonomic annotation, species abundance calculation, functional prediction, and statistical evaluation, and has gradually become automated and standardized. Report writing at the end of the workflow is a crucial step connecting data analysis results with user interpretation, directly affecting the interpretability and clinical or application value of the test results. However, report writing is still mostly done manually.
[0004] Currently, the generation of microbial gene testing reports mainly relies on manual writing or semi-automated template-filling systems. In manual writing, technical professionals must write test conclusions, clinical significance interpretations, and recommended measures based on multi-dimensional information such as sequencing data, species annotation, functional prediction, and drug resistance gene analysis, combined with medical or industry knowledge. In semi-automated systems, preset report templates are typically used, and rule engines or scripts fill the analysis results into corresponding fields, such as species abundance tables, pathogenic bacteria lists, and drug resistance gene detection information.
[0005] Rule-based or template-based report generation systems can only provide fill-in-the-blank outputs and cannot perform comprehensive reasoning based on the logical relationships between test results. For example, when multiple potential pathogens and their corresponding drug-resistant genes are detected simultaneously, the system struggles to automatically generate comprehensive assessments with clinical guidance, such as infection risk level judgments and treatment recommendation priority rankings. Furthermore, different application scenarios (such as hospital ICUs, food companies, or environmental monitoring stations) have significantly different focuses regarding report content, and existing template-based systems cannot dynamically adapt to user needs. In addition, when test results are abnormal or contradictory (such as the detection of low-abundance but highly pathogenic bacteria), the system cannot provide reasonable explanations or confidence levels, making it difficult for users to determine the reliability of the results.
[0006] With the emergence of new pathogens, the evolution of drug resistance mechanisms, and the updating of clinical guidelines, report content needs continuous iteration. However, traditional rule engines rely on manually writing and maintaining a large number of if-else logic or keyword mapping tables, making it difficult to respond quickly to knowledge updates. Once a new detection indicator or interpretation dimension is added, it is often necessary to redesign the template structure or even reconstruct the entire report generation process, severely limiting the system's flexibility. Summary of the Invention
[0007] To address existing technical problems, this invention provides a method and computing device for automatically generating gene testing reports. This method enables standardized output, eliminates subjective differences in interpretation among different experts, ensures that all users receive reports of uniform quality, and generates different prompt words based on different users, thereby generating reports of varying depths according to different user needs.
[0008] In a first aspect, a method for automatically generating a gene testing report is provided, comprising: acquiring gene testing result data; forming prompt word input data for a pre-trained report generation model based on the gene testing result data; and outputting a target interpretation report based on the prompt word input data and the report generation model.
[0009] In a second aspect, a computing device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the gene testing report automatic generation method provided in the embodiments of this application.
[0010] This application generates interpretation reports corresponding to gene testing results data through a pre-trained report generation model. Since the report generation model is a trained model, it can achieve end-to-end output, large-scale concurrent processing, and easily handle a large number of users, thus achieving efficient output. Moreover, the trained report generation model can integrate multi-dimensional knowledge such as the latest scientific literature, clinical guidelines, and pharmacogenomics, and can be updated in real time to achieve standardized output, eliminating subjective differences in interpretation among different experts and ensuring that all users receive reports of consistent quality. Furthermore, it generates different prompt words based on different users, thereby generating reports of varying depths according to different user needs. Attached Figure Description
[0011] Figure 1 This is an application environment diagram of a gene testing report automatic generation method in one embodiment; Figure 2 This is a flowchart of a method for automatically generating gene testing reports in one embodiment; Figure 3 This is a flowchart illustrating the process of generating prompt input data in an automatic gene testing report generation method according to one embodiment; Figure 4 This is a flowchart illustrating the output of a target interpretation report in an automatic gene testing report generation method in one embodiment; Figure 5 This is a flowchart of a method for automatically generating gene testing reports in another embodiment; Figure 6 This is a schematic diagram of an automatic gene testing report generation device in one embodiment; Figure 7 This is a schematic block diagram of a computing device provided in one embodiment; Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in one embodiment. Detailed Implementation
[0012] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0014] In the following description, the expression “some embodiments” refers to a subset of all possible embodiments. However, it should be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] Existing report generation systems suffer from weak semantic understanding and repetitive, mechanical content. Traditional template-based or rule-driven report generation methods can only fill in fields, failing to perform semantic-level correlation analysis and logical reasoning on microbial test results (such as species composition, functional pathways, drug resistance genes, virulence factors, etc.), resulting in reports lacking professional depth and clinical readability. In traditional template-based replacement methods: the report generation program defines several report templates according to sample type or testing purpose, then sets conditional rules (e.g., "If E. coli is detected and abundance >1%, then insert 'Indicate possible intestinal flora imbalance'"), and finally performs field replacement to generate the final document. Such reports are generic and lack personalized analysis for each sample.
[0016] Traditional report generation systems suffer from a lack of personalization and scenario adaptability in their generated reports. Existing report generation systems struggle to dynamically adjust the report's focus and presentation based on different user roles (such as clinicians, disease control personnel, researchers, or corporate quality inspectors) or application scenarios (such as infection diagnosis, hospital infection monitoring, and food safety assessment).
[0017] Furthermore, traditional report generation systems suffer from outdated knowledge and high maintenance costs. The rules used in traditional report interpretation systems rely on manual maintenance of interpretation logic (such as if-else statements) and pathogen lists, making it difficult to integrate the latest pathogen information, drug resistance mechanisms, clinical guidelines, or industry standards in a timely manner. When new evidence needs to be added, the entire discrimination logic must be considered, making local adjustments impossible and requiring continuous manual intervention, as automatic updates and upgrades are not possible.
[0018] like Figure 1 As shown, Figure 1 This diagram illustrates the application environment of an automatic gene testing report generation method in one embodiment. A gene sequencer 20 communicates with a computing device 10. The gene sequencer 20 outputs gene testing result data, for example, in a microbial testing scenario, to obtain raw microbial gene testing data. The computing device 10 receives the gene testing result data and, based on this data, generates a corresponding target interpretation report using a pre-trained report generation model. The computing device 10 stores the trained report generation model, and the training of the report generation model can also be implemented within the computing device 10. In an optional implementation, the automatic gene testing report generation method can also be implemented within the gene sequencer 20.
[0019] Please see Figure 2 This is a flowchart of a method for automatically generating gene testing reports according to an embodiment of this application. The method for automatically generating gene testing reports is applied in a computing device and includes the following steps: S10. Obtain gene testing results data.
[0020] In this embodiment, the gene detection result data represents the detection data output by the gene sequencer. Subsequent analysis of the gene detection result data yields bioinformatics analysis results.
[0021] S11. Based on the gene testing results data, generate prompt word input data for a pre-trained report generation model.
[0022] In this embodiment, the prompt input data represents speech data that the report generation model can understand and that is associated with the gene testing result data. The gene testing result data can be preprocessed, key information extracted, and formatted to generate prompts suitable for the report generation model. The report generation model can be a large language model.
[0023] S12. Based on the input data of the prompt words, the target interpretation report is output through the report generation model.
[0024] In this embodiment, the report generation model is trained on a training dataset derived from a large amount of genetic data and relevant medical literature, thus providing data-driven, scientifically based interpretations. The report generation model can be periodically retrained with the latest research data to maintain the timeliness of its knowledge and provide interpretations based on the latest scientific findings, in contrast to traditional human writing, where human consultants need to continuously learn to keep up with the latest research. Because the training dataset comes from real-world case scenarios, the report generation model can learn medical testing logic during training. In genetic testing scenarios, it can learn to understand medical terminology (such as "heterozygous mutation" and "pathogenicity") and, using report tags corresponding to real cases as targets, learn how to organize a coherent report. For example, it learns that "TP53 gene mutation" is often associated with "increased cancer risk," rather than simply using a template. The report generation model can also learn the knowledge structure from professional literature and clinical guidelines. For example, it can discover the typical sections included in a report from massive amounts of data. During training, the report generation model learns to identify genotype-phenotype relationships and expresses them according to the strength of evidence. The report generation model also learns the standardized expressions used by experts in writing.
[0025] In the above embodiments, a pre-trained report generation model generates interpretation reports corresponding to gene detection results data. Since the report generation model is a trained model, it can achieve end-to-end output, large-scale concurrent processing, and easily handle a large number of users, thereby achieving efficient output. Moreover, the trained report generation model can integrate multi-dimensional knowledge such as the latest scientific literature, clinical guidelines, and pharmacogenomics, and can be updated in real time, achieving standardized output, eliminating subjective differences in interpretation among different experts, and ensuring that all users receive reports of uniform quality. Furthermore, different prompt words are generated based on different users, thereby generating reports of different depths according to different user needs.
[0026] In some embodiments, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating the process of generating prompt input data in an automatic gene testing report generation method according to one embodiment. The step of generating prompt input data for a pre-trained report generation model based on the gene testing result data includes the following steps: S30. Analyze the gene detection results to obtain bioinformatics analysis results.
[0027] In this embodiment, the bioinformatics analysis results represent key information extracted from the gene detection results data. The gene detection results data is high-throughput sequencing data. After preprocessing, filtered sequencing data is obtained. Each read length in the filtered sequencing data is compared with a standard database (such as Minikraken, Standard, etc.) to obtain the species in the gene detection results data. Then, the Kraken2 results are used for correction to estimate the relative abundance of species and obtain a relative abundance table. Megahit annotation is used, and then ABRicate is used to obtain bacterial resistance genes and virulence genes for analysis.
[0028] Bioinformatics analysis results include, but are not limited to: the species name of the detected microorganisms, the relative abundance of the detected species, whether they are pathogenic, the microbial category, the microbial diversity score, and whether lactobacilli are dominant. Microbial categories include fungi, bacteria, viruses, etc.
[0029] S31. Obtain the target metadata of the detection item.
[0030] In this embodiment, target metadata represents basic project data related to the testing project, including but not limited to sample type, testing purpose, user role, submitting institution, and clinical symptoms. This metadata serves as contextual information to guide the contextual adaptation and content focus of subsequent report generation. For example, sample types may include blood, sputum, vaginal secretions, food, and water. Testing purposes may include infection diagnosis, hospital infection control, and food safety assessment. User roles may include clinicians, patients awaiting testing, and researchers. Clinical symptoms may include textual descriptions such as abnormal vaginal discharge. In some embodiments, a user interface may be provided for selecting or filling in text related to the target metadata, such as providing a user role option, allowing users to select their role. Target metadata can be obtained based on the user interface. Users configure relevant testing project information on the user interface, and after configuration, target metadata is extracted from the configured testing project information.
[0031] S32. Based on the target metadata, query the target knowledge fragments associated with the target metadata.
[0032] In this embodiment, knowledge bases for various fields can be pre-established. Each knowledge base integrates authoritative medical guidelines, pathogen data, clinical interpretation rules, and historical expert-reviewed cases, among other authoritative knowledge. Furthermore, the knowledge bases are regularly updated after expert review to ensure they are up-to-date and authoritative. The knowledge base module provides authoritative, timely, and scenario-relevant medical knowledge support for the report generation model. The knowledge base is divided into three categories, all reviewed and confirmed by domain experts before being added: authoritative guidelines and consensus statements, such as the "Expert Consensus on the Clinical Application of Vaginal Microecological Evaluation"; structured pathogen databases, such as NCBI's species classification; and authoritative research papers in related fields.
[0033] For example, through paragraph segmentation and entity recognition, the knowledge required for original interpretation is transformed into a unified structured format. A specific example in the knowledge base is as follows: { "id": "consensus_2023_sec3.2_1", "source": "Chinese Expert Consensus on the Diagnosis and Treatment of Vaginal Microecology (2023)", Section 3.2 Molecular Diagnosis of Bacterial Vaginosis "text": "In metagenomic testing, a relative abundance of Gardnerella vaginalis >10% and Lactobacillus <70% suggests a high risk of vaginal viral infection (BV), but a comprehensive judgment should be made in conjunction with clinical symptoms." "entities": { "pathogens": ["Gardnerella vaginalis"], "conditions": ["bacterial vaginosis"], "populations": ["non-pregnant women"] }, "embedding": [0.23, -0.45, ..., 0.89] / / 1024 dimensions }
[0034] Optionally, querying the target knowledge fragment associated with the target metadata based on the target metadata includes: Based on the target metadata, determine the target knowledge base that matches the domain corresponding to the target metadata; The target metadata is concatenated with the gene detection result data to obtain the query statement; The target knowledge segment is obtained by searching the target knowledge base for one or more knowledge segments that are most semantically similar to the query statement.
[0035] By selecting the knowledge base using target metadata, a high degree of relevance between the knowledge base and the current testing project is ensured, avoiding searches in irrelevant knowledge bases and improving retrieval accuracy. Concatenating the target metadata with gene testing results into a query statement ensures that the query includes basic information about the testing project and specific test results. This allows for a more comprehensive expression of the query intent during knowledge base retrieval, resulting in more relevant knowledge fragments and enriching the context. Using semantic similarity retrieval, rather than keyword matching, captures the deep semantic similarity between the query statement and knowledge fragments. Even with different expressions, relevant content can be found, improving retrieval recall and accuracy. Since the target knowledge base is determined first through metadata, the search scope is narrowed, thus improving retrieval efficiency. When new testing projects or new knowledge bases are introduced, the system can be easily expanded through the mapping relationship between metadata and the knowledge base without modifying the core retrieval mechanism. Target metadata can include basic patient information, testing types, etc., making queries more personalized and retrieval results closer to specific scenarios. Target knowledge fragments provide authoritative and relevant knowledge support for generating interpretation reports, making the report content more accurate and reliable.
[0036] For example, when processing a vaginal discharge sample, the knowledge base module concatenates the detection results with metadata into a query statement, selects the 10 most relevant statements, and outputs them as a plain text list for the report generation module to use. Examples of knowledge base fragments extracted during retrieval include: "A healthy vaginal microecology is characterized by a predominance of lactobacilli (>70%); co-occurrence of Gardnerella vaginalis and Atopobium vaginae suggests a high risk of bacterial vaginosis (BV)." and "Detection of Candida albicans during pregnancy should raise suspicion of vulvovaginal candidiasis (VVC), but low abundance in non-pregnant women may indicate colonization."
[0037] S33. Based on the bioinformatics analysis results, target metadata, and the target knowledge fragments, generate prompt word input data.
[0038] In this embodiment, the prompt input data represents the input data fed into the report generation model.
[0039] Optionally, generating the prompt word input data based on the bioinformatics analysis results, the target metadata, and the target knowledge fragment includes: Match the target prompt word template that matches the target metadata from the prompt word template library; Based on the bioinformatics analysis results, the target metadata, and the target knowledge fragments, a portion of the content in the target prompt word template is filled in to obtain the prompt word input data.
[0040] In this embodiment, to improve the efficiency of report generation, various types of prompt word templates can be pre-configured. The types of prompt word templates can be combined based on at least one of the following data: user role, testing item, testing field, etc. The template library can contain multiple templates, and different templates can be selected for different testing items, different audiences (such as clinicians and patients), or different report types (such as detailed reports and summary reports). By matching templates with target metadata, the most suitable template can be automatically selected, making report generation more intelligent. Integrating bioinformatics analysis results, target metadata, and retrieved target knowledge fragments into the template ensures that the report content includes both raw data and relevant background knowledge, making the report more comprehensive and accurate. The template can be designed to include fixed and variable parts. The fixed part ensures that key information is not omitted, while the variable part allows the content to be adjusted according to specific data.
[0041] For example, based on the target metadata, the corresponding target prompt word template is selected. The prompt word template in this example is as follows: You are a gynecological microecology expert, writing a professional interpretation of a vaginal secretion metagenomic testing report. Patient's chief complaint: {clinical_indication}, age {patient_age} years, {pregnancy_status}.
[0042] Based on the following test results and the latest clinical guidelines, please generate a well-structured, professionally worded vaginal microecology assessment report that provides valuable guidance for diagnosis and treatment. The report should include the following sections: [Microbial Ecosystem Status Assessment] Determine whether it is a Lactobacillus-dominant type.
[0043] [Key Pathogen Interpretation] The clinical significance of Gardnerella vaginalis, Atopobium vaginae, and Candida albicans detected was explained, with particular attention paid to co-occurrence patterns and abundance thresholds.
[0044] Therefore, the patient's chief complaint, age, microecological status assessment, and interpretation of key pathogens can be filled in, and the variable parts can be adjusted.
[0045] In the above embodiments, the bioinformatics analysis results, target metadata, and retrieved target knowledge fragments are integrated into the template to ensure that the report content includes both the original data and relevant background knowledge, thus making the report generated by the report generation model more comprehensive and accurate.
[0046] In some embodiments, such as Figure 4 As shown, Figure 4 This is a flowchart illustrating the process of automatically generating a gene testing report in one embodiment, whereby the step of outputting a target interpretation report based on the prompt input data and through the report generation model includes: S40. Based on the input data of the prompt words, generate an interpretation report to be fine-tuned through the report generation model.
[0047] In this embodiment, the initial interpretation report refers to a report that needs modification. It can be the initially generated report, or a report that has been fine-tuned and still requires further adjustments.
[0048] S41. Obtain the verification prompt word template.
[0049] In this embodiment, the verification prompt template includes, but is not limited to, at least one of the following: consistency check, completeness check, accuracy check, logic check, language style check, format check, etc. The consistency check checks whether the report content is consistent with the original data and knowledge fragments. The completeness check checks whether the report includes all necessary parts, such as test results, clinical significance, and recommendations. The accuracy check checks whether the professional terminology, numerical values, and explanations in the report are accurate. The logic check checks whether the report's logic is reasonable, such as whether the risk assessment matches the results. The language style check checks whether the report's language is suitable for the target audience's (e.g., patients or doctors) understanding level and whether it is objective and clear. The format check checks whether the report's format meets the requirements, such as titles, paragraphs, and lists.
[0050] S42. Based on the interpretation report to be fine-tuned, bioinformatics analysis results, target metadata and target knowledge fragments, fill in the verification prompt word template and generate verification input data.
[0051] In this embodiment, the validation input data includes the parts of the interpretation report that need to be adjusted. For example, the interpretation report mentions "Candida albicans detected" but does not explain its low abundance of 1.8%. The validation input data can be supplemented by: Please rewrite only the part about Candida albicans in the "Key Pathogen Interpretation" section, which must include the following elements: - Clearly indicate "low abundance (1.8%)"; - Explain that it may be colonization when there are no typical symptoms; - Use the standard name "Candida albicans".
[0052] S43. Based on the test input data, form the input of the report generation model and output a fine-tuned interpretation report.
[0053] In this embodiment, the input data to be tested is input into the report generation model again, and the report generation model adjusts the interpretation report to be fine-tuned based on the verification input data.
[0054] S44. Determine whether the revised interpretation report needs further adjustments.
[0055] If the fine-tuned interpretation report needs further adjustments, return to execute S41. If the fine-tuned interpretation report needs further adjustments, execute S45 and use the final adjusted output report as the adjusted interpretation report.
[0056] In this embodiment, the adjustment process can be carried out by fine-tuning thematic content, that is, adjusting one theme at a time. After all themes have been adjusted in sequence, it can be determined that the fine-tuned interpretation report does not need further adjustment. Alternatively, multiple themes can be adjusted at a time. Multiple fine-tunings can also be performed on the same theme. Whether the fine-tuned interpretation report needs further adjustment can be determined by the user selecting the corresponding option on the user interface. After one or more adjustments, the final adjusted interpretation report is obtained.
[0057] S46. Based on the initial interpretation report and the adjusted interpretation report, generate a verification interpretation report, and based on the verification interpretation report, generate a target interpretation report.
[0058] In this embodiment, the initial interpretation report and the adjusted interpretation report are compared for text similarity. The initial interpretation report is then modified, and the modified sentences are marked to generate a verification interpretation report. This allows users to visually observe whether the fine-tuned parts are correct in the verification interpretation report.
[0059] Optionally, generating the verification interpretation report based on the initial interpretation report and the fine-tuning report includes at least one of the following: Calculate the text similarity between sentences in the initial interpretation report and sentences in the adjusted interpretation report under the same topic content, and replace sentences in the initial interpretation report with corresponding sentences in the adjusted interpretation report whose text similarity is lower than a preset similarity threshold; Replace non-standard terms in the initial interpretation report with standard terms; The sentences that were modified in the initial interpretation report were marked.
[0060] In this embodiment, the initial interpretation report and the adjusted interpretation report are segmented into sentences. For each sentence under the same topic, preprocessing is performed to construct a TF-IDF vector generator. The sentences from both reports are merged and vectorized together. Then, the vectors of each sentence in both the initial and adjusted interpretation reports are obtained. The cosine similarity between each sentence in the initial interpretation report and the corresponding sentence in the adjusted interpretation report is calculated. Sentences with similarity scores below a preset similarity threshold are replaced. A terminology mapping dictionary can be constructed, and each sentence is checked for non-standard terms, which are then replaced to achieve localized modification. The locally modified sentences are then labeled for easy user observation of the fine-tuned content.
[0061] In the above embodiments, after generating the initial interpretation report, the input data of the report generation model can be adjusted to fine-tune the initial interpretation report multiple times. The content of the interpretation report to be fine-tuned is then reviewed multiple times to obtain an adjusted interpretation report. A verification interpretation report is then obtained by comparing the initial interpretation report and the adjusted interpretation report, and locally modified sentences are marked to facilitate users' intuitive observation of the fine-tuned content.
[0062] Optionally, generating the target interpretation report based on the verification interpretation report includes at least one of the following: At least a portion of the numerical data in the gene testing results is converted into charts and embedded into the verification and interpretation report; Based on the species data in the gene detection results, a species composition map is generated; When the verification interpretation report triggers the conditions for manual review, the associated data of the verification interpretation report is sent to the expert user's terminal device so that the expert user can review the verification interpretation report and obtain the modified verification interpretation report.
[0063] In this embodiment, structured microbial detection results are transformed into intuitive, professional charts that conform to clinical reading habits, and automatically embedded into the verification and interpretation report to obtain the target interpretation report. For species data in the gene detection results, a stacked bar chart of species composition can be generated to visually display the relative abundance of dominant bacteria and pathogenic bacteria. For example, only species with an abundance ≥1% are displayed, and the rest are categorized as "Others". Following medical visualization standards, beneficial bacteria (Lactobacillus): dark green; BV-related bacteria (G. vaginalis, A. vaginae, etc.): orange-red; fungi (Candida); other: light gray (#D3D3D3).
[0064] In this embodiment, the conditions for manual review include, but are not limited to, at least one of the following: detection of a high-risk pathogen, contradictory results or high uncertainty, treatment recommendations involving contraindications, and the verification report containing keywords indicating the need for reconfirmation. Keywords requiring reconfirmation include, but are not limited to, keywords such as "uncertainty" and "recommendation verification." If no manual review conditions are triggered, the report can be output directly. The expert user's terminal device can be a web application that displays the original detection data, the initial interpretation report generated by the large model, and allows relevant professionals to modify the text and submit the revised interpretation report, thus obtaining the final target interpretation report. Through this human-machine collaboration mechanism, while maximizing automation efficiency, the stringent requirements of safety, traceability, and compliance in medical scenarios are met.
[0065] In the above embodiments, the numerical data and species data in the gene testing results can be graphically represented in the target interpretation report, making it easier for users to intuitively observe the interpretation report. When the verification of the interpretation report triggers the conditions for manual review, it is directly sent to experts for compliance, resulting in the final target interpretation report. Through a human-machine collaboration mechanism, the automation efficiency is maximized while meeting the stringent requirements of medical scenarios for safety, traceability, and compliance.
[0066] In some embodiments, the method further includes: Obtain a training dataset, wherein each training sample in the training dataset includes detection sample data and a label report for the detection sample data; Based on the detected sample data, the input to the report generation model in training is formed, and the current training report corresponding to the detected sample data is output. The report generation model is trained based on the textual differences between the current training report and the labeled report.
[0067] In this embodiment, the training dataset may include training samples from various medical fields. The detection sample data are real-world detected data, and the label report can be a standard report written by an expert. For example, vaginal microecology detection sample data and a real vaginal microecology detection report written by a doctor form an instruction-response pair, which is used as the training dataset. In some embodiments, some large, pre-trained models can be acquired first, and then the large models can be fine-tuned using the training dataset to obtain a report generation model. This updates a small number of adapter parameters, retains the original model knowledge, and reduces GPU memory consumption.
[0068] The current training report represents the report generation model's output based on the detection sample data during the current iteration. Since the detection sample data corresponds to a standard label report, the textual difference between the current training report and the label report can be calculated, and this textual difference can be used as part of the loss value. The textual difference can be calculated using text similarity calculation methods. The label report is used as the training target, aiming to make the text of the output current training report as close as possible to the text of the label report.
[0069] In this embodiment, to further suppress model illusions and improve the rigor of clinical descriptions, in addition to the aforementioned fine-tuning, the system analyzes the presence of non-standard terms in the current training report and further penalizes the loss in the current iteration based on these non-standard terms, i.e., increasing the current term penalty loss. The more non-standard terms appear, the greater the current term penalty loss. The current total loss value is then calculated based on the current text difference loss and the current term penalty loss.
[0070] Optionally, during each training process, the text difference between the current training report and the label report is calculated to obtain the current text difference loss; Iterate through each word in the current training report, obtain words that belong to non-standard terms, and determine the current term penalty loss based on the words of non-standard terms; Calculate the current total loss value based on the current text difference loss and the current term penalty loss; When the current total loss value is less than or equal to the preset loss threshold, training is stopped, and the report generation model after training is stopped is used as the pre-trained report generation model; when the current total loss value is greater than the preset loss threshold, backpropagation is performed to update the parameters in the report generation model during training, and the report generation model continues to be trained.
[0071] In this embodiment, textual difference loss aims to generate reports that are as close as possible in content to the tagged reports, ensuring accuracy in facts and descriptions and improving overall report accuracy. Terminology penalty loss directly penalizes non-standard terminology, allowing the model to gradually learn to use industry-standard terminology during training, improving the professionalism and readability of the reports and ensuring terminology standardization. Combining these two losses, considering both content accuracy and terminology standardization, optimizes the model in both aspects, generating reports that are both accurate and professional. A preset loss threshold serves as a criterion for judging whether model training has converged. Training stops when the total loss is below the threshold to avoid overfitting and save training time. Parameters are updated via backpropagation; when the loss exceeds the threshold, backpropagation updates the model parameters, allowing the model to gradually improve.
[0072] For example, for 1000 real test data points, two interpretation reports are generated for each. Obstetricians and gynecologists then select 800 reports that can be distinguished as good or bad. Typical problems with bad reports include: fabricating undetected pathogens (e.g., "Trichomonas detected" when no data is available), absolute statements (e.g., "diagnosed BV"), ignoring low-abundance annotations, and recommendations that do not conform to guidelines. Subsequently, a mapping table is constructed from the *Handbook of Clinical Microbiology* and *Obstetrics and Gynecology (10th Edition)*. In addition to the standard cross-entropy loss, a current terminology penalty loss is added to penalize the model for using unrecommended terms, thus forcing the model to use standard medical terminology. The mapping table is shown below. { Standard Terminology: ["Bacterial vaginosis", "Vulvovaginal candidiasis", "Candida albicans", "Lactobacillus"], Non-standard terminology: ["vulvovaginal candidiasis", "vulvovaginal candidiasis", "good bacteria", "bad bacteria"] } In the above embodiments, the text difference loss requires the model to generate a report that is as close as possible to the labeled report in terms of content. This ensures the accuracy of the report in terms of facts and descriptions, and improves the accuracy of the report. The terminology penalty loss directly penalizes non-standard terms, so that the model gradually learns to use industry standard terms during training, which improves the professionalism and readability of the report and ensures the standardization of terminology. Through training on a large dataset, the model can efficiently output interpretation reports that conform to medical standards.
[0073] like Figure 5 As shown, Figure 5 Here is a flowchart of a method for automatically generating gene testing reports in another embodiment, which includes the following steps: S51, Obtain gene testing results data.
[0074] S52, based on the gene testing results data, obtain the target metadata.
[0075] S53, based on gene detection results, yields bioinformatics analysis results.
[0076] S54, Match the target prompt word template based on the target metadata.
[0077] S55, Match target knowledge fragments in the knowledge base based on target metadata.
[0078] S56. Based on the bioinformatics analysis results, target metadata, and target knowledge fragments, fill in part of the content in the target prompt word template to obtain the prompt word input data.
[0079] S57, based on the input data of prompt words, generates a report generation model, which needs to be fine-tuned to interpret the report.
[0080] S58, obtain the verification prompt word template.
[0081] S59, based on the interpretation report to be fine-tuned, the bioinformatics analysis results, the target metadata, and the target knowledge fragment, fill in the verification prompt word template to generate verification input data.
[0082] S510, based on the test input data, forms the input to the report generation model and outputs a finely tuned interpretation report.
[0083] S511, determine whether the fine-tuned interpretation report needs further adjustments.
[0084] If the fine-tuned interpretation report needs further adjustments, return to execute S58. If the fine-tuned interpretation report needs further adjustments, execute S512 and use the final adjusted output report as the adjusted interpretation report.
[0085] S513, Based on the initial interpretation report and the adjusted interpretation report, generate a verification interpretation report.
[0086] S514, determine whether the verification and interpretation report triggers the conditions for manual review.
[0087] If the verification and interpretation report triggers the conditions for manual review, execute S515 to send the associated data of the verification and interpretation report to the expert user's terminal device so that the expert user can review the verification and interpretation report and obtain the modified verification and interpretation report. Visualize the data in the gene testing results and embed it into the modified verification and interpretation report to obtain the target interpretation report. Alternatively, the data in the gene testing results can be visualized first and embedded into the verification and interpretation report before being sent to the expert for review. If the verification and interpretation report does not trigger the conditions for manual review, execute S516 to visualize the data in the gene testing results and embed it into the verification and interpretation report to obtain the target interpretation report.
[0088] Understandably, the step numbers in the flowchart above only indicate different steps. Some steps can be executed in any order. For example, S52 and S53 can be executed simultaneously or sequentially.
[0089] It is understood that one or more of the above embodiments can be applied to report output for various testing scenarios, including but not limited to at least one of the following: female reproductive tract testing, intestinal microecology testing, oral microecology testing, nasal microecology testing, male genital microecology testing, and other testing scenarios.
[0090] The above one or more embodiments have at least the following characteristics: Data input and analysis stage: The system receives raw microbial gene detection data from the high-throughput sequencing platform, and completes analysis such as sequence quality control, species annotation, functional prediction, and identification of drug resistance genes / virulence factors through the built-in bioinformatics analysis pipeline, generating a structured detection result data table.
[0091] In the target metadata acquisition stage, the interpretation system obtains metadata related to the test by having the user select and fill in text boxes on the webpage. This metadata includes, but is not limited to, sample type (e.g., blood, sputum, vaginal secretions, food, water), testing purpose (e.g., infection diagnosis, hospital infection control, food safety assessment), user role (e.g., clinician, patient, researcher), submitting institution, and clinical symptoms (e.g., textual description of abnormal vaginal discharge). This metadata serves as contextual information, guiding the subsequent report generation in terms of contextual adaptation and content focus.
[0092] Knowledge Enhancement and Reasoning Stage: The system accesses a regularly updated knowledge base that has been reviewed by experts. This base integrates authoritative medical guidelines, pathogen data, clinical interpretation rules, and historical expert-reviewed cases. Through search-enhanced generation, the most relevant knowledge fragments to the current test results are dynamically injected into the input prompts of the large language model, ensuring the scientific accuracy and compliance of the generated content. In this step, the report generation system searches for newly published microbiology-related research, books, clinical guidelines, and industry standards based on keywords. After expert review confirms the professionalism of the documents, search-enhanced generation extracts the key findings from the documents.
[0093] The large language model report generation stage: Through structured prompting engineering and a controllable text generation mechanism, and using pre-set multi-stage prompt templates, the model performs inference tasks to guide the large language model in generating a professionally standardized, logically rigorous, and context-appropriate microbial testing report based on multi-source inputs (detection results, metadata, and knowledge base). The prompts for each stage include: ① a summary of key findings; ② pathogenicity risk assessment; ③ drug resistance and treatment recommendations; ④ explanation of public health significance; and ⑤ uncertainty and limitation prompts. The model outputs a draft report in natural language format, with content that is professional, logical, and readable.
[0094] Data visualization and text verification stage: While generating the text portion of the test report through large model question answering, the interpretation system calls the data visualization module. Based on the test results obtained in the first step and the metadata of the test items obtained in the second step, this module calls the corresponding Python visualization code under the test items in the system to generate visualization charts, such as species composition bar charts, heatmaps, drug resistance gene distribution maps, functional pathway enrichment maps, etc., and embeds them into the report to enhance the information delivery effect.
[0095] Meanwhile, the text verification module, based on the initial report draft obtained in step four, fills the prompt word template, allowing the large model to perform multi-dimensional verification on the generated text, including: grammatical correctness check, terminology consistency verification (compared to the standard medical terminology database), sensitive word filtering, and logical contradiction detection (such as "no pathogens detected" but the conclusion is "high risk of infection"), to ensure that the report content is accurate, compliant, and unambiguous.
[0096] Human-machine collaborative review and output stage: The generated draft report can be pushed to the expert review terminal for quick review or modification by professionals. Review comments can be fed back to the knowledge base for continuous model optimization. The final report is output in PDF, HTML, or structured JSON format.
[0097] This application significantly improves the intelligence level of microbial testing report generation by deeply integrating bioinformatics analysis, domain knowledge base, context-aware prompting engineering, and intelligent post-processing mechanisms. Compared with existing technologies, this application significantly improves report quality and professional accuracy. Traditional template-based systems can only perform field filling, resulting in mechanical content and a lack of logical reasoning. In contrast, this application introduces a large language model finely tuned with microbiology and obstetrics / gynecology / infectious diseases knowledge, combined with a retrieval-enhanced generation (RAG) mechanism, enabling the report content to possess semantic coherence, clinical logic, and guideline compliance. For example, in vaginal secretion testing, the system can not only identify the risk of BV co-occurrence of "Gardnerella vaginalis + Atopobium vaginae", but also combine lactobacillus abundance, patient symptoms, and the latest expert consensus to generate individualized assessments and cautious recommendations, avoiding overdiagnosis.
[0098] Compared to the typical 1-2 hours required for manual drafting of a complete microbial test report, this application achieves end-to-end automated generation, from raw sequencing data input to structured PDF report output, significantly improving test interpretation efficiency. It is suitable for high-throughput testing scenarios (such as disease control screening and health check centers), supporting the generation of tens of thousands of reports per day, significantly alleviating the bottleneck of professional manpower. Through a dynamic template mechanism driven by test item metadata, this application can automatically adapt to different sample types (such as blood, sputum, vaginal secretions), user roles (doctors, disease control personnel, researchers), and testing purposes (diagnosis, monitoring, research). For example, the same vaginal flora data can highlight treatment recommendations when presented to gynecologists, while emphasizing indicators such as diversity and flora interaction networks when presented to researchers. This ability to use data for multiple purposes and generate reports on demand avoids the redundant investment of traditional systems that require developing separate templates for each type of user, improving system flexibility.
[0099] In another aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the automatic gene testing report generation method described in any embodiment of this application.
[0100] In the computer program product, the optional implementation form of the program module architecture of the computer program that implements each step of the automatic gene testing report generation method can be an automatic gene testing report generation device.
[0101] Please see Figure 6 One embodiment of this application provides an automatic gene testing report generation device, comprising: an acquisition module 61 for acquiring gene testing result data; a generation module 62 for generating prompt word input data for a pre-trained report generation model based on the gene testing result data; and an output module 63 for outputting a target interpretation report based on the prompt word input data and the report generation model.
[0102] Optionally, the generation module 62 is also used for: The gene detection results data are analyzed to obtain bioinformatics analysis results; Obtain the target metadata of the detection project; Based on the target metadata, query the target knowledge fragments associated with the target metadata; Based on the bioinformatics analysis results, the target metadata, and the target knowledge fragments, the prompt word input data is generated.
[0103] Optionally, the generation module 62 is also used for: Based on the target metadata, determine the target knowledge base that matches the domain corresponding to the target metadata; The target metadata is concatenated with the gene detection result data to obtain the query statement; The target knowledge segment is obtained by searching the target knowledge base for one or more knowledge segments that are most semantically similar to the query statement.
[0104] Optionally, the generation module 62 is also used for: Match the target prompt word template that matches the target metadata from the prompt word template library; Based on the bioinformatics analysis results, the target metadata, and the target knowledge fragments, a portion of the content in the target prompt word template is filled in to obtain the prompt word input data.
[0105] Optionally, output module 63 is also used for: Based on the input data of the prompt words, the report generation model generates a fine-tuning interpretation report, wherein the fine-tuning interpretation report includes an initial interpretation report and a report after fine-tuning the initial interpretation report; Get the verification prompt word template; Based on the interpretation report to be fine-tuned, the bioinformatics analysis results, the target metadata, and the target knowledge fragment, the verification prompt word template is filled in to generate verification input data; Based on the test input data, the input of the report generation model is formed, and the fine-tuned interpretation report is output. If it is determined that the fine-tuned interpretation report needs to be further adjusted, the verification prompt word template is retrieved. If it is determined that the fine-tuned interpretation report does not need to be further adjusted, the final adjusted output report is used as the adjusted interpretation report. Based on the initial interpretation report and the adjusted interpretation report, the verification interpretation report is generated, and based on the verification interpretation report, the target interpretation report is generated.
[0106] Optionally, output module 63 is also used for: Calculate the text similarity between sentences in the initial interpretation report and sentences in the adjusted interpretation report under the same topic content, and replace sentences in the initial interpretation report with corresponding sentences in the adjusted interpretation report whose text similarity is lower than a preset similarity threshold; Replace non-standard terms in the initial interpretation report with standard terms; The sentences that were modified in the initial interpretation report were marked.
[0107] Optionally, output module 63 is also used for: At least a portion of the numerical data in the gene testing results is converted into charts and embedded into the verification and interpretation report; Based on the species data in the gene detection results, a species composition map is generated; When the verification interpretation report triggers the conditions for manual review, the associated data of the verification interpretation report is sent to the expert user's terminal device so that the expert user can review the verification interpretation report and obtain the modified verification interpretation report.
[0108] Optionally, a training module 64 is also included for: Obtain a training dataset, wherein each training sample in the training dataset includes detection sample data and a label report for the detection sample data; Based on the detected sample data, the input to the report generation model in training is formed, and the current training report corresponding to the detected sample data is output. The report generation model is trained based on the textual differences between the current training report and the labeled report.
[0109] Optionally, training module 64 is also used for: During each training process, the text difference between the current training report and the label report is calculated to obtain the current text difference loss; Iterate through each word in the current training report, obtain words that belong to non-standard terms, and determine the current term penalty loss based on the words of non-standard terms; Calculate the current total loss value based on the current text difference loss and the current term penalty loss; When the current total loss value is less than or equal to the preset loss threshold, training is stopped, and the report generation model after training is stopped is used as the pre-trained report generation model; when the current total loss value is greater than the preset loss threshold, backpropagation is performed to update the parameters in the report generation model during training, and the report generation model continues to be trained.
[0110] Those skilled in the art will understand that the structure of the automatic gene testing report generation device does not constitute a limitation on the device itself, and the various modules can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the controller in the computing device, or stored in software in the memory of the computing device, so that the controller can invoke and execute the operations corresponding to each module. In other embodiments, the automatic gene testing report generation device may include more or fewer modules than those shown in the figures.
[0111] like Figure 7 As shown, the computing device 10 includes a processor 13, a memory 14, and a communication interface 15. The processor 13, memory 14, and communication interface 15 communicate with each other via a bus. The computing device 10 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 10. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. The bus can include a path for transmitting information between various components of the computing device 10 (e.g., memory 14, processor 13, communication interface 15). The processor 13 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0112] Memory 14 may include volatile memory, such as random access memory (RAM). Processor 13 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0113] The memory 14 stores executable program code, and the processor 13 executes this executable program code to implement the functions of the aforementioned modules, thereby realizing the automatic gene testing report generation method. That is, the memory 14 stores instructions for executing the automatic gene testing report generation method. Alternatively, the memory 14 stores executable code, and the processor 13 executes this executable code to implement the functions of the aforementioned automatic gene testing report generation device, thereby realizing the automatic gene testing report generation method. That is, the memory 14 stores instructions for executing the automatic gene testing report generation method.
[0114] The communication interface 15 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 10 and other devices or communication networks.
[0115] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone. Figure 8 As shown, the computing device cluster includes at least one computing device 10. The memory 14 of one or more computing devices 10 in the computing device cluster may store the same instructions for executing the automatic gene testing report generation method. In some possible implementations, the memory 14 of one or more computing devices 10 in the computing device cluster may also each store a portion of the instructions for executing the automatic gene testing report generation method. In other words, a combination of one or more computing devices 10 can jointly execute the instructions for executing the automatic gene testing report generation method.
[0116] It should be noted that the memory 14 in different computing devices 10 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the automatic gene testing report generation device. That is, the instructions stored in the memory 14 of different computing devices 10 can implement the functions of one or more modules.
[0117] In another aspect, this application provides a computer-readable non-volatile storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to perform the steps of an automatic gene testing report generation method provided in any of the above embodiments of this application.
[0118] In another aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of an automatic gene testing report generation method as described in any embodiment of this application and / or a training method based on a joint model.
[0119] Those skilled in the art will understand that all or part of the processes in the methods provided in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0120] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automatically generating gene testing reports, characterized in that, include: Obtain genetic testing results data; Based on the gene detection results data, prompt word input data is generated for a pre-trained report generation model; Based on the input data of the prompt words, the target interpretation report is output through the report generation model.
2. The method for automatically generating gene testing reports as described in claim 1, characterized in that, The prompt word input data for forming the pre-trained report generation model based on the gene detection results data includes: The gene detection results data are analyzed to obtain bioinformatics analysis results; Obtain the target metadata of the detection project; Based on the target metadata, query the target knowledge fragments associated with the target metadata; Based on the bioinformatics analysis results, the target metadata, and the target knowledge fragments, the prompt word input data is generated.
3. The method for automatically generating gene testing reports as described in claim 2, characterized in that, The step of querying the target knowledge fragment associated with the target metadata based on the target metadata includes: Based on the target metadata, determine the target knowledge base that matches the domain corresponding to the target metadata; The target metadata is concatenated with the gene detection result data to obtain the query statement; The target knowledge segment is obtained by searching the target knowledge base for one or more knowledge segments that are most semantically similar to the query statement.
4. The method for automatically generating gene testing reports as described in claim 2, characterized in that, The process of generating the prompt word input data based on the bioinformatics analysis results, the target metadata, and the target knowledge fragment includes: Match the target prompt word template that matches the target metadata from the prompt word template library; Based on the bioinformatics analysis results, the target metadata, and the target knowledge fragments, a portion of the content in the target prompt word template is filled in to obtain the prompt word input data.
5. The method for automatically generating gene testing reports as described in claim 1, characterized in that, The process of outputting a target interpretation report based on the input data of the prompt words, through the report generation model, includes: Based on the input data of the prompt words, the report generation model generates a fine-tuning interpretation report, wherein the fine-tuning interpretation report includes an initial interpretation report and a report after fine-tuning the initial interpretation report; Get the verification prompt word template; Based on the interpretation report to be fine-tuned, the bioinformatics analysis results, the target metadata, and the target knowledge fragment, the verification prompt word template is filled in to generate verification input data; Based on the test input data, the input of the report generation model is formed, and the fine-tuned interpretation report is output. If it is determined that the fine-tuned interpretation report needs to be further adjusted, the verification prompt word template is retrieved. If it is determined that the fine-tuned interpretation report does not need to be further adjusted, the final adjusted output report is used as the adjusted interpretation report. Based on the initial interpretation report and the adjusted interpretation report, a verification interpretation report is generated, and based on the verification interpretation report, a target interpretation report is generated.
6. The method for automatically generating gene testing reports as described in claim 5, characterized in that, The process of generating the verification interpretation report based on the initial interpretation report and the adjusted interpretation report, and generating the target interpretation report based on the verification interpretation report, includes at least one of the following: Calculate the text similarity between sentences in the initial interpretation report and sentences in the adjusted interpretation report under the same topic content, and replace sentences in the initial interpretation report with corresponding sentences in the adjusted interpretation report whose text similarity is lower than a preset similarity threshold; Replace non-standard terms in the initial interpretation report with standard terms; The sentences that were modified in the initial interpretation report were marked.
7. The method for automatically generating gene testing reports as described in claim 5, characterized in that, The generation of the target interpretation report based on the verification interpretation report includes at least one of the following: At least a portion of the numerical data in the gene testing results is converted into charts and embedded into the verification and interpretation report; Based on the species data in the gene detection results, a species composition map is generated; When the verification interpretation report triggers the conditions for manual review, the associated data of the verification interpretation report is sent to the expert user's terminal device so that the expert user can review the verification interpretation report and obtain the modified verification interpretation report.
8. The method for automatically generating gene testing reports as described in claim 1, characterized in that, The method further includes: Obtain a training dataset, wherein each training sample in the training dataset includes detection sample data and a label report for the detection sample data; Based on the detected sample data, the input to the report generation model in training is formed, and the current training report corresponding to the detected sample data is output. The report generation model is trained based on the textual differences between the current training report and the labeled report.
9. The method for automatically generating gene testing reports as described in claim 8, characterized in that, The training of the report generation model based on the text difference between the current training report and the label report includes: During each training process, the text difference between the current training report and the label report is calculated to obtain the current text difference loss; Iterate through each word in the current training report, obtain words that belong to non-standard terms, and determine the current term penalty loss based on the words of non-standard terms; Calculate the current total loss value based on the current text difference loss and the current term penalty loss; When the current total loss value is less than or equal to the preset loss threshold, training is stopped, and the report generation model after training is stopped is used as the pre-trained report generation model; when the current total loss value is greater than the preset loss threshold, backpropagation is performed to update the parameters in the report generation model during training, and the report generation model continues to be trained.
10. A computing device cluster, characterized in that, The system includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method according to any one of claims 1-9.