Method for generating medical reports

A computer-aided method using pattern recognition and generative language models integrates diverse medical data sources to generate precise and consistent medical reports, addressing the inefficiencies and quality issues in existing manual processes.

WO2025229162A1PCT designated stage Publication Date: 2025-11-06MYSCRIBE GMBH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
PCT/EP2025/062032
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-04
Filing Date
2025-05-02
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

The creation of medical reports, particularly discharge summaries, is time-consuming and prone to errors due to the fragmented and inconsistent nature of electronic health record data, requiring significant manual effort and medical expertise, leading to variations in quality and completeness.

Method used

A computer-aided method involving pattern recognition, machine learning, and generative language models to analyze and integrate medical information from multiple sources, normalize timestamps, and enrich data with external knowledge bases to generate coherent and precise medical reports.

Benefits of technology

This method streamlines the report generation process, ensuring high-quality, comprehensive, and consistent medical documentation that supports efficient communication and decision-making among healthcare professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000085_0000
    Figure 00000085_0000
  • Figure 00000086_0000
    Figure 00000086_0000
  • Figure 00000087_0000
    Figure 00000087_0000
Patent Text Reader

Abstract

A computer-implemented method for creating medical reports, in particular physician's letters, comprises the acquisition of patient information from various sources of electronic health records, with the latter being available in natural-language sampling points. Medical units are determined by computer-aided pattern recognition and classified according to their event type. Event groups that have temporal, thematic or treatment-related relationships are formed. Relationship information is derived and stored in a structured data representation. In addition, these data are enriched with an external medical knowledge base and converted into a format that can be processed by a generative language model in order to generate the continuous text for the medical report.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR PROCEEDING MEDICAL REPORTS

[0002] TECHNICAL AREA

[0003] The present invention relates generally to the field of medical data processing and in particular to computer-aided methods using generative language models and specifically to a method for generating medical reports using a trained generative language model.

[0004] BACKGROUND

[0005] Medical reports are essential in healthcare. Medical reports, especially doctors' letters, form the basis for communication between different medical professionals and institutions and serve as essential sources of information in a patient's treatment process.

[0006] In everyday clinical practice, vast amounts of medical data are generated and recorded in electronic health records (EHRs). This data originates from diverse sources such as clinical examinations, laboratory tests, imaging procedures, medication prescriptions, and treatment documentation. The information is often fragmented, frequently presented as bullet points or brief notes, scattered across various systems and modules within the electronic health record.

[0007] Creating summary medical reports, such as discharge summaries, requires healthcare professionals to review, evaluate, and coherently integrate disparate pieces of information. This process is time-consuming and ties up valuable personnel resources. In clinical practice, this leads to delays in the creation and transmission of critical patient documents. Furthermore, manually compiling information from various sources carries the risk of overlooking relevant aspects or creating inconsistent presentations. The varying structures and terminology across different medical specialties further complicate the consistent integration of all relevant information.

[0008] The creation of summary medical reports often places a significant additional burden on medical staff. This extra work can lead to a decrease in the quality of patient care and / or an increase in the workload of medical assistants. The increased workload can, for example, result in the creation of lower-quality medical reports, meaning that subsequent patient treatment is based on incomplete information. The low quality of medical reports can be identified, for instance, by the fact that some aspects are only inadequately described or are completely missing.

[0009] The temporal classification and causal linking of medical events presents a further challenge. Recognizing connections between symptoms, diagnoses, treatments, and outcomes requires medical expertise and analytical thinking, which cannot always be optimally implemented under time pressure.

[0010] Furthermore, the quality of medical reports varies depending on the experience, available time, and individual skills of the person writing them. This can lead to variations in the completeness, precision, and comprehensibility of the documentation.

[0011] Transforming medical bullet points and structured data into a fluent, natural-language text presents an additional task that requires linguistic competence. Selecting relevant information and prioritizing it for a report tailored to the recipient is a complex process.

[0012] In medical practice, there is a tension between the need for comprehensive, precise documentation and the limited time resources of healthcare professionals. While increasing digitalization in healthcare has improved access to patient data, the efficient processing and consolidation of this information remains a significant challenge.

[0013] Therefore, an objective of the present invention is to provide a method for generating medical reports, in particular discharge summaries, which at least partially overcomes the aforementioned disadvantages of the prior art.

[0014] SUMMARY OF THE INVENTION

[0015] This is achieved by the subject matter defined in the independent claims. Advantageous modifications of embodiments of the present disclosure are defined in the dependent claims, as well as in the description and the figures. In a first aspect, the present invention relates to a computer-implemented method for generating medical reports, in particular physician letters.

[0016] The procedure may include the following step a): capturing medical information about a patient from at least two different sources of an electronic health record, wherein the medical information received is formulated at least partially in natural language bullet points.

[0017] The procedure can include the following step b): Identifying medical units contained in the recorded medical information and classifying each identified medical unit according to a medical event type using computer-aided pattern recognition. Accordingly, the information is analyzed using computer-aided pattern recognition to identify the medical units it contains. These units can then be classified according to their specific medical event type. This classification enables a precise assignment of the information, which can facilitate further handling of the data.

[0018] One component of the process can therefore be event detection and categorization, enabling the precise identification of medical events, diagnoses, and medication orders. Specialized technology can be employed to accurately determine and classify these elements. This facilitates the collection and categorization of information on medical conditions and treatment processes. The application of machine learning (ML) and large language modeling (LLM), such as through artificial intelligence, is an optional extension of this phase that can improve the precision and efficiency of the detection process. ML / LLM models offer advanced analytical capabilities that allow patterns to be identified in complex datasets.

[0019] The procedure can include the following step c): forming event groups by computer-aidedly grouping medical units that are related at least in terms of time, treatment pathway, or thematic relevance. This grouping helps ensure that related information is considered within a unified context.

[0020] This allows for the clustering of events and event flows. Relevant events can be grouped together based on their chronological sequence, their connection to the patient pathway, or their overall relevance. This grouping contributes to a structured presentation of information, enabling the consideration of related medical units within a coherent context and supporting clarity in the reporting process. This methodical grouping not only simplifies data presentation but also subsequent data processing and analysis. In particular, considering the time course and pathway of a patient's treatment makes it possible to understand and prepare for historical and future developments in medical care.

[0021] The procedure can include the following step d): deriving relationship information that represents at least one causal relationship or correlation between elements of the event groups, and storing the elements of the event groups and the relationship information in a structured data representation. This enables a clear and concise presentation of the collected information.

[0022] The procedure may include the following step e): Enriching at least one element of the structured data representation by linking it to datasets from at least one external medical knowledge base. These links help ensure that the information in the report is current and relevant.

[0023] Enriching the report with external knowledge bases enables a direct link between medical events and established medical knowledge bases such as UMLS (Unified Medical Language System) and SNOMED (Systematized Nomenclature of Medicine). This integration ensures that the information in the report is supplemented and supported by scientifically validated and up-to-date medical data. Using such databases enhances the report's professional validity.

[0024] The procedure may include the following step f): converting the enriched structured data representation into an input format that can be processed by a generative language model. This conversion enables the technical prerequisites for the subsequent processing.

[0025] The process may include the following step (g): processing the input format with the generative language model to generate natural language flow text that reflects the captured medical information for a medical report. This final phase ensures that the generated reports are precise and understandable, which is beneficial for communication between healthcare professionals. It may be stipulated that the sources according to step (a) include at least one source from the group consisting of clinical reports, laboratory results, imaging reports, medication prescriptions, diagnostic lists, procedure documentation, or progress reports.

[0026] The sources can therefore include a wide range of medical documentation to create a detailed data basis for report preparation:

[0027] For example, the information may include treatment histories and physiological assessments. Such reports often reflect direct observations and therapeutic steps taken during patient care and contribute to describing the patient's overall health status. The clinical reports in this layer provide insights into the course of the patient's medical care, including direct observations and therapeutic interventions. They are conducive to understanding the underlying health problems and serve as a primary source for collecting qualitative data about the patient.

[0028] Laboratory results represent another potential source of data for the procedure. These results provide valuable information about a patient's chemical and biological states, which can be beneficial for diagnosis and treatment adjustment. These results offer quantitative metrics that complement and expand upon clinical reports.

[0029] Imaging reports are also among the relevant sources that can be considered. They provide image-based information about physiological and pathological changes in the body. Evaluating such reports allows access to visual diagnostics that are helpful for many medical decisions. This visual data source helps to decipher complex situations and provides a basis for further treatment planning.

[0030] Another component is medication prescriptions. These contain information about prescribed medications, their dosage, and administration. They can be helpful for tracking therapeutic measures and adjusting treatment plans, helping to avoid medical errors and maximize treatment success. Diagnostic lists can also be used as a source of information. These lists contain the medical diagnoses documented for the patient. They allow for a targeted assessment of the patient's health status and support the prioritization of therapeutic measures. These lists are often the starting point for identifying treatment goals.

[0031] Procedure records collect information on medical interventions and other clinical procedures. Analyzing these records allows for detailed information about specific treatment steps and their chances of success.

[0032] Finally, the possible sources also include progress reports that document a patient's condition during their hospital stay and provide further information on the traceability of treatments and decisions.

[0033] Integrating these diverse sources provides a comprehensive overview of the patient, their health status, and progress. Advanced algorithms are used to analyze the collected data, classify events, and generate an informative report. This approach enables evidence-based documentation that supports medical decision-making.

[0034] It may be stipulated that, according to step a), medical information is collected from at least three different sources. This expanded data collection can provide deeper and more comprehensive insights into a patient's medical status. Utilizing diverse sources promotes the thoroughness of the collected medical information. It allows for the mapping of different aspects of treatment and diagnosis, taking into account the diversity of medical documentation. Such an approach can support the development of a holistic picture of the patient's health status, based on detailed and diverse datasets.

[0035] For example, information from clinical reports, laboratory results, and imaging reports can be captured simultaneously. Each of these sources provides specific information that, taken together, allows for a deeper analysis of the patient's status. Clinical reports offer direct insights into the patient's treatment history and therapeutic management, while laboratory results provide quantitative information about their biological condition. Furthermore, medication prescriptions can be included as a fourth source, offering a detailed overview of medication treatment and dosage. Diagnostic lists can provide additional structured insight into the patient's health challenges, while procedural documentation supplements this with comprehensive information about medical interventions performed.

[0036] Integrating information from these various sources allows for the creation of coherent data that supports the reporting process. This diversity can not only enable the accurate recording of a patient's health status but also assist in deciding on further treatment steps.

[0037] The systematic collection of data from a wide variety of sources makes it easier for the process to efficiently utilize the complexity of medical data and precisely feed the generative language model in order to create coherent and detailed medical reports. This promotes comprehensive communication among healthcare professionals and contributes to high-quality patient care.

[0038] It may be necessary for the data capture process according to step a) to include normalizing timestamps to a uniform time format. This normalization can help to accurately synchronize the captured medical information and present it within a uniform time format.

[0039] Timestamps, often provided by various sources within an electronic health record, frequently use different formats and time zones. Such inconsistencies can lead to difficulties in interpreting the data and may impair the comparability of events or processes. Therefore, normalizing the timestamps can contribute to a clear and coherent representation of the timeline, which is helpful for further processing.

[0040] Normalization allows all recorded timestamps to be displayed in a defined and uniform format. By standardizing this time information, a consistent chronology of medical events and treatments becomes possible, which is advantageous for analysis and the creation of clear medical reports. This normalization proves particularly useful when grouping events and deriving causal relationships or correlations between elements. It enables a precise temporal linking of the data and can contribute to generating meaningful relationship information.

[0041] Normalizing timestamps makes a patient's treatment pathway transparent and allows for the precise chronological sequence and prioritization of medical interventions. Structuring and synchronizing the time information contributes to clear and accurate medical reports, thereby promoting diagnostic and therapeutic efficiency. Normalization is particularly beneficial in situations where multiple healthcare professionals work with the data or where reports are used for shared decision-making, as it enhances clarity and unambiguity.

[0042] It may be envisaged that the pattern recognition according to step b) includes a Named Entity Recognition (NER) model. This model enables the precise identification and classification of specific medical entities within the captured data.

[0043] Named Entity Recognition (NER) is a natural language processing technique that aims to identify and categorize specific entities within a text, such as diagnoses, medications, treatment-related events, or other medically relevant information. Using an NER model allows for more accurate and optimized pattern recognition, as it facilitates the targeted extraction of relevant information from large datasets.

[0044] With a NER model, natural language processing can be optimized so that words or phrases can be identified as specific medical units. These units can be categorized within specific contexts such as patient conditions or treatment plans, which can be advantageous for reporting.

[0045] Using a NER model makes it possible to present the collected medical information in a more structured way. It supports appropriate classification by assigning medical entities to corresponding categories, which can be reflected in the organization and presentation of the data in the medical report.

[0046] This pattern recognition method provides an effective way to systematically capture key information to ensure that relevant medical data is included in reporting. The NER model enhances the accuracy of automated identification of relevant entities, thereby supporting further processing and analysis suitable for generating detailed and informative medical reports.

[0047] Implementing a named entity recognition model in the process enhances the system's processing capacity for natural language data and can help increase the relevance and accuracy of the generated reports. This can positively impact the quality of medical documentation and support clinical decision-making.

[0048] It is possible that the named entity recognition model is transformer-based. Transformer models, which are among the modern techniques of natural language processing, allow for detailed analysis and classification of text data.

[0049] Transformer-based models offer the ability to process information efficiently and accurately by employing a mechanism called self-awareness, which facilitates the understanding of the context of individual words within large datasets. This is helpful for identifying and categorizing specific medical entities from complex natural language structures.

[0050] A transformer-based NER model can be used to analyze large amounts of medical information and accurately identify relevant entities, even when they occur in variable contexts. This supports more accurate data capture and classification, which can be beneficial for further processing and report generation.

[0051] Transformer models have the ability to process linguistic context at a comprehensive level, which improves the quality and accuracy of the recognized and classified medical entities. The efficiency of these models makes it possible to represent complex relationships within the text, which can be useful for medical reports.

[0052] The use of a transformer-based NER model in the overall process can expand the possibilities for natural language processing and support a detailed and structured evaluation of medical data. This modern technique contributes to the precision and informative value of the generated reports by improving the process of extracting information and creating clear, coherent data structures.

[0053] Such a model offers the potential to optimize the collection and processing of medical entities, thereby promoting communication in healthcare through precise and informative reports. Advanced algorithms can filter and categorize highly complex datasets, enabling the effective use of relevant information to support medical decisions. This contributes to comprehensively supporting the treatment process and improving patient outcomes.

[0054] It may be possible to fine-tune the transformer-based named entity recognition model using a domain-specific medical text corpus. This fine-tuning can help to better adapt the model to the specific terminology and special terms used in the medical field.

[0055] Fine-tuning with a medical text corpus enables the model to recognize medical entities with high precision and relevance. When trained with specific medical texts, the model can reliably distinguish between subtle nuances of meaning that are important in the professional context.

[0056] A domain-specific medical text corpus comprises a selection of texts specifically compiled for the medical field. These texts include clinical reports, scientific articles, studies, guidelines, and other relevant medical documentation, encompassing a wide range of specialized terminology and concepts.

[0057] Fine-tuned models based on such a corpus are suitable for capturing subtle differences in terminological usage, thereby improving the performance of the NER model. The optimized recognition of medically relevant terms and the precise classification of this information help to reduce the risk of misinterpretations and ambiguities in reporting.

[0058] Through the application of a specialized training process, the Transformer model can be used in medical reporting to provide deeper insights and more accurate analyses of the collected medical information. The finished report can benefit from a high-quality representation of medical terminology and clinical processes. This process enables documented healthcare communication that is both precise and understandable, contributing to quality assurance in medical practice and patient care. This technology facilitates access to specialized information, which can support decision-making and streamline processes in clinical settings.

[0059] It may be stipulated that the medical event type according to step b) be assigned to at least one of the types diagnosis, medication administration, procedure, laboratory value, imaging finding, or progress report. By specifically assigning the identifiable medical units to these clearly defined event types, a structured and more precise analysis of the collected health information is promoted. This assignment simplifies the understanding of complex medical processes and provides a clearer presentation in the report.

[0060] Diagnoses represent specific health conditions or medical problems that can be identified from the collected data. Assigning a diagnosis to this type allows for a targeted presentation of the clinical pictures and systematic documentation of patient health.

[0061] Medication administration records include information on prescribed or administered medications and their dosages. Categorizing medication administration can facilitate informed tracking of medication therapy, which can be useful for updating and adjusting treatment plans.

[0062] Procedures encompass detailed information about medical interventions and clinical procedures. Classifying them as such allows for the chronological recording and evaluation of complex treatment processes and supports reliable documentation of surgical or therapeutic interventions.

[0063] Laboratory values ​​refer to specific diagnostic parameters resulting from laboratory analyses. These values ​​provide quantitative metrics for assessing a patient's biological status and support a clear presentation of test results in the report.

[0064] Imaging findings encompass information that can be obtained from diagnostic image data. Classifying findings as such facilitates visual analysis of medically relevant changes, which can be useful in diagnosis and treatment monitoring.

[0065] Progress reports comprise information collected by medical staff that depicts the inpatient course of a patient's illness. This clear categorization of medical units promotes efficiency in report generation by supporting a structured organization of information that can be reflected in the report. The use of specialized pattern recognition algorithms results in a more precise and informative representation of medical facts and provides a solid basis for informed clinical decisions.

[0066] It may be possible to form the event groups according to step c) using hierarchical clustering. This method enables efficient organization and summarization of medical information by involving a structured grouping of data in a multi-stage approach.

[0067] Hierarchical clustering is a technical approach that allows data to be represented in a hierarchically structured way, where similar or related medical units can be grouped based on criteria such as time reference, treatment pathway, or thematic relevance. This method is suitable for representing event groups from the most detailed to the most general levels, which can facilitate the visualization of connection patterns within medical information.

[0068] By applying a hierarchical clustering approach, a coherent picture of the patient's current and historical medical situation can be generated. This method makes it possible to visualize close and distant relationships within the data, thereby gaining detailed and comprehensive insights.

[0069] This type of clustering is particularly helpful for presenting the complex and diverse information from medical sources in a clear and organized manner. The hierarchical structure promotes an understanding of the connections and relationships between different medical entities and event groups.

[0070] Furthermore, hierarchical clustering supports the flexible classification of data, allowing both the highlighting of individual local details and the identification of overarching trends. The method is adaptable and can be specified with various criteria and parameters to extract beneficial insights from the data.

[0071] This clustering-based approach to forming event groups can be a valuable component in the systematic creation of medical reports. It brings structure and clarity to complex medical information and facilitates the identification and documentation of relevant relationships, thereby improving the quality and comprehensibility of the generated reports. This methodology supports the reliability of medical documentation and can strengthen clinical decision-making.

[0072] It may be possible to form the event groups according to step c) using density-based clustering. This technique offers the possibility of grouping data based on its spatial density, thereby enabling the efficient organization of related medical information.

[0073] Density-based clustering is able to identify groups of medical entities that are ordered and compact within a high density, while simultaneously filtering out low-density areas as separate clusters or as noise. This method can be helpful for analyzing complex datasets and identifying data patterns.

[0074] This clustering allows the formation of event groups based on the natural grouping of information, regardless of the data distribution pattern. Point density analysis ensures that the most interconnected data points are grouped together, supporting a more accurate and reliable understanding of the relevant information and its relationships.

[0075] Density-based clustering is suitable for identifying clusters with varying shapes and sizes, which can be advantageous in a medical context when addressing how different data categories and relationships can be grouped differently. It can support the analysis and visualization of data ranges that possess a certain relevance and frequency, in a way that reveals a coherent and meaningful grouping.

[0076] By employing a density-based clustering approach, insights from medical information can be systematically summarized, revealing both fine details and broader patterns. This method can contribute to a well-founded and clear presentation in the report by optimizing the structuring and documentation of event groups, which can then be integrated into the reporting process.

[0077] In practice, this clustering approach can help improve the quality and efficiency of healthcare reporting by enabling the organization of complex data in a simple and accessible way, facilitating the understanding of interrelational information and the derivation of medical insights. This methodological approach supports accurate documentation and can lead to better clinical decisions and consistent patient care.

[0078] It may be stipulated that when forming event groups, a maximum permissible time interval between medical units to be included in the same group is specified. This temporal restriction ensures precise organization and structuring of medical information by indicating that related medical units can be grouped within a defined timeframe.

[0079] By using a maximum time interval, it is possible to specifically capture medical events that occur consecutively or in close association. This approach is helpful in illustrating the chronology and relevance of treatment events and changes in condition, thus improving the accuracy and coherence of the medical report.

[0080] Establishing a time interval ensures that related medical units do not have excessively large gaps in time, and the event groups can thus provide a consistent representation of the treatment sequence. This structuring allows for the precise recording of temporal progress and highlights significant moments in the treatment process.

[0081] Methodically considering the temporal interval when grouping units allows medical events to be integrated into their natural sequence, making it possible to derive relationship patterns and causal relationships. This dynamic classification supports both the medical analysis and the interpretation of the diagnostic recommendations in the report.

[0082] By strategically applying this temporal constraint when forming event groups, a comprehensive presentation of medical information is enabled, which can positively influence the quality of the report. The structured and timely organization of the data aims to provide comprehensible medical documentation that can facilitate clinical decision-making and contribute to an improved patient experience. This approach increases the efficiency of reporting and ensures that the chronological sequence of medical events and their significance for treatment are clearly depicted. It supports the strategic planning of treatment measures and the precise monitoring of treatment progress in healthcare.

[0083] It may be possible to derive the relationship information according to step d) using a structural causal model. This provides deeper insights into the direct and indirect effects of medical interventions and events. This technique makes it possible to understand and present cause-and-effect relationships more clearly.

[0084] The use of a structural causal model can support the identification of causal relationships and their influence on treatment outcomes. This focuses on the analysis of dependencies between groups of events, enabling the creation of medical reports that reflect the context of information in clinical scenarios.

[0085] These models simplify the analysis of medical processes by providing a graphical and mathematical representation of causal relationships. They enable a deep understanding of the interaction and interplay of medical units, which serves as a basis for documentation and can improve clinical decision-making.

[0086] Integrating a structural causal model into the process makes relationship information and its influence on patient care and prognosis more readily apparent. It offers a way to support the adaptation and optimization of treatment processes by strengthening and accurately representing relevant connections between data.

[0087] Such an approach can help increase the quality and informative value of medical reports by providing data-driven insights into health trajectories and treatment outcomes. The precise mapping of causal relationships enables a comprehensive and evidence-based perspective that can serve as a foundation for clinical decisions and an improved patient experience.

[0088] The structural causal model can be provided as a directed acyclic graph. This representation offers a clear depiction of the causal relationships between medical units and event groups. A directed acyclic graph (DAG) provides a structured way to map the direction and hierarchy of causality between different elements. In a DAG, connections run along directed edges from one node to another, excluding cycles. This prevents an element from being reverted to a previous state, thus promoting a clear structure.

[0089] Using a DAG can contribute to the transparency and clarity of causal relationships. Each node in the graph represents a medical entity or a group of events, and the directed edges indicate the direct influences or effects between the entities. This graph representation simplifies the understanding of the sequence and interactions of medical events.

[0090] Such a graph is suitable for clearly illustrating the relationships between treatment processes, diagnoses, medical interventions, and other relevant criteria. It opens up the possibility of a detailed analysis and representation of correlations that are relevant in the context of patient care.

[0091] Providing a DAG (Data Access Group) offers the opportunity to place reporting on a visual basis, supporting accurate findings and decision-making processes. This can make documentation more understandable and contribute to supporting clinical decisions.

[0092] A DAG (Dynamic Assessment Group) as a model for causal relationships is well-suited for structured reporting. It can enable medical information to be carefully processed to provide structured insights that optimize the treatment process. This method can contribute to the quality of medical documentation and support the promotion of transparent communication in healthcare.

[0093] It is possible to store the structured data representation according to step d) as a knowledge graph. This storage option offers an advantageous way to represent the collected and analyzed medical information in a linked and semantically rich format.

[0094] A knowledge graph is a model that makes the relationships and attributes of entities within a specific context representable and linkable. Using a knowledge graph in the medical field can enable an organized and efficient representation of the relationships between different medical units and event groups.

[0095] By storing structured data representations as knowledge graphs, the connections between medical information can be depicted not only linearly, but also in a way that considers extensive relationships and offers broader insights. This can contribute to a better understanding of treatment processes and patient histories.

[0096] A knowledge graph is suitable for visually and logically representing detailed and complex relationships between medical data. The nodes represent medical entities, and the edges provide the means to depict the relationships and interactions between them. Such representations enable efficient navigation and querying of the data, thus supporting the creation of precise and informative reports.

[0097] Using a knowledge graph allows medical information to be flexibly updated and expanded, thus promoting the development of a dynamic and adaptable reporting process. This method of storing structured data represents a comprehensive representation that improves both the quality and accessibility of the generated report.

[0098] A knowledge graph, as a storage format for structured data representation, offers a deep and interconnected view that fosters clear and evidence-based insights into clinical decision-making and optimizes medical documentation in an advanced way. This supports diverse clinical scenarios and contributes to improved patient care by providing a narrative and comprehensive overview of medical information.

[0099] One exemplary embodiment includes the implementation of a mechanism for linking causal relationships, which aims to make causal relationships within medical data identifiable and structurable. This embodiment demonstrates how a knowledge graph can be used specifically for the creation of discharge summaries by visualizing causal connections between different clinical factors and events.

[0100] The mechanism for establishing causal relationships serves as a functional module within the computer-implemented process for optimizing medical reporting. It creates a detailed and dynamic representation of causal interactions by systematically capturing and modeling cause-and-effect relationships. This allows dependencies and influencing factors between medical entities to be clearly represented and integrated into the knowledge graph, thus providing sound clinical decision support.

[0101] The introduction and integration of a specially developed knowledge graph for creating medical reports provides an advanced method that goes beyond purely informative presentation and enables expanded, evidence-based insights. The knowledge graph functions as a semantic network, allowing the nodes and edges to be structurally organized in such a way that causal relationships are not only identifiable but also comprehensible within a narrative context.

[0102] This implementation extends the functionality of the knowledge graph by supporting the recognition and representation of causal relationships in medical documentation. This can lead to a significant improvement in the quality and precision of discharge summaries by not only making existing information linkable but also revealing new and relevant connections that contribute to a more comprehensive understanding of the medical situation.

[0103] In summary, this exemplary implementation offers a complete and specialized use of the knowledge graph to represent causal relationships in medical reports. This significantly increases the integrity and value of clinical information for decision-making and improves the efficiency of the reporting process, ultimately leading to better medical care and patient support.

[0104] It may be possible to persist the knowledge graph as RDF triples. RDF (Resource Description Framework) is a standard model for describing information on the web and offers a flexible way to represent structured data.

[0105] Persisting the knowledge graph as RDF triples enables a clear and formalized representation of medical information in the form of subject-predicate-object statements. These triples can contribute to the specification potential of semantic web applications and offer a standardized method for storing and querying interconnected data. This representation enhances the interoperability of medical data, as RDF triples allow structured and linked information to be efficiently exchanged between different systems and applications. This openness and flexibility contribute to the meaningful use of the captured medical information in a broader context.

[0106] RDF triples enable a high degree of data accuracy and clarity, which can be particularly beneficial for creating precise and informative reports. The triple-based structure not only captures data but also visualizes the relationships and interactions between them, which can positively influence the quality of the reporting.

[0107] Persisting the knowledge graph as an RDF triple allows for the integration of data into relational or non-relational databases, as well as graph-based storage systems. This facilitates efficient data processing and querying and can be characterized by flexible scalability suitable for dynamic reporting requirements.

[0108] The use of RDF triples helps increase the value and relevance of medical information by creating a structured and well-linked data landscape that can strengthen the evidence and information base for clinical decisions. This not only supports the efficiency of the reporting process but also improves access to qualified medical insights in healthcare.

[0109] It may be stipulated that the external medical knowledge base according to step e) is the Unified Medical Language System (UMLS). UMLS comprises a comprehensive collection of medical terms and their relationships to standardize and integrate the terminology of diverse health data.

[0110] Using the UMLS as an external knowledge base can support the linking of recorded medical units with an established network of terms and their semantic meanings. This extension can increase the information content and accuracy of reports by providing clinically relevant and standardized information.

[0111] The UMLS provides resources such as thesauri, lexical word forms, and metathesaurus, which enable the effective categorization of medical information and the clarification of its meanings. Integrating this knowledge base into the process can help enrich the collected data with additional references and definitions, thereby improving the comprehensibility and validity of the report.

[0112] By using UMLS, medical reports can be formatted with validated terms that are recognized within the scientific and clinical community. This standardized format can facilitate communication among healthcare professionals and support the reliability of reporting.

[0113] Furthermore, the use of UMLS offers the possibility of methodically presenting complex medical information, which can improve diagnostic support and positively influence decision-making processes. The provision of metadata from UMLS can promote the creation of informative and consistent reports that are beneficial in clinical settings and contribute to the quality of patient care.

[0114] The integration of UMLS into the process can thus contribute to more comprehensive and efficient report generation by providing a network of domain-specific definitions and semantic references that optimizes the expressiveness and precise communication of medical information.

[0115] It may be stipulated that at least one external medical knowledge base, as per step e), is SNOMED CT, ICD, ICD10GM, ICD11GM and / or OPS. Combinations are also possible. SNOMED CT, ICD, ICD10GM, ICD11GM and OPS, as a comprehensive clinical terminology, comprise detailed sets of medical terms used worldwide to standardize and clarify medical documentation.

[0116] The integration of SNOMED CT, ICD, ICD10GM, ICD11GM, OPS, or various combinations thereof, is suitable for linking the recorded medical units with a globally recognized and established system of clinical terminology. This enhancement can improve the depth and accuracy of reporting by enabling the presentation of medical information in a clear and standardized context.

[0117] SNOMED CT, ICD, ICD-10GM, ICD-11GM, and OPS offer a hierarchical structure of terms and concepts capable of codifying and categorizing complex medical information. Utilizing this knowledge base within the process supports the application of precise clinical terminology, which contributes to improved comprehensibility and communication in the report.

[0118] By linking collected data with SNOMED CT, ICD, ICD10GM, ICD11GM, and OPS, medical reports can be enriched with validated and internationally recognized terminology. This approach promotes the quality and consistency of medical documentation and can facilitate communication between healthcare facilities and international professionals.

[0119] The application of SNOMED CT, ICD, ICD10GM, ICD11GM, and OPS in reporting provides a robust and detailed method for structuring clinical data, which can support both decision-making in clinical settings and research and continuing education processes. The precise coding of medical information enables the creation of more in-depth and evidence-based reports that are beneficial in healthcare and can optimize the flow of information.

[0120] The use of SNOMED CT, ICD, ICD10GM, ICD11GM, and OPS in the reporting process facilitates the creation of comprehensive and understandable reports by raising the diversity of language and meaning to a globally standardized level and simplifying the sharing and discussion of clinical information in an international context. This enhancement of the process can contribute to the precision and usefulness of the reporting process and expand its efficiency and potential in the medical field.

[0121] It may be possible to perform the linking according to step e) using a vector similarity measure based on embedded representations. This method offers an advantageous way to analyze semantic relationships between medical entities through machine learning and natural language processing.

[0122] The vector similarity measure is based on the idea that words and concepts are represented in a high-dimensional space, and their semantic relationships can be identified through mathematical calculations between their corresponding vector representations. This technique utilizes deep neural networks and embedded representations to enable an understanding of the context and meaning of medical terms. By applying a vector similarity measure, captured medical entities can be precisely enriched with relevant datasets from external knowledge bases. This method allows for the identification of hidden relationships and semantic similarities that might be missed by simple text-based approaches.

[0123] Embedding representations supports complex semantic comparisons and allows medical terms to be linked within the context of their meaning and relationships, despite varying terminology. This approach enriches structured data and can significantly improve the quality and informative value of the generated reports.

[0124] For example, medical units such as diagnoses, treatment options, or outcomes can be enriched with relevant knowledge from established systems like UMLS.SNOMED CT, ICD, ICD10GM, ICD11GM, or OPS, where semantic similarity can serve as the basis for linking. This leads to a comprehensive and meaningful presentation of information in the medical report.

[0125] The use of a vector similarity measure supports the harmonious linking of medical information with the terms and standards in knowledge bases, thus facilitating the provision of informative and reliable reports. This advanced technique offers the potential to improve accuracy and efficiency in the medico-informatics documentation process and can enable deeper, evidence-based clinical insights that can enhance patient safety and quality of care.

[0126] It is possible for the input format according to step f) to be structured as a JSON document. This option allows for efficient and flexible organization of the structured data for further processing.

[0127] JSON, JavaScript Object Notation, is a lightweight data format that supports readability for both humans and machines. Its straightforward structuring and semantic clarity make it ideal for representing complex data. JSON provides a simple syntax for describing hierarchical data and enables a structured presentation of information.

[0128] The JSON format allows medical information to be presented comprehensively and transparently in the input format by organizing it as coherent structures of key-value pairs. This clear structure is advantageous for advanced data processing and facilitates integration into systems that perform further analyses or utilize the generative language model.

[0129] A JSON document offers the possibility of efficiently representing nested and hierarchical information, which can be helpful for implementing medical event groups and their relationships. The various elements from the captured data are arranged clearly and logically, which can support coherence in further processing.

[0130] This form of representation simplifies both the import and export of structured data and improves the interoperability of the reporting process with various applications and technologies. JSON documents are suitable for use in real-time applications or for data exchange between different services.

[0131] Using a JSON document as an input format can support flexible, scalable, and compatible medical data processing with many modern IT systems. It provides a solid foundation for creating informative, accurate, and understandable medical reports that can improve the quality and efficiency of medical documentation and support the coordination of clinical workflows.

[0132] The JSON document may include a prompt template with a role description, instructions, and context. This structure allows for specific guidance and contextualization of the data for use in the generative language model.

[0133] The prompt template can be designed to describe the relevant roles, including defining which tasks and perspectives can be included in the generative language model. This approach helps to precisely specify the target output and ensures that the generated communication meets the required criteria.

[0134] The instruction section of the prompt template includes detailed instructions on how to interpret and process the data from the JSON document. This promotes a focused and consistent processing approach, which can be geared towards generating a precise and coherent linguistic representation of medical information.

[0135] The context section of the prompt template provides a framework for interpreting the data within the generative language model. This section ensures that the generated language is not only based on the information itself, but also reflects its meaning and relevance within the medical context.

[0136] Such a structured JSON document is suitable for supporting precise data flow management and facilitating the creation of reports that can be used by both experienced clinical professionals and in patient communication. The use of a clear and defined template creates a foundation for implementing a standardized and reproducible reporting process in the medical setting. This contributes to the consistency and comprehensibility of medical documentation and can positively influence the quality and efficiency of healthcare systems and patient care.

[0137] It can be envisaged that the generative language model according to step g) is an autoregressive transformer, e.g., with at least one billion model parameters. This advanced model architecture enables complex and nuanced natural language processing and is suitable for generating medical information precisely and coherently.

[0138] An autoregressive transformer uses its architectural structure to sequentially predict elements in the input format through self-awareness, enabling it to understand long context dependencies and subtle linguistic structures. Extensive parameterization helps the model capture fine distinctions within medical information and generate detailed, multi-layered text outputs.

[0139] The large number of model parameters offers enhanced capacity for in-depth pattern recognition and semantic analysis, thus facilitating the creation of structured and content-rich medical reports. Thanks to its broad information scope, the generative language model can handle both the complexity and the context of the captured medical entities.

[0140] Such a parametric model supports report generation by providing multidimensional and contextually relevant language patterns, which can improve the quality of reporting and increase information accuracy. The use of this model helps ensure that the generated reports are technically precise and that communication in the medical field can be improved.

[0141] This powerful model architecture enables reliable and comprehensive accountability in the generation of medical texts by offering seamless integration and adaptation to the input format. Even complex or specialized terms can be presented in a meaningful and understandable flow of text, increasing the transparency and value of medical reports.

[0142] An autoregressive transformer with this parameterization represents an effective option for generating high-quality reports and can address diverse healthcare requirements, thereby significantly improving the accuracy, efficiency, and clarity of communication. This structure offers potential for long-term benefits in medical documentation and clinical decision support, contributing to the optimization of patient outcomes and the promotion of interdisciplinary collaboration.

[0143] It may be intended that the generative language model is fine-tuned with regard to medical summaries through Reinforcement Learning from Human Feedback.

[0144] It may be possible to specify a maximum token length, e.g., 1,024 tokens, for the output during step g). This approach allows the model to be adapted to specific requirements and can increase the quality and relevance of the generated reports.

[0145] Reinforcement Learning from Human Feedback is a method that trains the model through reward systems and feedback from human experts. It allows the model to better understand specific patterns and writing styles, which can enable more targeted generation of medical summaries.

[0146] Using RLHF, the generative language model can make optional decisions for results relevant within a specific domain, while simultaneously avoiding erroneous or inefficient outcomes. Human feedback supports continuous improvement of the model and allows for precise fine-tuning of the output process for the natural language texts.

[0147] The approach of creating optimizations through human feedback enables the model to process complex medical information precisely and understandably, and to present it in a useful format. Particularly in terms of technical explicitness and clarity, the process indicates an increase in quality while simultaneously improving comprehensibility.

[0148] Through fine-tuning using RLHF, with particular attention to the generation of medical summaries, the model can learn to communicate specific medical context conditions, terminology, and meanings correctly and effectively. This improves accuracy and can increase the applicability of reporting and enhance the relevance of documentation in the clinical setting.

[0149] Integrating RLHF into model training helps to expand the adaptability of the generative language model and to meet human expectations as well as the demands of healthcare. This can help improve reporting, minimize errors, avoid misunderstandings, and promote the quality and clarity of medical communication.

[0150] This approach enables the generative language model to meet the increasing demands and new challenges in the medical field and to systematically optimize documentation processes. Reporting is supported to improve patient care and facilitate standardized communication in healthcare.

[0151] It may be possible to define a maximum token length, which helps to keep the generated reports within a manageable scope, positively impacting both the processing speed of the language model and the readability of the reports. Tokens are basic units of text length and include words and punctuation marks.

[0152] This limit structures the reports so that they are comprehensive enough to convey relevant medical information while remaining precise and focused. This framework prevents reports from becoming overly long or unwieldy and facilitates the rapid acquisition of core information.

[0153] Setting a maximum token length also offers advantages in resource control of the generative language model. It limits the size of the output and contributes to more efficient use of the model's internal resources, thereby optimizing the entire creation and processing process.

[0154] Applying this limit improves the creation of medical reports by ensuring that communication remains clear, concise, and direct without omitting essential medical content. This approach promotes the comprehensibility and accessibility of reports, enabling healthcare professionals to quickly and effectively access relevant information. With a defined maximum token length, this method supports both the consistency and quality of generated medical texts, as it encourages a balanced presentation of relevant information and ensures that reporting meets the specific requirements of the healthcare sector. This facilitates reporting and focuses documentation on significant clinical aspects that contribute to patient care and decision support.

[0155] It may be possible to automatically write the generated medical report back into the patient's electronic health record. This functionality can facilitate the seamless integration of new medical information into the existing record and support the consistency and completeness of patient-related documentation.

[0156] The automatic write-back of the report facilitates access to all relevant data, which could improve efficiency in managing the health record. Centralized storage of the reports also improves accessibility, allowing healthcare professionals to have a comprehensive overview of the patient's medical history at any time.

[0157] Integrating the generated report into the electronic health record offers the opportunity to promptly record and update the patient's progress and changes in their health status. This can improve the accuracy of documentation and support treatment decision-making by maintaining the continuity and timeliness of medical information.

[0158] The automatic integration of medical reports into the existing system aims to create an efficient workflow designed to improve information processing and minimize administrative tasks. This results in optimized patient documentation that is less prone to manual errors and promotes consistency in the health record.

[0159] Focusing on the automatic integration of the report into the electronic health record can further enhance the quality of medical documentation by providing comprehensive and well-structured data that can positively impact both treatment outcomes and organizational processes. In the long term, this functionality could contribute to the efficiency of healthcare systems and support high-quality patient care as well as standardized communication among healthcare professionals. The process may also include providing a user interface through which healthcare professionals can edit and approve the generated report. This interface provides a platform that can support the adaptation and validation of reports before their integration into the electronic health record.

[0160] By providing a user interface, professionals have access to report editing options, allowing them to review the content and make adjustments as needed. This opportunity for control helps optimize the accuracy and relevance of the medical information in the report.

[0161] The interface may include tools for commenting and correction, thereby promoting interactive collaboration with the generated documents. Medical professionals can tailor the reports to reflect specific patient care needs and preferences and adequately represent the clinical context.

[0162] Furthermore, the user interface potentially supports the release of edited reports, allowing them to be integrated into the electronic health record and made accessible to other relevant parties in the healthcare system. The release process can include security and quality controls that can provide evidence of the report's accuracy and authenticity.

[0163] The functionality of the user interface can facilitate collaboration and the

[0164] Supporting information exchange between healthcare professionals by creating a platform for shared decision-making and communication. These features promote the patient record as a central, up-to-date, and reliable repository of information that supports comprehensive healthcare.

[0165] Providing this interface simplifies the process by facilitating the integration of technology and professional input, making reporting more effective and flexible. It positively impacts the quality of documentation and supports the coordinated use and adaptability of healthcare, which could ultimately have a beneficial effect on care outcomes and the patient experience.

[0166] It may be stipulated that all patient-identifying information be pseudonymized or anonymized before data is transmitted to the generative language model. This approach can support patient data protection and promote compliance with data protection regulations.

[0167] Pseudonymization or anonymization involves concealing or removing identifiable characteristics from medical information, thus protecting patient privacy and preventing direct identification. This practice can be beneficial for both data security and acceptance of the procedure in clinical settings.

[0168] Regular application of such data protection measures makes it possible to achieve a high level of data security when processing data with the generative language model. The risks of unauthorized data use or data loss are reduced, thus serving as a measure to protect sensitive health data.

[0169] Furthermore, anonymization or pseudonymization ensures that the generated reports are based solely on medical information, without the unintentional disclosure of personal data. This separation between content and identity maximizes confidentiality in reporting.

[0170] This security feature can help ensure that the process complies with national and international data protection regulations, such as the GDPR in Europe. Adherence to these standards fosters a trustworthy foundation for the management, processing, and storage of medical information.

[0171] Implementing data anonymization or pseudonymization enables data protection-compliant reporting without compromising the quality or integrity of medical information. This security measure promotes more efficient and secure medical documentation and communication processes, thereby improving reliability in healthcare.

[0172] It may be planned that all calculation steps are carried out in a local IT infrastructure operated within the healthcare facility.

[0173] It may be possible to execute step g) in a hardware-based security environment with protected memory. This configuration facilitates keeping sensitive medical data under the direct control of the institution while supporting recommendations for enhanced security and data protection. Local deployment of the computational steps minimizes access to external networks, which could reduce the risk of unintended data leaks or cyberattacks. This implementation can help preserve the integrity and confidentiality of medical information.

[0174] Furthermore, processing medical data locally improves the efficiency and responsiveness of data processing, as there are no delays caused by external communication channels. This enables fast and reliable report generation and supports time-critical work in clinical settings.

[0175] Furthermore, switching to a locally operated IT infrastructure facilitates compliance with specific institutional or national data protection guidelines by providing direct control over how and where medical information is processed and stored.

[0176] By processing all steps within the facility, complete control over the hardware and software is maintained, thus promoting adaptation and optimization to the individual needs of the healthcare institution. This creates the opportunity to continuously improve the infrastructure and respond to new challenges in medical documentation.

[0177] This local implementation serves both to protect sensitive data and to provide a central hub for information-related activities in healthcare, thus supporting the quality and efficiency of reporting processes. Direct access to and management of reporting within the institution promotes the effectiveness and integrity of medical documentation and facilitates the smooth flow of information in healthcare.

[0178] It may be stipulated that step g) is executed in a hardware-based security environment with protected memory. This security measure helps to ensure that touch-sensitive medical data processing can be carried out securely and confidentially.

[0179] A hardware-based security environment is suitable for making processes of the generative language model feasible in a controlled environment where physical and virtual security mechanisms are in place. Protected memory provides an additional layer of security, as it makes manipulation and unauthorized access less likely.

[0180] Performing processing steps in such an environment makes it possible not only to protect critical data from external threats but also to maintain the internal integrity of the processes. This method helps ensure that sensitive medical information is securely stored and processed, and that the risks of data loss or misuse are mitigated.

[0181] By isolating the calculation steps in protected memory, it is possible to make the key information accessible to the intended process. These security mechanisms offer the potential to increase the confidentiality and reliability of the generated reports by ensuring that the process outputs are dependable and secure.

[0182] The application of this security environment facilitates compliance with relevant data protection regulations and supports the production of trustworthy reports. Such an infrastructure can strengthen the foundation for secure medical communication and the quality of the reporting process in the medical field.

[0183] Implementing a hardware-based security environment during step g) enhances the security of reporting without significantly impacting the efficiency or quality of the generated medical reports. This security measure offers valuable benefits in supporting the responsible and effective management of medical information and the overall quality of patient care.

[0184] In another aspect, the invention relates to a computer-implemented method for generating medical reports, in particular medical letters. The method can include a step of receiving medical information about a patient. The received medical information can be formulated as natural language bullet points. The method can then include a step of generating natural language flow text, reflecting the received medical information, for a medical report using a trained generative language model. This enables the automatic generation of high-quality medical reports using the generative language model.The fact that the system can be generated from natural language bullet points reduces the associated additional workload for medical staff, as they only need to provide the bullet points in natural language. These can be provided, for example, via a text field in an app or via a speech recognition interface (e.g., a smartphone, tablet, smartwatch, dictation device, etc.). This provides a system that addresses the growing need for efficient, standardized, and timely medical documentation in healthcare facilities by significantly reducing the administrative burden for medical professionals.

[0185] The further aspects of the invention described below can be combined with all the methods and further aspects disclosed herein.

[0186] According to another aspect, the generative language model can be configured so that the generated natural language flow text reflects all and only the received medical information.

[0187] This prevents the model from generating medical information that is not included in the input. It also prevents the model from omitting medical information during generation. This ensures that all relevant medical information is part of the generated report, and that no incorrect information is included. This guarantees subsequent patient treatment based on complete information. As a result, the procedure can be approved as a medical device.

[0188] According to another aspect, the generative language model can be configured so that the natural language flowing text deterministically reflects the received information.

[0189] Since in most cases the medical report is reviewed by a medical professional after generation, it is helpful if the model reflects the received information deterministically. If the model did not operate deterministically, the generated text could differ each time it was created, requiring medical staff to review the entire report every time. However, if the model is deterministic, all parts of the report only need to be reviewed once (i.e., when new medical information is added, only the corresponding newly generated section of the medical report needs to be reviewed).

[0190] According to another aspect, the process of generating natural language text can include: providing all information contained in the received medical information, collected and / or in one step, to the generative language model in order to generate the natural language flow. This technique, which can also be described as "one-shot prompting," allows for particularly efficient generation of the resulting natural language flow, since the generative language model receives all the information in a single input prompt and therefore only needs to process a single input prompt to generate the natural language flow.

[0191] As an alternative to the above approach, the process of generating natural language text can include: individually providing the information contained in the received medical information to the generative language model to generate an intermediate result; combining the intermediate results into a final intermediate result; and providing the final intermediate result to the generative language model to generate the natural language flow. The intermediate results can be generated as natural language flow. The final intermediate result can be generated by concatenating the intermediate results. This technique, also known as "Chain of Summary" or "Chain of Density Prompting," allows the resulting natural language flow to be more detailed in some application scenarios without being too dense or difficult to understand.

[0192] Furthermore, the procedure or a program executing the procedure may support both types of natural language text generation described above. For example, the generation method to be selected for an individual input containing medical information may be configurable by the user.

[0193] According to another aspect, the step of receiving medical information can include receiving at least one user input via a graphical user interface, particularly from medical professionals during a patient visit. Additionally or alternatively, the step of receiving medical information can include receiving at least one electronic data record from a hospital information system. This electronic data record can be Health Level 7 compliant and / or Fast Healthcare Interoperability Resources compliant. Additionally or alternatively, the process can include generating the medical report as an electronic file containing the generated natural language text.

[0194] This provides a fully digital workflow for generating medical reports, which can be easily integrated into existing hospital information systems. The fully digital workflow improves the quality of medical reports (e.g., eliminating copying errors) while simultaneously reducing the effort associated with their generation.

[0195] According to another aspect, the medical information received can include at least one piece of medical information belonging to one of the following categories: gender, leading symptom, diagnoses, treatments, medications, findings, and progress documentation.

[0196] This medical information can be used to create a complete medical record.

[0197] According to another aspect, the procedure can include a step of categorizing the received medical information. This categorization can be carried out using a classifier, in particular a statistical classifier.

[0198] According to another aspect, the procedure can include a step of generating an input encoding for the generative language model based on the received medical information. The input encoding can encode the medical information in such a way as to take into account a maximum input length for the generative language model.

[0199] This provides a compact encoding of medical information, allowing, for example, consideration of the maximum input length of the generative language model. Furthermore, the compact encoding reduces the amount of data to be processed.

[0200] According to another aspect, the input coding can be generated in such a way that each piece of medical information in the input coding is assigned to a predefined category abbreviation, which indicates the respective category.

[0201] This provides a particularly compact coding system, which, thanks to the category abbreviations, does not result in any loss of information. Furthermore, the input coding can be generated in such a way that it contains multiple pieces of medical information from the same category for at least one of the categories. The category abbreviation can only appear once in the input coding. The input coding can also be generated so that the multiple pieces of medical information from the same category are separated by a predefined separator abbreviation.

[0202] This provides a particularly compact coding system. Because a single category can contain multiple pieces of medical information, the input coding allows for the mapping of a complete and comprehensive treatment process. The separator abbreviation ensures that the individual pieces of medical information remain identifiable and that no mixing of medical information occurs. Otherwise, this could lead to a medical report containing incorrect medical information.

[0203] According to another aspect, the input encoding can be generated in such a way that at least one piece of medical information in the input encoding is assigned a timestamp, in particular after the medical information and separated by a predefined separator abbreviation.

[0204] This provides a particularly compact encoding that can accurately represent a temporal progression.

[0205] According to another aspect, the category abbreviation and / or the separator abbreviation can contain exactly one character.

[0206] This provides a particularly compact coding system, which nevertheless enables the unambiguous assignment of medical information to categories as well as the unambiguous differentiation of multiple pieces of medical information from the same category.

[0207] In a further aspect, the invention relates to a generative language model for generating medical reports, in particular medical letters. The generative language model can be configured to generate natural language flow text for a medical report. The generated natural language flow text can reflect medical information about a patient. The medical information can be formulated in natural language bullet points. The generative language model can be configured for use in the method according to any of the preceding aspects. In a further aspect, the invention relates to an input encoding data structure for the generative language model according to the second aspect. The input encoding can be configured according to any of the preceding aspects.

[0208] In a further aspect, the invention relates to a method for training the generative language model according to the second aspect. The method may include a step of providing training datasets. A training dataset may comprise an input encoding data structure according to the third aspect. A training dataset may comprise natural language flow text for a medical report. The method may include a step of training the generative language model by supervised learning with the training datasets.

[0209] In another aspect, the invention relates to a data processing device comprising means for carrying out the methods according to any of the mentioned aspects.

[0210] In another aspect, the invention relates to a computer program or a computer-readable medium on which a computer program is stored, wherein the computer program includes instructions which, when the computer program is executed by a computer, cause the computer to execute the method according to any of the mentioned aspects.

[0211] The following terms are used solely for the purpose of understanding the invention.

[0212] Unless otherwise specified herein, the term "computer-implemented method" refers to a method that is executed via one or more computer systems or processors, with each step of the method relying on computational operations. Examples of specific embodiments include software-based workflows that run on a local server, cloud-based applications that perform the steps remotely, or hybrid systems where data is processed both locally and in the cloud.

[0213] Unless otherwise specified herein, the term "generating medical reports, in particular medical letters" refers to the automated or semi-automated creation of structured or unstructured text documents summarizing patient-related medical data, with particular emphasis on letters commonly exchanged between healthcare providers. Examples of specific implementations include the digital creation of discharge summaries, referral letters, or surgical reports. Unless otherwise specified herein, the term "medical information" refers to all patient-related health data, including but not limited to diagnoses, laboratory results, imaging reports, medications, and treatment plans. Examples of specific implementations include ICD-coded diagnoses, blood test results, MRI findings, medication prescriptions, or therapy protocols.

[0214] Unless otherwise specified, the term "patient" in this document refers to a person who receives or is registered for healthcare services and whose data is recorded, processed, and used to generate medical reports. Examples of specific instances include inpatients, outpatients, or individuals participating in a clinical trial.

[0215] Unless otherwise specified, the term "electronic health record" refers to a digital repository or platform where patient health data is stored, accessed, and managed electronically. Examples of specific implementations include electronic patient records in hospitals, regional health information exchange systems, or patient-controlled personal health records.

[0216] Unless otherwise specified herein, the term "natural language bullet points" refers to concise bullet points formulated in everyday language that summarize medical details or observations rather than presenting them in full sentences. Examples of specific examples include physician notes typed in short sentences (e.g., "BP 120 / 80, stable"), bullet point lists of diagnoses, or brief text references to laboratory findings without a complete grammatical structure.

[0217] Unless otherwise specified herein, the term "medical units" refers to individual medical contents recognized within the collected medical information, such as specific diagnoses, procedures, laboratory values, references to body parts, or other identifiable clinical concepts. Examples of specific embodiments include recognized terms such as "diabetes mellitus," "appendectomy," "hemoglobin 12 g / dL," or "left ventricle."

[0218] Unless otherwise specified herein, the term "medical event type" refers to a classification category assigned to each recognized medical entity, indicating its general nature or role in the clinical context. Examples of specific embodiments include categories such as "diagnosis," "drug administration," "procedure," "laboratory value," or "image finding." Unless otherwise specified herein, the term "computer-aided pattern recognition" refers to automated processes or algorithms executed on one or more computer systems to identify patterns or specific entities in text or structured data. Examples of specific embodiments include machine learning workflows, rule-based entity extraction systems, or transformer-based neural network models.

[0219] Unless otherwise specified, the term "event groups" in this document refers to collections of medical entities that are related to one another within a patient record based on time frame, type of treatment, or thematic relevance. Examples of specific implementations include grouping laboratory values ​​measured on the same day, medications prescribed for the same condition, or diagnoses associated with a particular treatment pathway.

[0220] Unless otherwise specified, the term "time reference" in this document refers to the temporal dimension or timestamp associated with medical events that allows for a relevant grouping or sequencing of those events. Examples of specific embodiments include date-related contexts such as "admission date," "date of symptom onset," "date of laboratory testing," or approximate timeframes for the course of the illness.

[0221] Unless otherwise specified, the term "treatment pathway" here refers to a sequence or course of clinical interventions, assessments, or procedures that constitute a coherent course of treatment for a patient. Examples of specific implementations include a chemotherapy schedule, a physical therapy plan, or an orthopedic treatment pathway from pre- to post-operative care.

[0222] Unless otherwise specified, the term "thematic relevance" here refers to the conceptual or thematic proximity of medical entities or events that indicates they address similar or related clinical problems. Examples of specific embodiments include data points related to a patient's cardiological examination or specifically targeting complications of diabetes.

[0223] Unless otherwise specified, the term "causal relationship" in this document refers to a connection indicating that one event or factor in a clinical context directly influences or leads to another event or outcome. Examples of specific embodiments include "administering a drug leads to a reduction in blood pressure" or "a surgical procedure leads to a reduction in the risk of infection."

[0224] Unless otherwise specified, the term "correlation" in this document refers to a statistical or observed association between two or more clinical events or entities, without necessarily implying direct causation. Examples of specific embodiments include "association between elevated blood glucose levels and reported fatigue" or "simultaneous occurrence of fever and rash."

[0225] Unless otherwise specified, the term "structured data representation" refers to an organized and machine-readable format of medical information that enables systematic querying, storage, and processing of the data. Examples of specific implementations include a knowledge graph that specifies relationships between diagnoses and tests, an XML-based record, or a JSON-based structure that details diagnoses, laboratory results, and treatments.

[0226] Unless otherwise specified herein, the term "enrichment" (or "enrichment") refers to the process of improving one or more data elements by linking them with additional information or metadata from external sources. Examples of specific implementations include appending UMLS Concept Unique Identifiers (CUIs) to recognized diagnoses or linking drug prescriptions to standardized drug codes.

[0227] Unless otherwise specified herein, the term "external medical knowledge base" refers to an external knowledge repository that provides domain-specific information that can be used to supplement or clarify the recognized medical data. Examples of specific implementations include the Unified Medical Language System (UMLS), SNOMED CT, ICD, ICD-10GM, ICD-11GM, OPS, and other specialized ontologies or drug databases.

[0228] Unless otherwise specified herein, the term "convert...into an input format" refers to the transformation or conversion of medical data from an internal structure into a representation that can be processed by a language model or similar system. Examples of specific implementations include converting a knowledge graph into a JSON-based payload, generating text-based input prompts, or packaging extracted data into a user-defined schema for natural language processing.

[0229] Unless otherwise specified herein, the term “generative language model” refers to a computational model, typically based on deep learning architectures such as transformers, that is capable of generating coherent text in human language. Examples of specific implementations include autoregressive transformers with billions of parameters, large language models specialized in generating medical texts, or domain-optimized GPT-like architectures.

[0230] Unless otherwise specified herein, the term "natural language flow" refers to continuous prose in everyday language that is free of structured bullet points or lists and is understandable to human readers. Examples of specific forms include paragraphs describing patient histories, discharge summaries, or descriptions of clinical findings.

[0231] Unless otherwise specified herein, the enumeration "clinical reports, laboratory results, imaging reports, medication prescriptions, diagnostic lists, or procedure documentation" refers to various types of patient-related health data typically found in an electronic health record system. Examples of specific formats include PDF discharge summaries, HL7-formatted laboratory data, structured DICOM image reports, electronic prescription records, coded problem lists, and surgical reports.

[0232] Unless otherwise specified herein, the term "normalization of timestamps" refers to the conversion of various date and time data recorded in different formats into a uniform or standardized temporal representation. Examples of specific implementations include the conversion of local or regional date formats to ISO 8601 or the alignment of different time zones to a single reference time.

[0233] Unless otherwise specified herein, the term "recorded speech or dictation recordings" refers to audio recordings in which medical personnel orally document patient information or instructions, which are subsequently converted into text. Examples of specific implementations include digital speech recordings captured via mobile devices or telephone-based dictation services and stored on hospital servers. Unless otherwise specified herein, the term "named entity recognition model" refers to a computer system used to identify and classify important entities (e.g., diagnoses, medications, anatomical locations) in unstructured text. Examples of specific implementations include transformer-based natural language processing pipelines or rule-based entity extraction models specifically used in healthcare.

[0234] Unless otherwise specified, the term "transformer-based" refers to neural network architectures based on the self-attention mechanism, commonly used in natural language processing. Examples of specific implementations include BERT, GPT, and other state-of-the-art transformer variants for processing sequence data.

[0235] Unless otherwise specified, the term "domain-specific medical text corpus" refers to a specialized collection of medical texts used to train or fine-tune language models, typically compiled from clinical notes, research articles, or other health-related sources. Examples of specific implementations include anonymized electronic text corpora from health records, PubMed articles, or specialized medical literature repositories.

[0236] Unless otherwise specified herein, the term "fine-tuned" refers to the process of adjusting the parameters of a pre-trained model to domain-specific data in order to optimize performance for a particular task. Examples of specific implementations include further training a BERT-based model using oncology notes or adjusting GPT-based parameters using an internal dataset of clinical summaries from a hospital.

[0237] Unless otherwise specified herein, the enumeration "diagnosis, medication administration, procedure, laboratory value, or imaging finding" refers to representative categories of medical event types that the system recognizes as relevant to patient care. Examples of specific embodiments include the diagnosis of pneumonia, the administration of insulin, a laparoscopic appendectomy, a potassium laboratory result, or an MRI evaluation suggesting a ligament tear. Unless otherwise specified herein, the term "hierarchical clustering" refers to an algorithmic approach to grouping data by creating a multi-level or tree-like structure of clusters.Examples of specific implementations include agglomerative clustering, where the nearest data points are combined step by step, or divisive clustering, where a single cluster is divided into subclusters.

[0238] Unless otherwise specified, the term "density-based clustering" in this document refers to clustering methods in which data points are assigned to clusters based on areas of higher density compared to the rest of the dataset. Examples of specific implementations include DBSCAN or OPTICS algorithms, which are used to identify clinically relevant events based on their frequency and distribution over time.

[0239] Unless otherwise specified herein, the term "maximum permissible time interval" refers to a threshold that defines how far apart in time two medical events may be before they are excluded from the same group. Examples of specific implementations include requiring all events to be grouped within 24 hours or setting a 30-day limit for clustering chronic diseases.

[0240] Unless otherwise specified herein, the term "structural causal model" refers to a formal representation of assumed cause-and-effect relationships in clinical data, often used to derive direct causal connections. Examples of specific implementations include Bayesian networks for modeling patient outcomes, structural equation models for disease progression, or specific domain-based causal graphs for drug-response analyses.

[0241] Unless otherwise specified herein, the term "directed acyclic graph" refers to a graph structure with directed edges and no cycles, representing unidirectional relationships between nodes. Examples of specific implementations include a causal DAG stating that "Drug A → Reduced Symptom X → Improved Outcome Y" holds true without loops.

[0242] Unless otherwise specified herein, the term "knowledge graph" refers to a data structure in which entities (nodes) and their relationships (edges) are represented and stored in a graph-based format that supports semantic queries. Examples of specific implementations include a knowledge graph that links diagnoses with treatments, or a specialized graph that links medications with side effects.

[0243] Unless otherwise stated herein, the term "RDF triples" refers to a standard format for representing data in a knowledge graph as subject-predicate-object statements. Examples of specific implementations are "Patient123 - has Diagnosis - DiabetesMellitus" or "Diagnosis_X - is TreatedBy - Medication_Y".

[0244] Unless otherwise specified herein, the term "Unified Medical Language System (UMLS)" refers to a comprehensive compendium of many controlled vocabularies in the biomedical sciences, maintained by the U.S. National Library of Medicine. Examples of specific implementations include the use of UMLS Concept Unique Identifiers (CUIs) to reference specific medical concepts or synonyms.

[0245] Unless otherwise specified herein, the terms "SNOMED CT, ICD, ICD10GM, ICD11GM and OPS" refer to systematically organized collections of medical terms that provide codes, terms, synonyms and definitions for clinical documentation and reporting. Examples of specific implementations are SNOMED CT, ICD, ICD10GM, ICD11GM and OPS codes for diseases, findings, procedures or body structures.

[0246] Unless otherwise specified herein, the term "vector similarity measure" refers to metrics that quantify the similarity between vector embeddings of medical data or text. Examples of specific implementations include cosine similarity for Word2Vec embeddings, Euclidean distance for BERT-derived sentence embeddings, or specialized domain embeddings for medical terms.

[0247] Unless otherwise specified herein, the term "embedded representations" refers to numerical vector representations of words, phrases, or documents generated by machine learning models to capture semantic and contextual relationships. Examples of specific implementations include BERT embeddings for sentences, Word2Vec vectors for tokens, or domain-specific embeddings for lab results.

[0248] Unless otherwise specified herein, the term "JSON document" refers to a text-based data format (JavaScript Object Notation) structured with name-value pairs and arrays, commonly used for data exchange between systems. Examples of specific implementations include hierarchical structures containing subsets for diagnoses, laboratory tests, or drug data.

[0249] Unless otherwise specified herein, the term "prompt template with role description, instruction part, and context part" refers to a structured input format for prompting a generative language model, typically containing specific roles (e.g., System, User), instructions for response, and context data. Examples of specific implementations include a JSON payload with "System" style instructions, "User" instructions containing specific clinical questions, and "Context" containing patient data.

[0250] Unless otherwise stated, the term "autoregressive transformer" refers to a model architecture that generates each token or word in a sequence based on previously generated tokens and is typically used in large language models. Examples of specific implementations include GPT-based models that generate clinical paragraphs or advanced summarizing systems based on predicting the next token.

[0251] Unless otherwise specified herein, the term "Reinforcement Learning from Human Feedback" refers to the optimization of model outputs by incorporating iterative feedback from human raters into the training or fine-tuning loop. Examples of specific implementations include adjusting weights when summaries are rated as incorrect by clinicians or reinforcing style preferences derived from subject matter experts.

[0252] Unless otherwise specified herein, the term "maximum token length" refers to a defined upper limit on the number of tokens (partial word units) that the generative model outputs or processes in a single inference. Examples of specific implementations include limiting the model to 1024 tokens for a single discharge summary or limiting the output to 2000 tokens to prevent excessively long narratives.

[0253] Unless otherwise specified herein, the term "automatically written back to the patient's electronic health record" refers to the direct insertion or saving of the generated medical report into the patient's digital health record without manual intervention. Examples of specific implementations include writing the summary text to an FHIR-based medical record or updating the patient record in an HL7-compliant system.

[0254] Unless otherwise specified herein, the term "user interface" refers to any software interface that enables healthcare professionals to interact with, review, correct inaccuracies, and ultimately confirm the generated report. Examples of specific implementations include a web-based platform, a mobile application, or a desktop client integrated into the hospital's information system.

[0255] Unless otherwise specified herein, the term "pseudonymized or anonymized" refers to processes of modifying or removing a patient's personal data (PH) so that individuals cannot be identified without additional data (pseudonymization) or cannot be identified at all (anonymization). Examples of specific implementations include replacing patient names with random identifiers or removing direct identifiers such as social security numbers.

[0256] Unless otherwise specified herein, the term "local, in-house IT infrastructure" refers to the local hardware and software environment that is fully under the control of a healthcare provider and does not rely on external data centers. Examples of specific implementations include local servers running the model behind the hospital firewall or an internal data center licensed for storing and processing patient data.

[0257] Unless otherwise specified herein, the term "hardware-based security environment with protected memory" refers to a secure computing environment in which the execution of sensitive code and the storage of data are protected by special hardware features that prevent unauthorized access or manipulation. Examples of specific implementations include trusted execution environments such as Intel SGX enclaves or special secure processors that isolate the calculations of the generative model.

[0258] BRIEF DESCRIPTION OF THE FIGURES

[0259] The invention can be better understood with reference to the following figures: Fig. 1: A flowchart of a computer-implemented method for generating medical reports according to an exemplary embodiment of the present invention.

[0260] Fig. 2: A flowchart of a procedure for training a generative processor

[0261] Language model according to an exemplary embodiment of the present invention.

[0262] Fig. 3: An architecture of a generative language model for generating medical reports according to an exemplary embodiment of the present invention.

[0263] Fig. 4a: An input encoding data structure for a generative language model according to an exemplary embodiment of the present invention.

[0264] Fig. 4b: A medical report according to an exemplary embodiment of the present invention.

[0265] Fig. 5: A data processing device according to an exemplary

[0266] embodiment of the present invention.

[0267] DETAILED DESCRIPTION

[0268] The following section describes representative embodiments illustrated in the accompanying drawings. It should be understood that the illustrated embodiments and the following descriptions are examples and are not intended to limit the embodiments to a preferred embodiment.

[0269] The disclosed embodiments aim to automatically generate high-quality, clinically reliable discharge summaries. In some embodiments, patient-specific event data are fed into a large language model in a form structured by causal graphs and enriched with external expertise, in order to correctly represent causal relationships between diagnoses, interventions, and findings, thereby increasing the reliability of care.

[0270] Fig. 1 shows a flowchart of a computer-implemented method 100 for generating medical reports according to an exemplary embodiment of the present invention. The medical reports can be physician letters. The medical report can be a medical report as shown in Fig. 4b.

[0271] Procedure 100 can include a step of receiving 102 medical information about a patient. The received medical information can be formulated as natural language bullet points. Procedure 100 can include a step of generating 104 natural language flow text, which reflects the received medical information, for a medical report using a trained generative language model. This can be a generative language model like the one shown in Fig. 3.

[0272] The generative language model can be configured so that the generated natural language text reflects all and only the received medical information. Alternatively, the generative language model can be configured so that the natural language text deterministically reflects the received information.

[0273] The step of receiving medical information (102) can include receiving at least one user input via a graphical user interface, particularly from medical professionals during a patient visit. Additionally or alternatively, the step of receiving medical information (102) can include receiving at least one electronic record from a hospital information system. This electronic record can be Health Level 7 compliant and / or Fast Healthcare Interoperability Resources compliant. Additionally or alternatively, the procedure (100) can include generating the medical report as an electronic file containing the generated natural language text.

[0274] Medical information about patients can be stored in a database. This database may be an internal hospital database. Receiving at least one electronic data record from the hospital information system may involve sending a request to a hospital server. The hospital server may be part of the hospital information system and / or have access to the internal hospital database. The request may be an HTTPS request. This allows for straightforward implementation and integration of the present invention into an existing hospital information system, since other interfaces (e.g., TCP / IP) between the hospital server or the internal hospital database and other components of the hospital information system (e.g., backend and / or frontend) can remain unchanged.

[0275] The medical information received may include at least one piece of medical information belonging to one of the following categories: gender, leading symptom, diagnoses, treatments, medications, findings, and progress documentation.

[0276] Procedure 100 can include a step of generating an input encoding for the generative language model based on the received medical information. The input encoding can encode the medical information in such a way as to accommodate a maximum input length for the generative language model. The input encoding can be generated such that each piece of medical information in the input encoding is assigned to a predefined category abbreviation indicating the respective category. The input encoding can be generated such that, for at least one of the categories, the input encoding contains multiple pieces of medical information from the same category. The category abbreviation can only be included once in the input encoding. The input encoding can be generated such that the multiple pieces of medical information from the same category are separated by a predefined separator abbreviation.The input encoding can be generated such that at least one piece of medical information in the input encoding is assigned a timestamp, in particular after the medical information and separated by a predefined separator abbreviation. The category abbreviation and / or the separator abbreviation can contain exactly one character. The input encoding can be as shown in Fig. 4a.

[0277] Regardless of whether the input coding described above is used or the medical information is categorized in another way, the use of Procedure 100 can significantly reduce the duration of a patient admission. Practical trials have shown the following: the duration of the medical history was reduced from 10 minutes to 2 minutes, and the duration of the physical examination from 10 minutes to 5 minutes. Assuming a duration of 5 minutes for recording medications, this results in a time saving of 13 minutes (i.e., 12 instead of 25 minutes).

[0278] By using Procedure 100, the duration of a ward round can be reduced as follows: preparation time from 5 minutes to 0 minutes, follow-up time from 5 minutes to 0 minutes. Assuming a duration of 5 minutes for ward round documentation, this results in a time saving of 10 minutes (i.e., 5 instead of 15 minutes). By using Procedure 100, the duration of a patient discharge can be reduced as follows: the duration of discharge summary writing from 30-60 minutes to 5 minutes, of which the time spent copying old findings is reduced from 15 minutes to 0 minutes and the time spent writing the discharge summary from 15-45 minutes to 5 minutes. Assuming a duration of 5 minutes for the treatment recommendation, this results in a time saving of 25-55 minutes (i.e., 5 instead of 35-65 minutes). This assumes an average length of stay of 4 days.

[0279] Furthermore, using method 100 can achieve a time saving of 65-72% (i.e. 42 instead of 120-150 minutes).

[0280] Fig. 2 shows a flowchart of a method 200 for training a generative language model according to an exemplary embodiment of the present invention. This can be a generative language model as shown in Fig. 3.

[0281] Method 200 can include a step of providing 202 training datasets. A training dataset can comprise an input encoding data structure according to aspects of the present invention. It can, for example, be an input encoding data structure as shown in Fig. 4a or otherwise categorized input information. A training dataset can comprise natural language flow text for a medical report. The medical report can be a medical report as shown in Fig. 4b. Method 200 can include a step of training 204 the generative language model by supervised learning with the training datasets.

[0282] During training, an initial set of training data can be generated. This first set of training data can be artificially generated. Using this first set, the generative language model can be pre-trained. This pre-trained generative language model can then be deployed (e.g., integrated into a hospital information system) and re-trained or fine-tuned using a second set of training data. This second set of training data can be real-world data. In this context, "real" means that it consists of medical information from patients in the relevant hospital information system. This allows for a language model that is better tailored to individual patient types (e.g., patients in specific specialist clinics).Before training the model, data preprocessing can be performed, in which the training data is divided into a training, a validation and a test data set.

[0283] A loss function can be used when training the model. This loss function can be suitable for sequential generation tasks. For example, it can be a cross-entropy loss function. The loss function can measure the inconsistency between the model's predicted output (e.g., the natural language flow of a generated medical report) and the actual target output (e.g., the corresponding, actual natural language flow of the medical report). Based on this measurement, the model can be trained to minimize the error between the predicted and target outputs, thereby improving its prediction accuracy.

[0284] An optimizer can be used when training the model. The optimizer can be specialized for large models (i.e., in terms of the model's size or depth) and large datasets (e.g., in terms of the amount of data and the file size of a single training file). Alternatively or additionally, the optimizer can be suitable for transformer models (i.e., in terms of its efficiency). For example, the optimizer could be an Adafactor optimizer.

[0285] A generation function can be used when training the model. A parameter set for the generation function can include: input token IDs (e.g., 512 IDs with a maximum input length of 300 of 512 for the generative language model), a number of the highest probabilities from which to select (e.g., 50), a threshold to limit the probability of the selected tokens (e.g., 0.95), a factor to influence the probability distribution (e.g., 0.3, where a value < 1 results in a deterministic output), an early-stopping indicator (e.g., True, so that generation is stopped early if the end-of-sequence condition is met), and a number of beams for a beam-search algorithm.

[0286] A set of hyperparameters can be used when training the model. This hyperparameter set can include, for example, a batch size and a learning rate. The hyperparameter set can also include a subset of parameters used to parameterize the optimizer. This subset of parameters can include a learning rate, a learning rate decay rate, a first epsilon value, a second epsilon value, a clipping threshold, an adaptive learning rate indicator, a weight decay value, a clipping norm value, a clipping value, a global clipping norm value, an EMA (Exponential Moving Averages indicator), or any combination thereof. The learning rate can be 0.001. The learning rate decay rate can be -0.8. The first epsilon value can specify a distance to prevent a denominator from becoming zero. For example, the first epsilon value can be 1e-30.The second epsilon value can specify a margin to prevent the learning rate from becoming too small. For example, the second epsilon value can be 1e-3. The clipping threshold can be 1. The adaptive learning rate indicator can be a Boolean value, where "True" indicates that an adaptive learning rate (i.e., the learning rate is adjusted based on the current training iteration) is used, and "False" indicates that no adaptive learning rate is used. The weight decay value can be a float value. Preferably, the weight decay value is zero, so no weight decay occurs. The clipping norm value can be a float value. The clipping norm value can indicate a value that the norm of a gradient for each weight must not exceed (i.e., the value at which the norm of a weight is clipped). The clipping value can be a float value.The clipping value can indicate a value that must not exceed the value of any weight's gradient (i.e., the value to which a weight's value is clipped). The global clipping norm value can indicate a value that must not exceed the global norm of all gradients of all weights. The EMA indicator can take a Boolean value, where "True" indicates that an EMA is used and "False" indicates that no EMA is used. With EMA, an exponential moving average is calculated based on the model's weights, periodically overwriting the actual weight values. The EMA momentum value can take a float value, preferably 0.99. The EMA momentum value can only be used if the EMA indicator is "True". The EMA overwrite frequency can be a positive integer, including zero.The EMA overwrite frequency can only be used if the EMA indicator is "True". If the EMA overwrite frequency is zero, no model variables (e.g., weights) are overwritten during training. If the EMA overwrite frequency is √(i) where i > 0, then the model variables (e.g., weights) are overwritten with the moving average every i-th iteration. The loss scaling factor can take a float value. If the value is zero, no loss scaling factor is applied. Otherwise, the loss scaling factor is multiplied by the loss before the gradients are calculated. The gradient accumulation step size can take an integer value, including zero. If zero, no accumulation step size is used (i.e., the model and / or optimizer variables are updated in every iteration). Otherwise, the variables (e.g., weights) are updated only every √(i-th) iteration.

[0287] Fig. 3 shows an architecture of a generative language model 300 for generating medical reports according to an exemplary embodiment of the present invention. The medical report can be a medical report as shown in Fig. 4b. The medical reports can be medical letters. The generative language model 300 can be configured to generate natural language flow text for a medical report. The generated natural language flow text can reflect medical information about a patient. The medical information can be formulated in natural language bullet points. The generative language model 300 can be configured for use in method 100. The generative language model 300 can be trained using method 200.

[0288] The generative language model 300 can be a transformer-based generative language model. For example, it can be a text-to-text transformer model such as a T5 model or a Mistral model. The generative language model 300 can be a model configured to process sequential tasks.

[0289] The exemplary architecture shown in Fig. 3 corresponds to the architecture of a generative language model 300 based on the T5 model. The generative language model 300 comprises an encoder and a decoder part. As shown, both the encoder and decoder parts can be based on the Transformer design. One or both parts can include multiple layers of multi-head self-attention mechanisms and feed-forward networks. The generative language model 300 can include a special embedding for tokens and positions. This special embedding can be added at the beginning of the model.

[0290] The encoder section can be configured to receive input (e.g., structured bullet points) and process it to create a context-rich representation of the input. This context-rich representation then serves as input for the decoder section, which generates output text (e.g., one or more sections of the medical report) based on this context-rich representation. When generating further output text (e.g., one or more subsequent sections of the medical report), the decoder section can use the previously generated output text as input, thus providing the decoder with additional contextual information.

[0291] The structured bullet points, which may be in the form of an input encoding (as explained in Fig. 4a), can be tokenized. Tokenization converts the input encoding into a numerical representation. The context-rich representation can then be created from this numerical representation.

[0292] To generate the programming code for the generative language model 300, the "Transformers" software package can be used, for example, which offers a collection of pre-trained models and functionalities for natural language processing. Additionally, a text tokenization library (e.g., "Sentencepiece") can be used.

[0293] Fig. 4a shows an input encoding data structure 400a for a generative language model 300 according to an exemplary embodiment of the present invention.

[0294] As shown, the 400a input code can contain structured bullet points representing various pieces of medical information. In other words, the 400a input code for the generative language model 300 can be generated based on received medical information. The 400a input code enables an efficient and compact representation of the medical information, thus ensuring both no information loss and efficient processing (e.g., with regard to the required processing resources).

[0295] The compactness of the 400a input code can be achieved in one possible embodiment by using appropriate input or category abbreviations, thereby reducing the required data volume (e.g., "g" for gender, "I" for leading symptom, "d" for diagnosis, "b" for treatment, and / or "e" for discharge summary). In other words, the 400a input code can be generated such that each piece of medical information in the 400a input code is assigned to a predefined category abbreviation that indicates the respective category. For example, the category abbreviation "g" can indicate the patient's gender. The use of category abbreviations can be particularly important if the generative language model 300 used has a maximum input length (e.g., 512 tokens). In other words...The 400a input encoding can encode medical information in a way that respects the maximum input length of the generative language model 300. This allows for the avoidance of redundant words / characters and instead provides a format with low memory requirements. This low memory requirement also enables a large amount of medical information to be transmitted in a compact, machine-readable format (i.e., the input encoding), allowing the generative language model 300 to generate detailed and nuanced medical reports.

[0296] The input code 400a can be generated such that it contains multiple pieces of medical information from the same category for at least one of the categories. For example, the input code 400a shown contains multiple pieces of medical information from category "b" (i.e., treatment).

[0297] The 400a input code can be generated in such a way that multiple pieces of medical information within the same category are separated by a predefined separator. For example, multiple pieces of medical information in category "b" are separated by the assigned separator "+". This allows for clear differentiation between the multiple pieces of medical information. This differentiation capability enables multiple treatments, medications, or progress documentation (observations) to be specified within a single 400a input code (e.g., an input string), while the medical information remains clearly separated.

[0298] The 400a input code can be generated such that at least one piece of medical information within it is assigned a timestamp, specifically placed after the medical information and separated by a predefined separator. The category code and / or the separator can contain exactly one character. For example, a treatment such as "Infusion Therapy" can be assigned a timestamp (e.g., 01.07.2021) indicating the date and time the treatment was administered. Although the timestamp is displayed after the medical information in the example shown, it can also be displayed before it. Additionally, a space can be used between the medical information and its corresponding timestamp. The timestamps (e.g., dates) allow for the representation of a chronological sequence or progression of symptoms, diagnoses, and treatments.

[0299] The category abbreviation can only appear once in the 400a input code. For example, the 400a input code shown contains the category abbreviations "g", "I", "d", "b", and "e" only once. The 400a input code can be generated from natural language bullet points. These bullet points can be provided by medical personnel (e.g., via a text field in an app and / or a speech recognition interface). The 400a input code shown can be generated from raw input (i.e., the natural language bullet points) that includes, among other information (e.g., gender): "Fever, fell out of bed, 2 units of packed red blood cells, vomiting".

[0300] Fig. 4b shows a medical report 400b according to an exemplary embodiment of the present invention. The medical report 400b can be a doctor's letter.

[0301] The medical report 400b can include generated natural language flow text that reflects all received medical information or only the received medical information. The natural language flow text can reflect the received information deterministically. The medical report can be generated as an electronic file containing the generated natural language flow text.

[0302] As can be seen, medical report 400b contains all the medical information from input code 400a. Furthermore, medical report 400b does not contain any medical information that is not also included in input code 400a. In addition, medical report 400b, or rather the corresponding natural language text, can deterministically reflect the medical information of input code 400a. In other words, based on input code 400a, the same natural language text is always generated for medical report 400b.

[0303] Fig. 5 shows a data processing device 500 according to an exemplary embodiment of the present invention.

[0304] The data processing device may include means for carrying out the methods (e.g., methods 100 and / or 200). The means may be a processor 502 and a memory 504. The processor 502 and the memory 504 may be operatively connected. A computer program may be stored in the memory 504, wherein the computer program comprises instructions which, when the computer program is executed by a computer, cause the computer to execute the method according to any of the aforementioned aspects (e.g., methods 100 and / or 200). In a further embodiment, the invention relates to a computer-implemented method for generating medical reports, in particular medical letters, wherein the method comprises at least the following steps, or a subset thereof: a) capturing medical information about a patient from at least two different sources of an electronic health record,where the received medical information is formulated at least partially in natural language bullet points; b) Identifying medical units contained in the recorded medical information and classifying each identified medical unit according to a medical event type by computer-aided pattern recognition; c) Forming event groups by computer-aidedly grouping medical units that are related at least with regard to temporal reference, treatment pathway, or thematic relevance; d) Deriving relationship information that represents at least one causal relationship or correlation between elements of the event groups,and storing the elements of the event groups and the relationship information in a structured data representation; e) enriching at least one element of the structured data representation by linking it to data records from at least one external medical knowledge base; f) converting the enriched structured data representation into an input format that can be processed by a generative language model; and g) processing the input format with the generative language model to generate natural language flow text that reflects the captured medical information for a medical report.

[0305] In some embodiments, the following phases, or suitable subsets thereof, are implemented:

[0306] 1. Raw data from the EHR (Electronic Health Record) (input layer) — Sources: clinical reports, lab results, imaging reports, medications, diagnoses, procedures, etc. Event detection and categorization

[0307] • (Preprocessing phase) — Identification of medical events, diagnoses, and medication orders. Procedures with or without the use of ML / LLM models.

[0308] • Configurable algorithm: e.g., NER (Named Entity Recognition) models, transformer-based models, or rule-based structures. Lustering of events and event streams.

[0309] • (Organization phase) — Grouping of events based on relevance, time course, or patient pathway.

[0310] • Configurable algorithm: Clustering models (e.g., hierarchical clustering, density-based methods, sequential models). Finding causal relationships

[0311] • (Inference phase) — Determining causal relationships (e.g., drug — > side effect, imaging — > finding).

[0312] • Configurable algorithm: Causal inference models (e.g., graph-based models, structural causal models). Preferred implementation: Knowledge graph for generating medical reports.

[0313] • This establishes a causal relationship or at least a correlation between the data in order to provide more metadata for the model and to better interpret the data. Enrichment through knowledge bases (KBs)

[0314] • (Knowledge enrichment phase) — Linking events with external knowledge bases (e.g. UMLS, SNOMED, ​​own knowledge bases).

[0315] Configurable algorithm: entity linking and enrichment modules. Transformation into an LLM-compliant structure (preparation phase) — conversion of events and knowledge into structured prompts or input formats for LLMs.

[0316] • Configurable algorithm: Structured document layout, JSON / graph formatter, triplets.

[0317] 7. Generating the summary

[0318] • (Distribution phase) — Creating understandable, accurate and clinically relevant summaries using LLMs.

[0319] • Configurable algorithm: Summary LLMs (fine-tuned models, tuned prompts, agent-based).

[0320] In some embodiments, the event recognition and categorization process includes a text processing module that receives incoming medical information, some of which is in the form of natural language bullet points. This module splits the text into sequences and performs preprocessing, normalizing timestamps to a uniform time format. The named entity recognition model, preferably transformer-based and fine-tuned with a domain-specific medical text corpus, identifies medical entities within the text. These are then assigned by the classification module to different medical event types: diagnoses, medication administrations, procedures, laboratory results, and imaging findings. The classified medical entities, along with their metadata such as timestamps and source references, are provided for further processing.

[0321] In some implementations, the classified medical units are fed into the clustering algorithm during the medical event clustering process. Depending on the configuration, this can be implemented as hierarchical or density-based clustering. The clustering module uses various parameters for grouping, including the time reference, where a maximum permissible time interval between medical units to be included in the same group can be specified. Further grouping parameters include the treatment pathway and thematic relevance. The process generates several event groups, which are represented in the diagram as connected clusters. Each event group contains medical units that are thematically related. The grouping results are then prepared for the subsequent inference phase.In some implementations, a knowledge graph depicting causal relationships between medical events is created by feeding the event groups to the inference module, which derives relationship information using a structural causal model. This structural causal model is implemented as a directed acyclic graph that represents causal relationships or correlations between the elements of the event groups. The diagram shows examples of derived causal relationships, such as the relationship between medication administration and observed side effects, or between diagnostic procedures and resulting findings. The event group elements and the relationship information are stored in a structured data representation designed as a knowledge graph.

[0322] In some implementations, medical data is enriched with external knowledge bases by feeding the knowledge graph, including event group elements and relationship information, to the enrichment module. This module links the knowledge graph elements with datasets from external medical knowledge bases, such as the Unified Medical Language System (UMLS), SNOMED CT, ICD, ICD-10GM, ICD-11GM, and OPS. The linking is performed using a vector similarity measure based on embedded representations. The figure illustrates various linking examples, such as enriching diagnoses with standardized codes, extending drug information with drug data and interaction profiles, and supplementing treatment procedures with evidence-based guidelines.The enriched structured data representation now contains additional semantic information that can be used for subsequent report generation.

[0323] In some implementations, the transformation of the structured data representation into an LLM-compatible input format involves feeding the enriched structured data representation to the transformation module, which converts it into a format processable by the generative language model. The knowledge graph is transformed into a structured text representation. The resulting input format is structured as a JSON document, comprising several sections: a role description that provides the language model with the context of being a medical reporter, an instruction section containing specific instructions for report generation, and a context section containing the transformed medical information. Different levels of structuring within the JSON document allow for a hierarchical organization of the medical data.In some implementations, the LLM-compatible input format is fed into a generative language model during the report generation process. This model can be implemented, for example, as an autoregressive transformer with at least one billion model parameters. The model's internal architecture can incorporate multi-head attention mechanisms and feed-forward networks. In this example, the model has been specifically fine-tuned for medical summaries using reinforcement learning from human feedback. A maximum token length of 1024 tokens is specified for output during processing. To ensure data security, all patient-identifying information is pseudonymized or anonymized before being transmitted to the model, and processing takes place in a hardware-based security environment with protected memory.The result is a natural language flowing text for the medical report, which can be edited and approved by medical professionals via a user interface before being automatically written back into the patient's electronic health record.

[0324] In some embodiments, an agent-based approach can be used to generate summary medical reports, where one or more specialized software agents coordinate the use of large language models, rule-based processing, and domain-specific requirements to produce coherent and clinically relevant documents. In these embodiments, individual agents can be responsible for tasks such as content selection, data validation, or style formatting, and cooperate with a central control agent that integrates the partial results and passes a unified prompt to a large language model, thus enabling modular responsibilities to improve the quality of the summary.In embodiments, a multi-stage agent pipeline can also be established, comprising (i) a content filtering agent that assesses the clinical importance of each detected entity, (ii) a validation agent that checks the selected information using external knowledge bases, and (iii) a synthesis agent that iteratively interacts with a large language model to create the final text, automatically adding institutional reporting standards and required legal notices before the final report is transferred to the electronic health record.

[0325] Application Example 1: Automated Discharge Summary Generation Using AI-Based Data Processing. This implementation example illustrates a hypothetical project for the development and validation of a computer-implemented method for the automated generation of medical reports, particularly discharge summaries (discharge summaries), using artificial intelligence and employing the concepts of the invention. The experiment addresses the challenges identified in the prior art, especially the considerable documentation burden on medical personnel, which can lead to reduced quality of care and increased workload. A multi-stage process was developed and tested that automatically transforms raw medical data from various sources into a coherent natural language report.

[0326] For the experiment, an IT infrastructure consisting of several components is set up within a healthcare facility. Electronic health records (EHRs) containing clinical reports, laboratory results, imaging reports, prescriptions, diagnosis lists, and procedure documentation serve as data sources. The infrastructure also includes computing units for the various processing steps and a hardware-based security environment with protected memory for executing the generative language model. External knowledge from the medical databases UMLS (Unified Medical Language System), SNOMED CT, ICD, ICD-10GM, ICD-11GM, and OPS is integrated. A named entity recognition model with a transformer architecture, fine-tuned with a domain-specific medical text corpus, is used for data processing.Hierarchical and density-based clustering algorithms are used to cluster the identified entities. The core of the text generation process is an autoregressive transformer with over one billion model parameters, optimized for medical summaries using reinforcement learning from human feedback.

[0327] The experimental process is conducted in seven phases, beginning with data collection from electronic health records. In the first phase, medical information is captured from at least three different sources, with timestamps normalized to a uniform format and, where necessary, audio or dictation recordings converted to text. The second phase involves event recognition and categorization using the transformer-based Named Entity Recognition model, which identifies medical entities and assigns them specific event types such as diagnosis, medication administration, procedure, lab result, or imaging finding. The third phase comprises event clustering, where medical entities are grouped based on temporal reference, treatment pathway, or thematic relevance. A maximum permissible time interval between the entities to be grouped is defined.In the fourth phase, causal relationships between the elements of the event groups are derived using a structural causal model in the form of a directed acyclic graph. The resulting structured data representation is stored as a knowledge graph with RDF triples. The fifth phase involves enriching the data by linking it to external knowledge bases, using vector similarity measures based on embedded representations. In the sixth phase, the enriched data is transformed into a JSON document containing a prompt template, role description, instruction section, and context section, which can be processed by the generative language model. In the seventh and final phase, the natural language text for the medical report is generated, with a maximum token length of 1024 tokens defined for the output.

[0328] The developed method is capable of generating a coherent and clinically relevant medical report from the various sources within a patient's electronic health record. The automatically generated reports comprehensively reflect the captured medical information and present it in a well-structured, natural language format. The integration of external knowledge bases enriches the reports with relevant medical contextual information. The use of the finely tuned generative language model results in high linguistic quality and clinical precision in the generated texts. The generated reports can be automatically written back to the electronic health record, with a user interface available for medical professionals to edit the reports before final release.

[0329] Application example 2: Implementation of AI-supported reporting in pediatric oncology

[0330] In the pediatric oncology department of a university hospital, medical information from five different sources is collected: electronic patient records, imaging studies, laboratory analyses, medication documentation, and electronic nursing notes. A specialized algorithm is used to normalize the data chronologically, converting both exact timestamps and relative time references (such as "three days ago," "since last week") into a standardized ISO 8601 format. Handwritten notes from ward physicians are captured using OCR technology, and a speech recognition system converts daily dictation from ward rounds into text.

[0331] The system uses a hybrid named entity recognition model that combines rule-based components with a BioBERT transformer specifically fine-tuned for pediatric oncology terminology. In addition to standard event types, event classification also captures pediatric-specific event types such as "growth measurement," "developmental milestone," and "family care situation." Event clustering is performed using a density-based method (DBSCAN), with the temporal tolerance for grouping dynamically adjusted to the treatment phase—tighter time limits are applied for acute interventions than for long-term observations.

[0332] To derive causal relationships, the system uses a hierarchical Bayesian network that takes into account pediatric characteristics regarding drug effects and therapy tolerability. The structured data representation is persisted as a property graph, which, in addition to RDF triples, also includes weighted edges to represent the strength of clinical correlations. The enrichment of the identified entities is achieved through a combination of UMLS, the pediatric sub-ontology of SNOMED CT, ICD, ICD-10GM, ICD-11GM, and OPS, and an internal clinical knowledge base of pediatric oncology protocols.

[0333] The input format for the generative language model is a structured XML document that organizes medical information in hierarchically nested elements and includes age-dependent relevance markers. A domain-specific model with 7 billion parameters, trained through contrastive learning between child-friendly and medical-technical formulations, is used to generate the medical letters. During the generation phase, the system automatically adjusts the level of detail and the complexity of the medical terminology depending on the intended audience (parents, referring pediatrician, specialists).

[0334] Application example 3: AI-supported report generation in radiology and nuclear medicine

[0335] In the radiology department, a system is used to automatically generate reports for CT, MRI, and PET scans. Medical information is extracted from the DICOM image archive, the radiology information system (RIS), previous reports, and the electronic patient record. Computer vision algorithms are used to process the imaging data, recognizing anatomical structures and detecting anomalies. This visual data is then combined with the textual metadata from the DICOM headers and the clinical information from the RIS.

[0336] The normalization of timestamps takes into account sequences within an examination and allows for the temporal classification in relation to contrast agent administration or physiological processes. For radiological interpretation, primarily written reports are processed, but also audio recordings of the interpreting radiologists, which are transcribed in real time during image analysis using a special medical speech recognition system.

[0337] For event recognition, a multimodal transformer model is used that can process both image and text data and links visual features with textual descriptions. The named entity recognition model was trained on over 50,000 anonymized radiological reports and recognizes specialty-specific entities such as "lesion," "contrast enhancement," or "signal alteration." The recognized events are classified according to radiological categories, including "normal," "acute pathology," "chronic change," "incidental finding," and "technical artifact."

[0338] The events are clustered using a hierarchical agglomerative clustering algorithm, with anatomical regions serving as the primary organizing principle and a maximum time interval of 90 days defined for follow-up observations of the same finding. Causal relationships are derived using a radiological inference model based on a directed acyclic graph, which links radiological findings with clinical diagnoses and differential diagnoses.

[0339] The structured data representation is stored in a special knowledge graph that can also depict spatial relationships between anatomical structures. To enrich the data, the radiological sub-ontology of RadLex and the anatomical reference database FMA (Foundational Model of Anatomy) are used, with the linking achieved via semantic vector representations of anatomical terms.

[0340] The input format for the generative language model consists of a JSON-LD document containing both structured diagnostic data and references to images. The generative language model is a medical transformer with 5 billion parameters, specifically optimized for generating structured radiological reports and aligned with the reporting standards of the European Society of Radiology.

[0341] Application example 4: AI-based progress reporting in chronic treatment with integration of patient data from wearables

[0342] At the Center for Chronic Diseases, the system is used to generate quarterly progress reports for patients with diabetes mellitus, COPD, and chronic heart failure. In addition to conventional clinical data from the electronic health record, data from wearables (smartwatches, blood glucose sensors), home healthcare devices (blood pressure monitors, spirometers), and patient apps are also collected. The system thus integrates data from over ten different sources, including electronic medication dispensers and smart pillboxes that monitor medication adherence.

[0343] Temporal normalization is particularly complex, as continuous data streams with varying sampling rates (from minute intervals to weekly measurements) must be converted into a uniform temporal scheme. The system uses a specialized time-series database that harmonizes different granularities and identifies data gaps using statistical methods. Audio recordings of telephone consultations are converted into text using an adaptive speech recognition system trained on the individual speech patterns of the medical staff.

[0344] For event detection, a multiparametric deep learning model is used that can process both regular and irregular time series data. The system recognizes not only individual measurements but also trends, fluctuation patterns, and threshold violations. In addition to standard clinical types, the event classification includes patient-reported events such as "symptom worsening," "adherence problems," and "lifestyle changes," which are extracted from structured questionnaires and free-text responses from patients.

[0345] Events are clustered using temporal sequence clustering, which considers not only temporal proximity but also causal sequences and recurring patterns. The algorithm automatically identifies episodes of clinical deterioration and associated factors. A dynamic Bayesian network is used for causal analysis, capable of modeling time-delayed effects and quantifying relationships between medication use, lifestyle changes, and clinical parameters. The structured data representation is stored as a heterogeneous time-series graph that integrates various time scales and can be queried using temporal logic. Data enrichment utilizes not only medical knowledge bases such as UMLS but also guideline databases and pharmacokinetic models. Integration is achieved through context-dependent embedding models that incorporate the clinical context into similarity calculations.

[0346] The input format for the generative language model is a hybrid format of structured data and natural language elements that visualizes the temporal development of patient parameters. The language model is specialized for generating longitudinal reports and is optimized through continuous learning from expert feedback. It automatically generates therapy recommendations based on guidelines and individual patient data, which can be reviewed and adjusted by the treating physician.

[0347] Application example 5: AI-supported reporting in psychiatric and psychotherapeutic care

[0348] In the psychiatric clinic, the system is used to generate therapy progress and discharge reports. The recorded medical information includes structured diagnoses according to ICD-10 / DSM-5, psychometric tests, medication data, therapy documentation, and free narrative entries from therapy sessions. A unique feature is the integration of audio recordings of therapeutic conversations, which, with the patients' consent, are transcribed using a speech recognition system specifically trained for therapeutic language. The system also processes paraverbal features such as pauses, intonation, and speech rate as additional diagnostic information.

[0349] Timestamp normalization, when processing therapy protocols, considers not only the session date but also retrospective time references from patients relating to biographical events. A temporal reasoning module places these events within a biographical timeline. Event recognition utilizes a BERT-based language model, specifically fine-tuned to psychological and psychiatric terminology and capable of recognizing subtle linguistic markers of mental states. Events are classified according to psychiatrically relevant categories such as "symptom," "life event," "coping strategy," "therapy intervention," and "therapy response." Additionally, nonverbal events, such as documented behavioral observations and results of standardized psychometric tests, are also processed.The clustering of events is carried out through thematic clustering using Latent Dirichlet Allocation, whereby psychologically relevant thematic complexes are identified and grouped chronologically.

[0350] For causal analysis, a specialized cognitive-behavioral therapy model is implemented, based on the cognitive triangle model, which maps relationships between thoughts, feelings, and behavior. The structured data representation uses a narrative knowledge graph that can also represent subjective experiences, patient perspectives, and therapeutic hypotheses as special node types.

[0351] The data is enriched by linking it to psychological ontologies such as the Mental Function Ontology and evidence-based psychotherapeutic manuals. The input format for the generative language model is a narrative schema that structures and represents therapeutic developmental trajectories and change processes. The language model is trained on empathetic, non-stigmatizing formulations and adheres to ethical guidelines for psychiatric documentation. It automatically adjusts the level of detail and the terminology used depending on the intended audience (patient, referring therapist, insurance provider).

[0352] Application example 6: Multilingual AI-supported reporting in international healthcare institutions

[0353] In an international hospital with patients and staff from various language backgrounds, the system is used for the cross-linguistic generation of medical reports. Medical information is collected from various sources in different languages, including electronic patient records in up to five primary languages ​​(German, English, French, Arabic, and Mandarin), laboratory reports with standardized international codes, and imaging studies with multilingual findings.

[0354] For cross-linguistic processing, a special normalization module is implemented that maps medical terminology from different languages ​​to a common standard (UMLS metathesaurus). Timestamp normalization takes into account different date formats and time zones and converts all data to an ISO 8601-compliant UTC standard. Audio data from patient histories in various languages ​​is processed by a multilingual speech recognition system that can recognize and transcribe 12 different languages.

[0355] Event recognition is performed using a multilingual XLM-RoBERTa model, which was trained in parallel in all five primary languages ​​and can identify cross-linguistic medical entities. A language-agnostic approach is used for event classification, based on international classification systems such as ICPC-2 (International Classification of Primary Care), enabling language-independent categorization.

[0356] The clustering of events is performed using a culture- and language-sensitive algorithm that also considers culture-specific descriptions of symptoms and experiences of illness. For causal analysis, a multilingual knowledge graph is used that integrates medical knowledge from various cultural contexts and takes culture-specific patterns of interpretation into account.

[0357] The structured data representation uses a multilingual graph that links concepts in different languages ​​and connects them through cross-linguistic embeddings. International knowledge bases such as the UMLS metathesaurus and the WHO ICD-11, as well as language-specific medical ontologies, are used for enrichment. The linking is achieved via cross-linguistic biomedical embeddings that recognize semantic equivalence across language boundaries.

[0358] The input format for the generative language model is a language-agnostic semantic representation format that encodes medical concepts independently of language. The generation process uses a multilingual language model with 175 billion parameters, capable of producing medical reports in any of the five primary languages. The system automatically adapts culture-specific formulations and medical terminology to the target language and cultural context.

[0359] Application example 7: AI-based emergency reporting with accelerated workflow

[0360] In the emergency department, a specially adapted version of the system is used, optimized for the rapid generation of emergency reports, transfer notes, and handover protocols. Medical information is captured in real time from the hospital information system, vital sign monitors, the emergency laboratory with point-of-care testing, and pre-hospital documentation from the ambulance service. Additionally, voice recordings of medical staff during emergency care are processed by a real-time speech recognition system that functions reliably even in noisy environments.

[0361] Timestamp normalization is performed in real time with highly precise granularity down to the second level to accurately document the time-critical course of emergency treatments. The system automatically synchronizes timestamps from various devices and systems. Event detection utilizes a highly efficient transformer model optimized for emergency terminology, identifying critical events such as "vital parameter changes," "drug administration," and "interventional measures" in real time.

[0362] Events are classified according to emergency medical priorities, with life-threatening conditions, urgent interventions, and secondary findings hierarchically ordered. Event clustering uses a streaming clustering algorithm that allows for continuous updates to event groups as new information becomes available. The maximum time interval for grouping is dynamically adjusted to the urgency of the case—shorter time windows are used for critical emergencies than for less urgent cases.

[0363] For causal analysis, a specialized emergency medical decision model is used, based on clinical treatment pathways and standardized emergency algorithms. The structured data representation is a dynamic graph that is continuously updated and represents priorities through weighted edges. Data enrichment occurs in real time through linking with emergency protocol databases and drug interaction databases.

[0364] The input format for the generative language model is a progressively expanded JSON document that incrementally integrates new information. A specially optimized language model is used for report generation, capable of producing reports within seconds and creating emergency reports structured according to the SOAP (Subjective, Objective, Assessment, Plan) schema. The system allows for continuous report updates during ongoing treatment and provides a dedicated mobile interface for emergency personnel.

[0365] Application Example 8: Modular Use of Clinical Trial Reporting in Small Medical Practices and Outpatient Facilities For smaller medical facilities such as general practitioner practices and outpatient centers, a scalable version of the system is implemented, adapted to the limited IT infrastructure. Medical information is captured from the practice management system, electronic forms, referral letters, and locally stored examination results. The system is modular and allows for the selective activation of required components depending on the practice's focus and available resources.

[0366] The timestamp normalization is tailored to the typical documentation patterns in medical practices and takes into account recurring appointments, chronological sequences of referrals and return referrals, as well as recall intervals for preventive check-ups. Audio data from doctor-patient conversations is processed by a data protection-compliant local speech recognition system that operates without a cloud connection and processes the data exclusively on local servers.

[0367] For event detection, a resource-efficient model is used that runs on standard practice computers without specialized hardware. The model employs a combination of rule-based methods and a compressed transformer model optimized for typical documentation patterns in primary care. The event classification is tailored to the needs of primary care and includes categories such as "presentation," "history," "examination," "treatment," and "progress."

[0368] Event clustering is case-based, focusing on treatment episodes and chronic diseases. A simplified inference model is used for causal analysis, mapping typical primary care treatment pathways and supplemented by practice-specific treatment patterns. The structured data representation is implemented as a lightweight graph, allowing it to run on a relational database system even without a dedicated graph database.

[0369] Data enrichment is achieved selectively by linking it to practically relevant parts of medical knowledge bases such as ICD-10 codes, drug databases, and laboratory reference ranges. The input format for the generative language model is a minimally structured template adapted to typical primary care medical report formats. The system uses a resource-efficient language model with 3 billion parameters, which can run on standard CPU hardware and is specifically optimized for the requirements of outpatient care. Application example 9: Integration of AI reporting into clinical research projects

[0370] In a university research center, the system is used for the automated generation of study reports and clinical case summaries for research purposes. Medical information is gathered from research databases, clinical study documents, patient records, and structured case report forms (CRFs). A particular requirement is the integration of raw data from various diagnostic procedures, genomic analyses, and biomarker measurements.

[0371] Timestamp normalization takes into account study-specific time points such as screening date, randomization, study visits, and follow-up appointments, and converts them into a uniform time grid. Audio data from interviews with study participants is transcribed using a speech recognition system optimized for scientific purposes, which also reliably recognizes subject-specific terminology.

[0372] For event detection, a specialized model is used that is trained to identify study endpoints, adverse events, and therapy-relevant outcomes. Events are classified according to study-specific categories such as "Primary Endpoint," "Secondary Endpoint," "Safety Event," and "Protocol Deviation," with the categorization being adaptable to the respective study protocol.

[0373] Events are clustered according to study-specific criteria, taking into account temporal, causal, and thematic relationships. A specialized biostatistical model is used for causal analysis, identifying potential confounders and quantifying the causal effects of interventions on outcomes. The structured data representation is implemented as a semantic research graph, capable of mapping complex relationships between clinical, genomic, and phenotypic data.

[0374] The data is enriched by linking it to scientific databases such as PubMed, clinical trial registries, and specialized ontologies like Gene Ontology or Disease Ontology. The input format for the generative language model is a scientific template that complies with the CONSORT guidelines for clinical trial publications. A language model specifically designed for scientific texts is used to generate reports, enabling precise formulations and statistical descriptions while ensuring the anonymization and aggregation of individual patient data.

[0375] Application example 10: Interdisciplinary AI-supported reporting for complex multimorbid patients

[0376] In a center for complex diseases, the system is used for the interdisciplinary documentation of multimorbid patients with complex treatment pathways. Medical information is gathered from documentation from various departments, multiprofessional case conferences, interdisciplinary consultations, and specialized diagnostic procedures. A key challenge is the integration and harmonization of different discipline-specific terminologies and documentation standards.

[0377] Timestamp normalization takes into account parallel treatment pathways in different departments and synchronizes the temporal documentation even in distributed treatment episodes. Audio data from interdisciplinary case conferences is processed by a speech recognition system with speaker identification, which can differentiate the contributions of various specialists and assign them to the corresponding disciplines.

[0378] For event detection, a multi-expert model is used that combines domain-specific sub-models for various medical specialties and identifies interdisciplinary relationships. The classification of events is performed using a cross-disciplinary scheme that includes both discipline-specific and integrative categories and considers interactions between diseases of different organ systems.

[0379] The clustering of events is performed using a multi-layered hierarchical method that captures both organ-specific and systemic relationships and identifies synergies and antagonisms between treatments. A complex multivariate causal model is used for causal analysis, which can also depict indirect pathways of action, mediator effects, and interactions between different diseases and therapies.

[0380] The structured data representation is implemented as a multidimensional knowledge graph that depicts different disciplinary perspectives as parallel levels and explicitly models their points of connection. Data enrichment is achieved through the integration of multiple knowledge bases from various disciplines, as well as specialized databases on multimorbidity and polypharmacy. The input format for the generative language model is a modular format that combines discipline-specific and integrative sections. The language model used generates a coherent overall report that integrates the various disciplines while simultaneously providing relevant detailed information for each discipline.

[0381] Further examples:

[0382] 1. A computer-implemented method (100) for generating medical reports, in particular medical letters, wherein the method comprises at least the following steps:

[0383] Receiving (102) medical information about a patient, wherein the medical information received is formulated in natural language bullet points; and

[0384] Generating (104) natural language flowing text that reflects the medical information received for a medical report using a trained generative language model.

[0385] 2. The procedure according to Example 1, wherein the generative language model is configured such that the generated natural language flow text reflects all and only the received medical information.

[0386] 3. The procedure according to any of Examples 1 or 2, wherein the generative language model is configured such that the natural language flowing text deterministically reflects the received information.

[0387] 4. The method according to any of the preceding examples, wherein the step of receiving medical information comprises: receiving at least one user input in a graphical user interface, in particular by medical professionals during a patient visit; and / or

[0388] Receiving at least one electronic data record from a hospital information system, wherein the at least one electronic data record is in particular “Health Level 7” compliant and / or “Fast Healthcare Interoperability Resources” compliant; and / or wherein the procedure further comprises the following step: generating the medical report as an electronic file containing the generated natural language flow text.

[0389] 5. The procedure according to any of the preceding examples, wherein the medical information received includes at least one piece of medical information belonging to one of the following categories: gender, leading symptom, diagnoses, treatments, medications, findings, and progress documentation.

[0390] 6. The procedure according to the preceding Example 5, which further comprises the following step:

[0391] Categorizing the received medical information into categories, preferably using a classifier.

[0392] 7. The procedure according to any of the preceding examples, which further comprises the following step:

[0393] Generating an input encoding for the generative language model based on the received medical information, wherein the input encoding encodes the medical information in such a way that a maximum input length of the generative language model is taken into account.

[0394] 8. The method according to Example 5 or 6, each combined with Example 7, wherein the input encoding is generated such that each piece of medical information in the input encoding is assigned to a predefined category abbreviation indicating the respective category; wherein the input encoding is generated such that the input encoding contains multiple pieces of medical information of the same category for at least one of the categories, wherein the category abbreviation is contained only once in the input encoding; wherein, optionally, the input encoding is generated such that the multiple pieces of medical information of the same category are separated by a predefined separator abbreviation.

[0395] 9. The method according to any of the preceding examples 7-8, wherein the input encoding is generated such that at least one piece of medical information in the input encoding is assigned a timestamp, in particular after the medical information and separated by a predetermined separator abbreviation. 10. The method according to any of the preceding examples 7-9, wherein the category abbreviation and / or the separator abbreviation contains exactly one character.

[0396] 11. A generative language model for generating medical reports, in particular medical letters, wherein the generative language model is configured to generate natural language flow text for a medical report, wherein the generated natural language flow text reflects medical information about a patient, the medical information being formulated in natural language bullet points; wherein, optionally, the generative language model is configured for use in procedure (100) according to any of the preceding examples 1-10.

[0397] 12. An input encoding data structure for the generative language model according to Example 11, wherein the input encoding is configured according to any of the preceding Examples 6-10.

[0398] 13. A procedure (200) for training the generative language model according to Example 11, wherein the procedure includes at least the following steps:

[0399] Providing (202) training datasets, each comprising: an input encoding data structure according to Example 12; a natural language flow text for a medical report; and training (204) the generative language model by supervised learning with the

[0400] Training data sets.

[0401] 14. A data processing device comprising means for carrying out the method according to any of Examples 1-10 or 13.

[0402] 15. A computer program or a computer-readable medium on which a computer program is stored, wherein the computer program includes instructions which, when the program is executed by a computer, cause the computer to perform the procedure according to any of Examples 1-10.

[0403] The term “and / or” used here includes all combinations of one or more of the listed aspects and can be abbreviated with “ / ”.

[0404] Although some aspects related to a device have been described, it is clear that these aspects also constitute a description of the corresponding process, where a block or device corresponds to a process step or a feature of a process step. Similarly, aspects described in connection with a process step also constitute a description of a corresponding block, element, or feature of a corresponding device.

[0405] Embodiments of the present disclosure can be implemented on a computer system. The computer system can be a local computing device (e.g., a personal computer, laptop, tablet computer, or mobile phone) with one or more processors and one or more memory devices, or a distributed computing system (e.g., a cloud computing system with one or more processors and one or more memory devices distributed across different locations, such as a local client and / or one or more remote server farms and / or data centers). The computer system can comprise any circuit or combination of circuits. In one embodiment, the computer system can comprise one or more processors, which can be of any type. The term "processor" as used herein can refer to any type of computing circuit, e.g.,a microprocessor, a microcontroller, a CISC (Complex Instruction Set Computing) microprocessor, a RISC (Reduced Instruction Set Computing) microprocessor, a VLIW (Very Long Instruction Word) microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a multi-core processor, an FPGA (Field Programmable Gate Array), or any other type of processor or processing circuit. Other types of circuitry that may be included in the computer system could be a custom-designed circuit, an application-specific integrated circuit (ASIC), or similar, such as one or more circuits (e.g., a communications circuit) for use in wireless devices like mobile phones, tablet computers, laptop computers, two-way radios, and similar electronic systems.The computer system may include one or more storage devices, which may comprise one or more storage elements suitable for the specific application, such as main memory in the form of random-access memory (RAM), one or more hard disks, and / or one or more drives that handle removable media such as compact discs (CDs), flash memory cards, digital video discs (DVDs), and the like. The computer system may also include a display device, one or more speakers, and a keyboard and / or a control device, which may include a mouse, trackball, touchscreen, speech recognition device, or any other device that enables a system user to input information into and receive information from the computer system.

[0406] Some or all of the process steps can be performed by (or using) a hardware device, such as a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the key process steps can be performed by such a device.

[0407] Depending on specific implementation requirements, embodiments of the present disclosure can be implemented in hardware or in software. The implementation can be carried out using a non-transferable storage medium such as a digital storage medium, for example, a floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, on which electronically readable control signals are stored that interact (or can interact) with a programmable computer system to execute the respective method. Therefore, the digital storage medium can be computer-readable.

[0408] Some embodiments according to the present disclosure include a data carrier with electronically readable control signals that can interact with a programmable computer system to perform one of the methods described herein. In general, embodiments of the present disclosure can be implemented as a computer program product with program code, wherein the program code serves to execute one of the methods when the computer program product is running on a computer. The program code can, for example, be stored on a machine-readable medium.

[0409] Other embodiments include the computer program for carrying out one of the methods described herein, which is stored on a machine-readable medium.

[0410] In other words, an embodiment of the present disclosure is therefore a computer program with program code for carrying out one of the methods described herein when the computer program runs on a computer.

[0411] Another embodiment of the present disclosure is therefore a storage medium (or a data carrier or a computer-readable medium) on which the computer program for carrying out one of the methods described herein is stored when executed by a processor. The data carrier, the digital storage medium, or the recorded medium is typically tangible and / or non-transferable. Another embodiment of the present disclosure is a device as described herein, comprising a processor and the storage medium.

[0412] Another embodiment of the present disclosure is therefore a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or sequence of signals can, for example, be configured to be transmitted via a data communication link, e.g., via the Internet.

[0413] Another embodiment comprises a processing means, e.g. a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0414] Another embodiment comprises a computer on which the computer program for carrying out one of the methods described herein is installed.

[0415] Another embodiment according to the present disclosure comprises a device or system configured to transmit a computer program for carrying out one of the methods described herein to a receiver (e.g., electronically or optically). The receiver may be, for example, a computer, a mobile device, a storage device, or the like. The device or system may, for example, include a file server for transmitting the computer program to the receiver.

[0416] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.

Claims

REQUIREMENTS 1. A computer-implemented method for generating medical reports, in particular medical letters, wherein the method comprises at least the following steps: a) capturing medical information about a patient from at least two different sources of an electronic health record, wherein the received medical information is formulated at least partially in natural language bullet points; b) identifying medical units contained in the captured medical information and classifying each identified medical unit according to a medical event type by computer-aided pattern recognition; c) forming event groups by computer-aided grouping of medical units that are related at least with regard to temporal reference, treatment pathway or thematic relevance;d) Deriving relationship information that represents at least one causal relationship or correlation between elements of the event groups, and storing the event group elements and the relationship information in a structured data representation; e) Enriching at least one element of the structured data representation by linking it to data records from at least one external medical knowledge base; f) Converting the enriched structured data representation into an input format that can be processed by a generative language model; and g) Processing the input format with the generative language model to generate natural language flow text that reflects the captured medical information for a medical report.

2. The method according to claim 1, wherein the sources according to step a) comprise at least one source from the group consisting of clinical reports, laboratory results, imaging reports, prescriptions, diagnostic lists, procedure documentation or progress reports.

3. The method according to one of the preceding claims, wherein, according to step a), medical information is collected from at least three different sources.

4. The method according to one of the preceding claims, wherein the acquisition according to step a) includes normalizing timestamps into a uniform time format.

5. The method according to any of the preceding claims, wherein the pattern recognition according to step b) comprises a named entity recognition model.

6. The method according to claim 5, wherein the named entity recognition model is transformer-based.

7. The method according to claim 6, wherein the transformer-based named entity recognition model is fine-tuned using a domain-specific medical text corpus.

8. The method according to any of the preceding claims, wherein the medical event type according to step b) is assigned to at least one of the types diagnosis, drug administration, procedure, laboratory value, imaging finding or progress report.

9. The method according to one of the preceding claims, wherein the formation of the event groups according to step c) is carried out by means of hierarchical clustering.

10. The method according to one of the preceding claims, wherein the formation of the event groups according to step c) is carried out by means of density-based clustering.

11. The method according to one of the preceding claims, wherein, when forming the event groups, a maximum permissible time interval is specified between medical units to be included in the same group.

12. The method according to one of the preceding claims, wherein the derivation of the relationship information according to step d) is carried out using a structural causal model.

13. The method according to claim 12, wherein the structural causal model is provided as a directed acyclic graph.

14. The method according to one of the preceding claims, wherein the structured data representation according to step d) is stored as a knowledge graph.

15. The method according to claim 14, wherein the knowledge graph is persisted as RDF triples.

16. The method according to any of the preceding claims, wherein the at least one external medical knowledge base according to step e) is the Unified Medical Language System (UMLS).

17. The method according to any of the preceding claims, wherein the at least one external medical knowledge base according to step e) is SNOMED CT, ICD, ICD10GM, ICD11 GM and / or OPS.

18. The method according to one of the preceding claims, wherein the linking according to step e) is performed using a vector similarity measure based on embedded representations.

19. The method according to one of the preceding claims, wherein the input format according to step f) is structured as a JSON document.

20. The method according to claim 19, wherein the JSON document comprises a prompt template with role description, instruction part and context part.

21. The method according to one of the preceding claims, wherein the generative language model according to step g) is an autoregressive transformer.

22. The method according to claim 21, wherein the generative language model is fine-tuned by Reinforcement Learning from Human Feedback with regard to medical summaries.

23. The method according to one of the preceding claims, wherein during step g) a defined maximum token length is specified for output.

24. The method according to one of the preceding claims, wherein the generated medical report is automatically written back into the patient's electronic health record.

25. The method according to any of the preceding claims, further comprising providing a user interface through which medical professionals can edit and release the generated report.

26. The method according to one of the preceding claims, wherein all patient-identifying information is pseudonymized or anonymized before data is transmitted to the generative language model.

27. The method according to one of the preceding claims, wherein all calculation steps are performed in a local IT infrastructure operated within the healthcare facility.

28. The method according to any of the preceding claims, wherein step g) is performed in a hardware-based security environment with protected memory.

29. The method according to any of the preceding claims, wherein the generative language model is configured such that the generated natural language flow text reflects all and only the received medical information.

30. The method according to any of the preceding claims, wherein the generative language model is configured such that the natural language flowing text reflects the received information.

31. The method according to any of the preceding claims, wherein the step comprises receiving medical information: Receiving at least one user input in a graphical user interface, particularly by medical professionals during a patient visit; and / or Receiving at least one electronic data record from a hospital information system, wherein the at least one electronic data record is in particular “Health Level 7” compliant and / or “Fast Healthcare Interoperability Resources” compliant; and / or wherein the procedure further comprises the following step: Generating the medical report as an electronic file containing the generated natural language flow text.

32. The method according to any of the preceding claims, wherein the medical information received includes at least one piece of medical information belonging to one of the following categories: gender, leading symptom, diagnoses, treatments, medications, findings and progress documentation.

33. The method according to the preceding claim 32, which further comprises the following step: Categorizing the received medical information into categories, preferably using a classifier.

34. The method according to any of the preceding claims, further comprising the following step: Generating an input encoding for the generative language model based on the received medical information, wherein the input encoding encodes the medical information in such a way that a maximum input length of the generative language model is taken into account.

35. The method according to claim 32 or 33, each combined with claim 34, wherein the input encoding is generated such that each piece of medical information in the input encoding is assigned to a predetermined category abbreviation indicating the respective category; wherein the input encoding is generated such that the input encoding contains multiple pieces of medical information of the same category for at least one of the categories, wherein the category abbreviation is contained only once in the input encoding; wherein, optionally, the input encoding is generated such that the multiple pieces of medical information of the same category are separated by a predetermined separator abbreviation.

36. The method according to any one of the preceding claims 34-35, wherein the input encoding is generated such that at least one piece of medical information in the input encoding is assigned a timestamp, in particular after the medical information and separated by a predetermined separator abbreviation.

37. The method according to any one of the preceding claims 34-36, wherein the category abbreviation and / or the separator abbreviation contains exactly one character.

38. A generative language model for generating medical reports, in particular medical letters, wherein the generative language model is configured to generate natural language flow text for a medical report, wherein the generated natural language flow text reflects medical information about a patient, the medical information being formulated in natural language bullet points; wherein, optionally, the generative language model is configured for use in the method according to any of the preceding claims 1-37.

39. An input encoding data structure for the generative language model according to claim 38, wherein the input encoding is configured according to any one of the preceding claims 33-37.

40. A method for training the generative language model according to claim 39, wherein the method comprises at least the following steps: Providing training datasets, each comprising: an input encoding data structure according to claim 39; a natural language flow text for a medical report; and training the generative language model by supervised learning with the Training data sets.

41. A data processing device comprising means for carrying out the method according to any one of claims 1-37 and / or 40.

42. A computer program or a computer-readable medium on which a computer program is stored, wherein the computer program comprises instructions which, when the program is executed by a computer, cause the computer to execute the method according to any one of claims 1-37 and / or 40.

Citation Information

Cited By

  • Intelligent medical record generation method and system

    CN121506525A

  • Medical text information processing method and device, electronic equipment and storage medium

    CN121809410A