A clinical pathology reporting system and method
The clinical pathology reporting system addresses format and accuracy issues in pathology reports by using a machine learning model that allows user correction and adaptation, enhancing data extraction and report generation accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BIO CONCEPTS PTY LTD
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
Interpreting pathology test reports is challenging due to varying formats, complex content, and the reliance on accurate data, which can lead to errors and misdiagnoses, especially when reports are in unstructured formats like PDFs.
A clinical pathology reporting system using a machine learning model to extract data, allow user correction, and continually improve by incorporating corrected data for enhanced accuracy and adaptability.
The system ensures accurate extraction and generation of healthcare reports, reducing errors and enabling informed decision-making by continuously learning from user corrections.
Smart Images

Figure AU2025051182_23042026_PF_FP_ABST
Abstract
Description
A CLINICAL PATHOLOGY REPORTING SYSTEM AND METHODTECHNICAL FIELD
[0001] The present invention relates to clinical pathology reporting. In particular, the present invention relates to interpretation of pathology test reports.BACKGROUND
[0002] Pathology test reports play a crucial role in healthcare decision-making, guiding physicians in the formulation of treatment plans, monitoring disease progression, and assessing responses to therapy. These reports usually contain information relating to one or more diagnostic measurements, e.g. blood tests, genetic tests, clinical imaging, or the like. Depending on the type of diagnostic measurement, different information is provided. Overall, however, they generally provide detailed pathology data which may be interpreted by a health professional to evaluate various aspects of an individual's health, facilitating the detection, diagnosis, and management of diseases and conditions.
[0003] However, effectively interpreting pathology test reports is not always straightforward. For example, each type of pathology report is generally presented in a different way. Furthermore, even the same types of pathology report but from different pathology clinics are generally presented in different ways.
[0004] Furthermore, functional pathology is a continuously evolving field. As a result, health professionals must actively work to keep up with the latest information to be able to accurately interpret pathology data according to best practice. Continued education in this regard is often time-consuming due to several factors. Firstly, it is very time consuming to process the vast volumes of information and research in the various fields required to stay abreast of developments. Additionally, pathology concepts and techniques are generally complex and require significant time to properly comprehend and be able to apply in practice. Moreover, health professionals often have demanding schedules, with limited time available for additional learning activities. As a result, the process becomes not only time-consuming but also inherently challenging and error prone.
[0005] Clinical decision software exists as a tool to assist a health professional in medical diagnosis and treatment. Research has demonstrated the clinical decision software is able to enhance the precision and efficacy of clinical decision-making processes. This may in turn mitigate errors and optimise patient care outcomes. For instance, clinical decision software may aid in identifying analytes that fall outside acceptable parameters, and provide informationrelating to potential conditions or underlying causes to facilitate informed clinical decisionmaking.
[0006] One problem with clinical decision software of the prior art is that clinical decision software depends heavily on the accuracy of the pathology data. Any inaccuracies or inconsistencies in the pathology data could lead to erroneous conclusions or misdiagnoses. This is particularly problematic when pathology test reports are provided as documents (e.g. in PDF format) or images, rather than structured data in a common format.
[0007] In this regard, extracting accurate pathology data from pathology test reports is complex, as pathology test reports from various laboratories often vary significantly in their formatting, reference ranges, and overall presentation. As a result, automatic extraction of pathology data from pathology test reports is prone to error.
[0008] As a result, there is a need for improved clinical pathology reporting systems and methods.
[0009] It will be clearly understood that, if a prior art publication is referred to herein, this reference does not constitute an admission that the publication forms part of the common general knowledge in the art in Australia or in any other country.SUMMARY OF INVENTION
[0010] Embodiments of the present invention relate to a clinical pathology reporting system and method, which may at least partially overcome at least one of the abovementioned disadvantages or provide the consumer with a useful or commercial choice.
[0011] According to a first aspect of the present invention, there is provided a clinical pathology reporting system for generating a healthcare report from a pathology test report, including: at least one processing device configured to: extract pathology data from the pathology test report using a machine learning model; provide, on a data interface, and to a user, the pathology data for review by the user; receive, on the data interface, and from the user, corrected pathology data; provide at least part of the corrected pathology data to the machine learning model to further train the machine learning model; retrieve auxiliary content according to the corrected pathology data; andgenerate the healthcare report from the corrected pathology data and the auxiliary content.
[0012] Advantageously, by providing the pathology data for review, receiving corrected pathology data, and providing corrected pathology data to the machine learning model, the system is able to continually improve its extraction of pathology data from pathology test reports. This in turn enables the system to adapt to different types of reports over time, including when reports change over time.
[0013] Furthermore, by providing the pathology data for review, and receiving and using corrected pathology data, the system is able to maintain accuracy, even when the machine learning model is unable to accurately extract the pathology data. This in turn enables more accurate healthcare reports to be generated, and reduces the risk of erroneous conclusions or misdiagnoses.
[0014] Yet further again, by retrieving auxiliary content according to the corrected pathology data, targeted information is able to be provided in the healthcare report. This targeted information may offer insight into the properties, functions, and clinical significance of various pathology test results, for example, thereby facilitating informed decision-making.
[0015] The pathology test report comprises a medical report that contains pathology data. The pathology test report may contain diagnostic measurements relating to a patient. The diagnostic measurements may relate to a specimen taken from the patient (e.g. a blood or urine test), or clinical imaging, for example. Any relevant information may be included in a pathology test report and in any suitable form.
[0016] The pathology test report may contain patient information in association with the pathology data.
[0017] The patient information may include patient name, patient date of birth, patient address, and / or any other suitable patient information.
[0018] The at least one processing device may be configured to extract the patient data from the pathology test report using the machine learning model. The at least one processing device may associate the patient data with the pathology data and / or the corrected pathology data.
[0019] The pathology data may include measurement data (e.g. red blood cell count). The measurement data may be associated with one or more reference values. The measurement data may be colour coded.
[0020] The pathology data may contain specimen analysis data, relating to one or more characteristics of a specimen.
[0021] The pathology test report may further include interpretation data, relating to an interpretation of the pathology data. The interpretation data may include categorisation according to one or more pre-defined categories (e.g. high, medium or low).
[0022] The pathology test report may further include comments, as well as possible recommendations for further evaluation, treatment, or management. For example, the comments might suggest additional diagnostic tests such as blood work or imaging studies, refer the patient to a specialist for a more detailed examination, propose adjustments to current medications or therapy regimens, or outline a follow-up plan to monitor the patient’s progress and response to treatment. Additionally, the recommendations could include lifestyle modifications, such as changes in diet, exercise, supplementation, medication, or stress management, aimed at improving overall health outcomes.
[0023] The pathology test report may be provided in a format of various possible formats, reflecting the diverse methods of documentation in medical reporting.
[0024] The at least one processing device may be configured to receive the pathology test report on the data interface. The pathology test report may be uploaded by a user for processing. The pathology test report may be received directly from a clinic or external system.
[0025] The at least one processing device may be configured to receive a plurality of pathology test reports. The plurality of pathology test reports may be individually processed to generate separate healthcare reports. The plurality of pathology test reports may be received over time, and processed as they are received.
[0026] The pathology test report may include the pathology data in image form.
[0027] The pathology test report may comprise an electronic document. The pathology test report may be defined in any suitable format, including in Portable Document Format (PDF), Microsoft Word, Microsoft Excel or the like.
[0028] The pathology test report may comprise or include an image. The image may comprise a scanned image of a report document. Alternatively, the image may comprise a computer generated image of the report.
[0029] In order to support informed decision-making, the healthcare report is generated from the corrected pathology data and the auxiliary content.
[0030] The healthcare report may be automatically generated. In other embodiments, however, the healthcare report may be generated according to one or more user inputs.
[0031] As indicated, the clinical pathology reporting system has at least one processing device configured to extract pathology data from the pathology test report using a machine learning model.
[0032] The at least one processing device may be any suitable device or combination of devices capable of executing program code or taking instructions. In general, the at least one processing device may be a computing device, such as a server, workstation, desktop, laptop, tablet, smartphone, smartwatch, or PDA, or a combination of a plurality of such computing devices. Similarly, the at least one processing device may comprise a virtual or distributed computing device.
[0033] The at least one processing device may be programmed to carry out various operations to provide functionality to the system. The at least one processing device may be configured to run software or an app to perform, coordinate and manage tasks of the clinical pathology reporting system.
[0034] The at least one processing device may be configured to pre-process the pathology test report prior to extracting the pathology data. Pre-processing the pathology test report may include noise reduction, contrast enhancement, and / or image normalisation, e.g. if the pathology test report or part thereof is provided in image form. The at least one processing device may analyse the image and apply one or more corrective measures based upon the analysis. For example, the at least one processing device may identify reference points or lines within the image, such as text lines or edges, to determine the degree and direction of skew. Based on this analysis, corrective transformations may be applied to the image to eliminate the skewing and to ensure that the text appears upright and straight.
[0035] The at least one processing device may be configured to select the machine learning model from a plurality of machine learning models. The at least one processing device may be configured to select the machine learning model according to a type of the pathology test report, or any other suitable criteria.
[0036] The at least one processing device may be configured to select the machine learning model according to accuracy, e.g. with reference to the type of the pathology test report. The at least one processing device may be configured to select the machine learning model according to scalability requirements, e.g. to ensure the selected model can process the pathology test report and others.
[0037] The at least one processing device may be configured to train the machine learning model using a training dataset. The training dataset may include examples of pathology test reports, and associated pathology data.
[0038] The at least one processing device may be configured to clean the training dataset prior to training the machine learning model. For example, the at least one processing device may remove duplicate data, identify outlier data and then remove or transform outliers if they are deemed erroneous or if they could unduly influence the model. The at least one processing device may divide a larger dataset into the training dataset and a test dataset, to train the machine learning model on the training dataset and evaluate the performance of the model using the test dataset.
[0039] In some embodiments, the machine learning model is a supervised learning model, such as a support vector machine (SVM), a decision tree, or a neural network for structured data extraction.
[0040] Preferably, the machine learning model includes an Artificial Neural Network based Optical Character Recognition (ANN-based OCR) model configured to extract the pathology data in the form of text from the pathology test report. In such case, the ANN-based OCR model may be able to perceive and recognise a character based on its topological features such as shape, symmetry, closed or open areas, and number of pixels.
[0041] The machine learning model may be further configured to automatically recognise and extract relevant information from the pathology test report. As an illustrative example, the machine learning model may be configured to recognise and extract the pathology data, but ignore other non-relevant text, such as branding or similar.
[0042] The ANN-based OCR model may comprise a neural network consisting of multiple layers of interconnected neurons. The neural network may be a convolutional neural network (CNN), a recurrent neural network (RNN) or a combination of both.
[0043] The machine learning model may be trained on a dataset comprising pathology test reports with associated labels, where the labels correspond to the pathology data and other text in the pathology test report.
[0044] In the case of an ANN-based OCR model, images or documents with associated labels, where the labels correspond to the characters or text present in the images or documents, may be used for training. The images or documents may comprise pathology test reports, or other text.
[0045] During the training, the ANN-based OCR model may learn to associate specific image features with corresponding characters, optimising its parameters through techniques such as backpropagation and gradient descent.
[0046] The machine learning model may comprise a combination of several sub-models, and each sub-model may be separately trained. As an illustrative example, one sub-model may comprise an ANN-based OCR model, trained on a wide range of text, and another sub-model may comprise a pathology test report specific model, trained on pathology test reports.
[0047] The system may identify any suitable features from the pathology test report to extract the pathology data therefrom. These features may include edges, shapes, or other patterns that are characteristic of different characters. It is envisaged that feature collection serves to extract properties that can identify a character uniquely, or to extract properties that can differentiate between similar characters.
[0048] The at least one processing device may analyse the extracted pathology data, and refine the extracted pathology data. The at least one processing device may employ any suitable postprocessing techniques to analyse and refine the extracted pathology data, such as language modelling, spell checking, or context-based correction, ensuring that the results are meaningful and error-free.
[0049] As indicated, the at least one processing device provides, on the data interface, and to the user (e.g. a general practitioner), the pathology data for review by the user. The data interface may be any suitable interface capable of communicating the pathology data.
[0050] The user is then able to make any necessary corrections, thereby creating the corrected pathology data. This enables the extracted pathology data to be reviewed and corrected remotely, e.g. on a user computing device that is remote to the at least one processing device.
[0051] The at least one processing device may be configured to generate a user interface to display the extracted pathology data to the user. The user interface may be user-friendly, and may include various tools and options enabling the user to view, edit, and / or annotate the extracted pathology data. As such, the user interface may provide a way for the user to perform a thorough review of the extracted pathology data and ensure the accuracy of the extracted pathology data.
[0052] The user interface may display the extracted pathology data in a structured manner, making it easy for the user to navigate and comprehend information from the extracted pathologydata.
[0053] For example, the extracted pathology data and / or the extracted patient data may be arranged into multiple elements, such as patient information, specimen information, test details, results, interpretations / comments, and / or recommendations.
[0054] The extracted pathology data and / or the extracted patient data within each element may be presented in a tabular format. The patient information may comprise fields of patient ID, name, date of birth, gender, medical record no and so on. The specimen information may comprise fields of specimen ID, collection data, specimen type, source and so on. The test details may comprise fields of test ID, test name, test method, test date and so on. The results may comprise fields of result ID, parameter, value, units, reference range and so on. The interpretations / comments may comprise fields of comment ID and comment. The recommendations may comprise fields of recommendation ID and recommendation.
[0055] The user interface is configurable to display the extracted pathology data and the pathology test report side-by-side. This enables the user to compare the extracted pathology data and the pathology test report in a manner that enables the user to quickly identify any discrepancies or inconsistencies.
[0056] In some embodiments, the user interface may provide interactive tools for viewing, editing, and annotating the extracted pathology data. The user may, using the user interface, view the extracted pathology data, cross-reference the extracted pathology data with the pathology test report, including text, images, and numerical data, to validate that all relevant information has been accurately captured and aligned with established medical standards and guidelines. This functionality allows the user to make necessary corrections or add pathology data directly within the interface. For instance, the user interface may display the pathology data as editable text fields to allow the user to rectify misinterpretations or fill in missing information, and / or drag-and-drop features to help realign data elements.
[0057] The user interface may include a search element, in which the user may enter a search query. Upon entering the search query, the user interface may select data according to the search query from the pathology data, and display the selected data.
[0058] Similarly, the user interface may include a filter element, with which the user may filter the pathology data for display.
[0059] The search element and filter element enable the user to quickly locate specific data points or sections within the pathology data, improving efficiency of the review of the pathologydata.
[0060] In some embodiments, the system may be coupled to a health record or patient management system. The system may be configured to retrieve the pathology test report, patient data, or any other suitable information from the health record or patient management system. Such configuration enables seamless access to relevant patient information, and may facilitate workflow integration for the user.
[0061] It is envisaged that the system may have robust security measures, such as encryption and user authentication, ensuring the confidentiality and integrity of patient data accessed through the data interface.
[0062] The user interface may be responsive, and adapt to different screen sizes and orientations to provide a good viewing experience on different devices. Whether accessed on a large desktop monitor in a clinical setting or on a handheld device in the field, the user interface may be capable of maintaining its functionality and ease of use, allowing for efficient and accurate data validation.
[0063] As indicated, the at least one processing device is configured to receive, on the data interface, and from the user, corrected pathology data. The corrected pathology data may be uploaded to the at least one processing device automatically from the user interface when the pathology data is corrected, or in response to the user finalising review of the pathology data using the user interface, for example.
[0064] In particular, if the user identifies any errors or inconsistencies in the extracted pathology data, the user interface may allow the user to correct the extracted pathology data. The user interface may then be configured to upload, to the at least one processing device, and on the data interface, the corrected pathology data.
[0065] In some embodiments, the at least one processing device may analyse the extracted pathology data to evaluate the quality of the extracted pathology data. In such case, the analysis may focus on identifying potential inconsistencies between the extracted pathology data and the original pathology report. This may be performed according to patient history, previous test results, diagnostic criteria, or any other suitable factors.
[0066] In other embodiments, the system may propose, using the user interface, potential corrections to the extracted pathology data based on patterns or data structures. The user interface may allow the user to accept or reject the proposed potential corrections through interaction with the user interface (e.g. by selecting an “accept” or “reject” button). In such case,the system may include validation rules including one or more valid analyte names and aliases, values (upper and lower bounds) as well as units that can be set and applied.
[0067] In some embodiments, the system may be configured to automatically correct errors in the data. For example, if the extraction process mistakenly identifies a unit as "mnol", which is an invalid or misspelled unit, the system may use a Levenshtein distance-based algorithm to find the closest valid unit based on similarity. In this scenario, the system may identify "mmol" (millimoles) as the closest valid unit to "mnol" and automatically substitute it. In other embodiments, the system may propose these corrections using the user interface, wherein the user may accept or reject the proposed corrections.
[0068] In order to maintain data integrity and for future reference, the system may generate an audit trail documenting the original data and the corrections applied. The audit trail may be in any suitable form, but typically includes timestamped entries, user information, original data, corrected data, and optionally, details such as the nature of the error, correction rationale, and linked references. The audit trail may be provided to the machine learning model to further improve the machine learning model.
[0069] In some embodiments, the system is configured to quantify the difference between the extracted pathology data and the corrected pathology data. This may involve calculating the magnitude of the difference, such as the difference (relative or absolute) between the extracted and corrected pathology data, or the number of characters or digits that differ between the two versions of the data. The system may further flag data points with significant differences or high delta values for further manual review, if necessary. This may be done on the user interface, for example, using colour coding, highlighting or similar.
[0070] In some embodiments, the system may generate a delta report including details of the differences between the extracted pathology data and the corrected pathology data. The delta report may include details such as the location of discrepancies, the nature of the differences, the magnitude of the delta for each data point, and any relevant information.
[0071] As indicated, the at least one processing device may provide at least part of the corrected pathology data to the machine learning model to further improve the machine learning model. The at least part of the corrected pathology data may comprise a delta report, for example.
[0072] The at least one processing device may provide the at least part of the corrected pathology data to the machine learning model for use as training data, to further train the machine learning model. In such case, the machine learning model may be trained using thepathology test report and the at least part of the corrected pathology data.
[0073] By learning from user corrections, the machine learning model may adapt and refine its extraction process to minimise errors and discrepancies over time. It is envisaged that the machine learning model may have been trained on a labelled dataset of labelled images or documents, where the labels correspond to the characters or text present in the images or documents. During training, the machine learning model adjusts its internal parameters to minimise a predefined loss function, typically categorical cross-entropy, which measures the disparity between predicted and actual characters.
[0074] In some embodiments, the system may create a new dataset including the at least part of the corrected pathology data. Then the system may retrain the machine learning model using the newly created dataset.
[0075] In preferred embodiments, the system may incorporate the at least part of the corrected pathology data into the existing training dataset, augmenting it with images and the corrected corresponding character labels. The system may then retrain the machine learning model with the updated dataset, allowing the model to learn from its misinterpretation and improve its accuracy. The retraining process may continue until a convergence criterion is met, such as achieving a desired level of accuracy on a test set or stabilising the loss function.
[0076] In such preferred embodiments, it is envisaged that the system may iteratively augment the training dataset by incorporating corrected pathology data over time.
[0077] In some embodiments, each set of pathology data in the training dataset may be assigned a timestamp when it is inserted into the dataset. The system may periodically remove the oldest data based on its timestamp. This process may help ensure that the training dataset remains manageable, that the training data is up to date and representative of recent pathology test reports, and that the retraining times are kept within an acceptable limit, preventing significant delays in system operations.
[0078] The system may continuously monitor the performance of the machine learning model, by analysing the corrected pathology data. The significance of any corrections may be assessed, and appropriate actions to take in response may be determined.
[0079] In some embodiments, the system may analyse the corrected pathology data to identify one or more conditions. The conditions may comprise medical conditions.
[0080] The system may include a predefined list of medical conditions.
[0081] Each condition may be associated with at least one marker and an associated state.
[0082] The markers may relate to analytes able to be measured. The markers may relate to substances like proteins, enzymes, hormones, or other measurable biological components. Each marker may have a measurable value derived from the corrected pathology data.
[0083] The state may comprise a threshold. The state may comprise a range. The state may comprise one or more other criteria, including a combination of thresholds.
[0084] The conditions may be identified at least in part according to the markers and the associated states.
[0085] The markers may be associated with more than one state. At least one state may be at least partly indicative of the condition.
[0086] Different states may be provided for different ages or demographics. Different states may be provided for different characteristics of the patient, including use of medication. The states may be dynamically generated according to characteristics of the patient.
[0087] Furthermore, different states may be used to indicate different levels of risk or patient health.
[0088] For example, one state may define an optimal range, representing a desired or healthy level fora typical individual. Additionally, a reference range, which is based on population data, may be used to contextualise data relating to the marker.
[0089] Furthermore, a state may relate to an alarm range, wherein an alert may be triggered according to the alarm range. Such state is useful in indicating the need for immediate medical attention or further investigation.
[0090] The states may represent various levels of deviation from normal, or various levels of risk associated with the condition. For example, a state "alarm_high" or "alarmjow" may indicate very high or low levels, signalling a potentially dangerous abnormality. Similarly, a state "high" or " low" may indicate high or low levels, but not to the extent of the "alarm_high" or "alarmjow" states. Finally, a state "supraoptimal" or "suboptimal" may indicate high or low levels, but not to the extent of the " high" or " low" states.
[0091] The markers may further be associated with categories. The categories may relate to the clinical relevance of the marker. The categories may include a category (e.g. “diagnostic marker”) indicating a necessary association between the marker and the condition.
[0092] The categories may include a category (e.g. “differential marker”) indicating an association between the marker and the condition, but without requiring the marker for classification. Such markers may help identify conditions when present, but need not be present to condition matching.
[0093] Finally, the categories may include a category (e.g. “consequential marker”) which are not used in condition matching, but may include data of secondary relevance to the condition.
[0094] The markers may further be associated with a flag to indicate the priority or importance of the marker in relation to the associated condition. In general, some markers are expected to receive greater attention during condition matching. The flag assigned to each marker may be defined with at least two levels, though three levels are preferred, such as "priority_required", “priority” and “non-priority”.
[0095] For example, a flag (e.g. “priority_required”) associated with a marker may indicate that data relating to the marker must be present in the pathology test report to be able to match the condition. Such flag is relevant where condition matching is not possible without data relating to the marker.
[0096] A flag (e.g. “priority”) associated with a marker may indicate that if data relating to the marker is present in the pathology test report, it must be in a predefined state for the system to identify the condition. Such flag is relevant where a marker is helpful, but not essential in identifying a condition.
[0097] In order to identify a condition, the system may apply any suitable rules based on one or more markers for the condition, and the state associated with each marker based on the measurable value from the corrected pathology data.
[0098] For example, if a condition is associated with diagnostic markers, data relating to all such markers must be present in the report and the data must correspond to a particular state to identify the condition. If the condition also includes differential markers, data relating to all “priority_required” differential markers must be present in the report and the data must correspond to a particular state to identify the condition. No data relating to any “priority” differential markers need not be present in the report, but if it is present, the data must correspond to a particular state to identify the condition.
[0099] If a condition is not associated with diagnostic markers, but is associated with differential markers, data relating to any “priority_required” differential markers must be present in the report and the data must correspond to a particular state to identify the condition. Datarelating to “priority” differential markers need not be present in the report, but if it is present, the data must correspond to a particular state to identify the condition. Data relating to “non-priority” differential markers need not be present in the report, but if it is present, the data of at least one “non-priority” differential marker must correspond to a particular state to identify the condition.
[0100] In some embodiments, the system may identify trends in pathology data. The trends may be useful in identifying potential conditions. Trends may be identified over any time period. For example, time periods may range from 1 month, 3 months, 6 months, 12 months, 24 months, or any other suitable duration.
[0101] In such embodiments, the markers may be associated with trend states, relating to changes in data overtime. Examples of trend states include an “up”, “flat”, or “down” trend state, indicating whether data associated with the marker has increased, remained stable, or decreased over time, e.g. across successive pathology test reports over a defined period. It is envisaged that the trending states enable the better analysis than one that focuses solely on a state at a single point in time.
[0102] The conditions may be identified at least in part according to a state of the marker at a single point of time, and a trend state of the marker over time. For example, data relating to a marker that is "normal" but trending "up" might be associated with a condition, whereas data of a marker in a "high" state but trending "down" might not.
[0103] The trend states may be associated with one or more trend deltas. A trend delta may be defined by a difference between data over time. The trend delta may be relative or absolute.
[0104] As an illustrative example, trend deltas may be associated with 0-5% differences, 5 - 10% differences, 10 - 20% differences, and 20+% difference. 0-5% trend deltas may be indicative of minimal change, indicating stability or slight fluctuation; 5-10% trend deltas as moderate change, suggesting clinical significance; 10-20% trend deltas as a larger fluctuation, potentially signalling ongoing processes such as disease progression or response to treatment; and trend deltas above 20% as a major shift, prompting further investigation or intervention.
[0105] The system may apply any suitable predictive model to forecast a condition over time based on markers associated with the condition, including the category, flag, and state of each marker, as well as the trend state and trend deltas calculated based on pathology data from at least two reports. The predictive model may use advanced algorithms, such as machine learning, statistical analysis, or rule-based systems, to analyse patterns in the marker data and provide informed predictions about the patient’s health trajectory in the healthcare report.
[0106] It is envisaged that if the corrected pathology data may also include user-provided comments or annotations from healthcare professionals, such as observations, interventions, treatments, or other relevant clinical notes, the system may incorporate the identified condition and predictive condition information into the user-provided comments or annotations, thereby leading to more informed and accurate predictive outcomes.
[0107] As indicated, the system retrieves auxiliary content according to the corrected pathology data. The auxiliary content may be any suitable content, such as educational or research. Preferably, the auxiliary content is selected to assist in understanding the pathology data, or impacts thereof. As an illustrative example, the auxiliary content may provide context, background information, and supplementary knowledge relevant to the interpretation and analysis of pathology test results.
[0108] The at least one processing device may be configured to determine one or more risks or attributes associated with the corrected pathology data. An example of a risk is risk for cardiovascular disease. The skilled addressee will, however, readily appreciate that any suitable risk may be identified.
[0109] The auxiliary content may be retrieved according to the determined one or more risks or attributes. For example, the auxiliary content may be categorised according to risks or attributes.
[0110] In general, the auxiliary content may contain a wide range of materials, including academic papers, textbooks, clinical guidelines, research articles, case studies, and educational modules.
[0111] The auxiliary content may be retrieved according to one or more analytes of the pathology data.
[0112] The auxiliary content may be retrieved according to one or more identified conditions of the pathology data.
[0113] The auxiliary content may also include detailed descriptions of analytes and medical conditions. These descriptions may be tailored to meet the specific needs of healthcare practitioners and researchers, offering comprehensive insights into the properties, functions, and clinical significance of various analytes and conditions.
[0114] Unlike publicly available information on the Internet, some auxiliary content may be proprietary data that is not accessible to the general public. This auxiliary content may be sourced from both internal Application Programming Interfaces (APIs) which connect toproprietary databases and systems within an organisation, and external APIs, which connect to third-party services and data providers. Together, these sources may contribute to constructing a comprehensive knowledge network.
[0115] The auxiliary content can be provided in any of a variety of formats including text, images, e-books, audio, and video, or a combination thereof.
[0116] In some embodiments, the system may convert the auxiliary content from non-text formats into text format, e.g. to make data more accessible and / or usable. This conversion may be facilitated by employing a range of machine learning techniques, including ANN-based OCR, image recognition algorithms, and Al-powered Speech to Text (STT) technology. As a result, text-based information may be extracted from diverse file formats, enabling retrieved content to be accessible and searchable in a text-based manner.
[0117] Once the auxiliary content is converted into text format, the text may undergo further processing and storage. The resulting text content may then be stored locally or remotely, and indexed using any suitable language processing techniques, such as text embedding. This indexing process may involve encoding the text data into high-dimensional vector representations, enabling efficient storage and retrieval within a specialised vector database.
[0118] By employing a combination of OCR, NLP, and / or text embedding techniques, the system may transform content from diverse sources into a unified and structured auxiliary content format. This may in turn enable information relating to the corrected pathology data to be accessed in a format that is easily searchable, navigable, and conducive to informed decisionmaking and analysis.
[0119] In some embodiments, Retrieval Augmented Generation (RAG) is used to augment a large language model (LLM) with external knowledge from the auxiliary content, content on the Internet and / or 3rdParty Data Providers. The LLM may then be used to generate information for inclusion in the healthcare report. Links to the original content from where the information was obtained may be provided in the healthcare report, e.g. to enrich the user experience and enhance the credibility of the information presented. The integration of auxiliary content with LLM allows the system to access a wealth of contextual information and leverage it to improve the quality and relevance of generated responses.
[0120] In general, the integration of the auxiliary content with the corrected pathology data is useful for providing insights in relation to the corrected pathology data. By further including links to the original content, users can verify and corroborate the information presented, ensuring transparency and trustworthiness.
[0121] As an illustrative example, a pathology test report may relate to a blood test result indicating low levels of Haemoglobin, low Red Cell Count, and low Haematocrit. After the relevant pathology data has been extracted and validated / corrected, the system may analyse the validated / corrected pathology data and determine one or more risks or attributes. This analysis may identify a potential risk of B6 deficiency anaemia from the validated / corrected pathology data.
[0122] The system may then use the LLM to generate content, or retrieve, from either local or external sources, auxiliary content, relevant to B6 deficiency anaemia as a possible condition, for inclusion in the healthcare report. Additionally, the system may use the LLM to generate additional content, or may retrieve additional auxiliary content in the form of information on Haemoglobin, Red Cell Count, and Haematocrit. The additional content may then be included in the healthcare report, together with the validated / corrected pathology data to enable the user to obtain a comprehensive understanding of the validated / corrected pathology data. This provides the user with easy access to information that assists in accurate diagnosis and treatment planning.
[0123] For example, if diabetes-related issues are detected in the corrected pathology data, the system may analyse parameters such as high fasting blood glucose and HbA1C levels. These indicators may suggest the presence of diabetes. Additionally, other auxiliary content, such as content related to elevated triglycerides, HOMA-IR, cholesterol levels, and the activation of antibodies, along with immune deficiency, can provide further insights into the specific subtype of diabetes.
[0124] It is envisaged that by using the LLM to generate the content, or by retrieving the auxiliary content, according to the corrected pathology data, the healthcare report is able to include comprehensive resources that facilitate deeper understanding, informed decisionmaking, and continuous learning in the field of pathology and related areas.
[0125] As indicated, the system generates the healthcare report from the corrected pathology data and the auxiliary content.
[0126] The healthcare report may include any suitable information relating to the pathology test report. For instance, the healthcare report may include the pathology data, potential matching conditions, and auxiliary content in the form of detailed information about specific medical conditions.
[0127] Preferably, the healthcare report is interactive. The healthcare report may include links to external data, such as auxiliary content, for example.
[0128] The healthcare report may comprise an interactive webpage.
[0129] The healthcare report may comprise a document.
[0130] The at least one processing device may further generate the healthcare report with reference to the patient information. For example, an age of the patient may be used to determine risk factors, and thereby to influence the auxiliary content.
[0131] The at least one processing device may further generate the healthcare report with reference to historical pathology data of the patient. For instance, a patient may undergo multiple pathology tests in a period, and the at least one processing device may receive multiple pathology test reports. Consequently, corrected pathology data from the multiple pathology test reports may be graphed and plotted to visualise the patient's health progression.
[0132] In such case, the multiple pathology test reports may be received at different times, processed separately, and data from same may be saved in a database with associated patient information. The data may then be retrieved from the database according to the patient information, for use with later healthcare reports.
[0133] Additionally, the healthcare report may include auxiliary content in the form of information such as dietary habits, lifestyle adjustments or integrative treatments. For example, the system may monitor changes in insulin resistance over time. The system may determine a risk for a specific subtype of diabetes based on a decline in insulin resistance, coupled with other parameters like age and life stage. The healthcare report may include such trends alongside the corrected pathology data, so the user can gain insights into the progression and management of the patient's condition.
[0134] In some embodiments, the healthcare report may include a search tool, enabling the user to enter queries relating to auxiliary content on various topics, including specific medical conditions and their relationships with different analytes. For example, the user may enter a query "ways to increase vitamin D levels" while reviewing a patient with low vitamin D levels. In response, the system may retrieve auxiliary content providing guidance on vitamin D supplementation, and also auxiliary content relating to other analytes within the corrected pathology data that may influence vitamin D levels, such as inflammation or activated immune system markers like CRP (C-reactive protein) or white blood cell count.
[0135] In this case, the healthcare report may be generated taking into account the patient's overall health condition and factors that may impact vitamin D levels. The auxiliary content may include information on how inflammation or an active immune response can lower vitamin Dlevels and suggest strategies in the healthcare report to address these underlying issues.
[0136] Preferably, the system may allow users to edit the healthcare report, e.g. to tailor it to the specific needs of their patient. For example, the system may provide an interactive user interface enabling the user can select or modify the healthcare report, and / or add auxiliary content, e.g. in the form of educational materials and lifestyle recommendations, to form a patient report. This is particularly useful in communicating health information to a patient, in view of the patient's health goals, preferences, and knowledge level.
[0137] In other embodiments, the user interface may include error identification functionality configured to identify errors in data inputted by the user. For example, the error identification function may identify misspellings and incorrect contact details and highlight the errors for correction by the user.
[0138] In some embodiments, the system may prompt users to verify details inputted by the user. For example, system may prompt users to review and confirm the accuracy of the entered information before submitting the information.
[0139] In some embodiments, the system may utilise secure channels to deliver the healthcare report to the user. This may involve encrypted communication methods to protect the confidentiality of the pathology data. Further, and / or alternatively, the system may implement authentication and authorisation procedures to ensure that only the authorised user has access to receive a healthcare report. Typically, these procedures may involve the use of secure login credentials, multi-factor authentication, or other secure verification methods.
[0140] As indicated, the clinical pathology reporting system comprises at least one processing device programmed to extract pathology data from a pathology test report, receive corrected pathology data from the user, retrieve auxiliary content and generate a healthcare report.
[0141] It is envisaged that various data is used and generated by the system during this process, including but not limited to, the original pathology test report, the extracted pathology data, the corrected pathology data, the initial and updated datasets for training the machine learning model, the retrieved auxiliary content, the generated healthcare report, and any input from the user.
[0142] Preferably, the system includes or is associated with a database configured to store some or all of the aforementioned data. The database may take various forms, such as a relational database, and is typically managed by a database management system (DBMS). ThisDBMS may ensure efficient storage, retrieval, and management of the data, enabling seamless integration with the clinical pathology reporting system.
[0143] Preferably, the pathology data is associated with patient data in the database.
[0144] The database may further include a plurality of user data, including name, address, and contact details. Each user may be assigned a unique identifier, which can be a profile name, number, or another form of identity. This unique identifier may be selected by the user or automatically assigned by the DBMS. Access to patient data, such as healthcare report and pathology test reports, may be restricted according to user. Such configuration ensures that data is only able to be accessed by appropriate users, such as the patient’s doctor.
[0145] In some embodiments, the database may incorporate regular backup and recovery procedures as part of its operational framework, to safeguard against potential data loss and mitigate the impact of unforeseen system failures.
[0146] In other embodiments, the system may engage suitable monitoring mechanisms to proactively track database performance and generate alerts for potential issues. The monitoring mechanisms may be instrumental in maintaining the overall health and efficiency of the database.
[0147] In general, the system may organise data in the database based on the data type and semantics, including patient information, specimen information, test details, results, interpretations / comments, recommendations, and other relevant data types. Preferably, the system uses different tables within the database to represent different types of the data. It is noted that each table may have well-defined columns that correspond to the attributes or characteristics of the data type.
[0148] The system may also assign semantic tags or labels to the data based on the meaning or context of the data. This may involve using standardised terminology or coding systems to categorise and classify data according to their semantics.
[0149] In some embodiments, the system may also incorporate metadata attributes into the database to provide additional context and information about the data. The metadata attributes can include details such as data source, data quality, authorship, and timestamps.
[0150] In order to navigate and query the database more effectively, the system may organise data hierarchically based on their relationships and dependencies. The system may also apply data validation rules, such as data type validation, range validation, and format validation, to ensure the integrity and accuracy of the catalogued data.
[0151] Further, the system may implement search and retrieval mechanisms to allow users to query the data within the database based on different criteria, including data types, semantics, and metadata attributes. This could involve using SQL queries, search indexes, or semantic search technologies. Additionally, the search functionality may leverage artificial intelligence (Al) technologies to further optimise the search process. For example, Al algorithms could be used to analyse user behaviour, preferences, and search patterns, providing personalised recommendations and enhancing the accuracy of results. The integration of natural language processing (NLP) could allow users to input queries in everyday language, making the system more accessible and user-friendly.
[0152] In general, the system is capable of indexing the pathology data along with their corresponding patient information, specimen information, test details, results, interpretations / comments and / or recommendations in the database to facilitate efficient filtering and retrieval of information.
[0153] Additionally, the system is also capable of indexing pathology data in the database using natural language processing (NLP) technologies, including symbolic NLP, statistical NLP or Neural NLP, among others. For example, text embedding may be applied in the current system, such that any textual data is transformed into a high-dimensional latent space, enabling intelligent and precise searching and querying capabilities in the system.
[0154] For example, some textual data is converted into embeddings which are stored within the database alongside the original textual data for quick retrieval. Such text embedding may enhance the system's ability to recognise semantic similarities, contextual nuances, and relationships within the textual data.
[0155] The database may be hosted by a processing device of the at least one processing device. In such case, the processing device may comprise a server. Alternatively, the database may be hosted by an external server, with which the at least one processing device communicates.
[0156] In some embodiments, the at least one processing device may be configured to transmit communications to and receive communications from the external server over a communications network, which may include, amongst others, the Internet, LANs, WANs, GPRS network, a mobile communications network, a radio network (UHF-band), etc., and may include wire and / or wireless communication links, preferably the latter.
[0157] In some embodiments, the communications may be received and transmitted via a private network connection established between the remote server and the at least oneprocessing device.
[0158] For example, in some embodiments, the private network connection may be a secure communication session across an encrypted communication channel such as Hypertext Transfer Protocol Secure (HTTPS), Transport Layer Security / Secure Sockets Layer (TLS / SSL) or some other secure channel.
[0159] In other embodiments, the private network connection may be a VPN connection established using an encrypted layered tunnelling protocol and authentication methods, including identifiers, passwords and / or certificates.
[0160] For example, in some such embodiments, the at least one processing device may be assigned a unique identifier that may be registered with the remote server. In use, the remote server may establish a VPN connection with the at least one processing device upon authenticating the identifier assigned to the at least one processing device.
[0161] In some embodiments, the communications between the remote server and the at least one processing device may require authentication, such as, e.g., identifiers, passwords, captcha and / or two-factor (2 FA) authentication.
[0162] According to a second aspect of the present invention, there is provided a clinical pathology reporting method including: extracting pathology data from a pathology test report using a machine learning model; providing, on a data interface, and to a user, the pathology data for review by the user; receiving, on the data interface, and from the user, corrected pathology data; providing at least part of the corrected pathology data to the machine learning model to further improve the machine learning model; retrieving auxiliary content according to the corrected pathology data; and generating a healthcare report from the corrected pathology data and the auxiliary content.
[0163] Any of the features described herein can be combined in any combination with any one or more of the other features described herein within the scope of the invention.
[0164] The reference to any prior art in this specification is not, and should not be taken as an acknowledgement or any form of suggestion that the prior art forms part of the common general knowledge.BRIEF DESCRIPTION OF DRAWINGS
[0165] Preferred features, embodiments and variations of the invention may be discerned from the following Detailed Description which provides sufficient information for those skilled in the art to perform the invention. The Detailed Description is not to be regarded as limiting the scope of the preceding Summary of Invention in any way. The Detailed Description will make reference to a number of drawings as follows:
[0166] Figure 1 illustrates a schematic of a clinical pathology reporting system, according to an embodiment of the present invention;
[0167] Figure 2 illustrates a clinical pathology reporting method, according to the embodiment of the present invention; and
[0168] Figures 3a, 3b and 3c together illustrate a pathology reporting method, according to an embodiment of the present invention.DETAILED DESCRIPTION
[0169] Embodiments of clinical pathology reporting systems and methods are provided below that assist users, such as medical practitioners, in interpreting pathology data and that offer insight to users, while maintaining accuracy.
[0170] Figure 1 illustrates a clinical pathology reporting system (100), according to an embodiment of the present invention.
[0171] The system (100) includes a processing device (110) in the form of a server, configured to receive and process pathology test reports (150) and extract pathology data therefrom using a machine learning model (115). The extracted pathology data is then provided to a user (190) for review and correction. In case of correction, the corrected data is fed back to the machine learning model (115). The processing device (110) then analyses the corrected pathology data, and retrieves auxiliary content according to the analysed pathology data. Finally, a healthcare report (180) is generated from the corrected pathology data and the auxiliary content.
[0172] The system (100) is able to continually improve its interpretation of pathology test reports over time, and to adapt to different types of reports over time through feedback to the machine learning model (115). The system (100) is also able to maintain accuracy, even when interpretation errors occur, through user review. Furthermore, the system (100) is able to offer insight in relation to pathology test results through the auxiliary content, to facilitate informeddecision-making.
[0173] In use, the processing device (110) receives, on a data interface thereof, the pathology test report (150). The processing device (110) may receive the pathology test report (150) from any suitable source, including from the user, from an associated pathology provider, or from a health record or patient management system. As an illustrative example, the processing device (110) may include a user interface, which enables the user to upload a pathology test report. Similarly, the processing device (110) may interface with a pathology provider, a health record or patient management system, e.g. using an Application Programming Interface (API).
[0174] The pathology test report (150) is a medical report that contains patient information and associated pathology data, such as diagnostic measurements relating to the patient, or clinical imaging, for example. The diagnostic measurements may relate to a specimen taken from the patient (e.g. a blood or urine test), and the measurement data may include one or more measurements (e.g. red blood cell count) relating to the specimen. The skilled addressee will, however, readily appreciate that any suitable type of specimen or patient analysis may be included.
[0175] The pathology test report (150) may further include interpretation data, relating to an interpretation of the pathology data, and the interpretation data may be in one or more predefined categories (e.g. high, medium, low, supraoptimal or suboptimal). For example, the report may note that cholesterol is high.
[0176] The pathology test report (150) may further include comments, as well as possible recommendations for further evaluation, treatment, or management.
[0177] The pathology test report (150) may take any suitable form, but is generally an electronic document comprising or including an image, such as a scanned image of a report.
[0178] In such case, upon receipt of the pathology test report (150), the processing device (110) may start by analysing the image and pre-processing the image based on the analysis. This may result in one or more corrective measures being applied to the image based upon the analysis. For example, the processing device (110) may identify reference points or lines within the image, such as text lines or edges, to determine skew, and a degree and direction of skew. Based on this analysis, corrective transformations are then applied to the image to eliminate the skew and to ensure that the text appears upright and straight. The processing device (110) may, however, perform any suitable pre-processing steps, such as noise reduction and contrast enhancement, prior to extracting the pathology data.
[0179] The processing device (110) then uses the machine learning model (115) to extract the pathology data from the pre-processed pathology test report (150), together with associated information, such as patient information.
[0180] The machine learning model (115) includes an Artificial Neural Network based Optical Character Recognition (ANN-based OCR) model which is able to perceive and recognise text (characters) based on topological features such as shape, symmetry, closed or open areas, and number of pixels. The ANN-based OCR model consists of multiple layers of interconnected neurons in its neural network.
[0181] The machine learning model (115) is trained prior to the extraction of pathology data from the pathology test report (150). In particular, the machine learning model (115) is trained using a training dataset comprising examples of pathology test reports, and associated pathology data.
[0182] Prior to training the machine learning model (115), the processing device (110) may prepare the dataset for training, e.g. by removing duplicate data, and removing or transforming outliers if they are deemed erroneous or if they could unduly influence the model. The processing device (110) may also divide the dataset into training and test sets to train the machine learning model on the training set but evaluate the model’s (115) performance on the test set.
[0183] Once the pathology data and associated information is extracted from the pre- processed pathology test report (150), it is provided to the user (190) for review.
[0184] In particular, the processing device (110) generates a user interface (118), which includes the pathology data and associated information for review. This enables the extracted pathology data to be reviewed remotely on a user computing device. The user interface (118) provides various tools and options enabling the user (190) to view, edit, and annotate the extracted pathology data.
[0185] The user interface may include a variety of different configurations for displaying the extracted pathology data.
[0186] The user interface (118) may allow for a side-by-side comparison between the extracted pathology data and the pathology test report. It is envisaged that this arrangement enables the user to quickly identify any discrepancies or inconsistencies.
[0187] The user interface (118) may display the extracted pathology data in a structured manner by arranging the extracted pathology data into multiple elements, such as patient information, specimen information, test details, results, interpretations / comments, andrecommendations.
[0188] The extracted pathology data within each element may be presented in a tabular format. The patient information may comprise fields of patient ID, name, date of birth, gender, and / or medical record no. The specimen information may comprise fields of specimen ID, collection data, specimen type, and / or source. The table of test details may comprise fields of test ID, test name, test method, and / or test date. The table of results may comprise fields of result ID, parameter, value, units, and / or reference range. The table of interpretations / comments may comprise fields of comment ID and / or comment. The table of recommendations may comprise fields of recommendation ID and recommendation.
[0189] The user interface (118) displays the extracted pathology data as editable text fields allowing the user (190) to rectify errors or fill in missing information. The user interface (118) may further include drag-and-drop features to help realign data elements.
[0190] The user interface (118) may display one or more suggestions for corrections based on patterns or known data structures. Such configuration is useful, as the user (190) only needs to accept or reject the corrections as necessary on the user interface (118).
[0191] Once the extracted pathology data has been reviewed and corrected using the user interface (118), it is uploaded, to the processing device (110). This may be performed as corrections are made, or upon selection of an element (e.g. a save button) of the user interface.
[0192] Upon receipt of the corrected pathology data, on the data interface and from the user (190), the processing device (110) provides at least part of the corrected pathology data to the machine learning model (115) to further improve the machine learning model (115). In particular, the system incorporates the at least part of the corrected pathology data into the dataset for training the machine learning model (115), augmenting the dataset with images and the corrected corresponding character labels.
[0193] The system (100) then either retrains the machine learning model (115) with the updated dataset, or additionally trains the machine learning model (115) with the corrected pathology data and the pathology test report (150) allowing the model (115) to learn from its errors and improve its accuracy.
[0194] The retraining process may continue until a convergence criterion is met, such as achieving a desired level of accuracy on a test set or stabilising a loss function. It is envisaged that the system (100) may iteratively augment the data set by incorporating corrected pathology data over time as pathology test reports are processed. Accordingly, the system (100)periodically removes data from the data set to ensure that the data set remains manageable and that the retraining times are kept within an acceptable limit.
[0195] The system (100) may quantify the difference between the extracted pathology data and the corrected pathology data by calculating the magnitude of the difference, such as the number of characters or digits that differ between the two versions of the data. The system may then further flag data points with significant differences or high delta values for further manual review on the user interface, for example, using colour coding, highlighting or similar.
[0196] As outlined above, is able to offer insight in relation to pathology test results through the auxiliary content, to facilitate informed decision-making.
[0197] Upon receipt of the corrected pathology data, the system (100) further retrieves auxiliary content (170) according to the corrected pathology data. The auxiliary content can contain a wide range of materials, including academic papers, textbooks, clinical guidelines, research articles, case studies, and educational modules.
[0198] The at least one processing device (110) may be configured to determine one or more risks or attributes associated with the corrected pathology data. An example of a risk is risk for cardiovascular disease. The skilled addressee will, however, readily appreciate that any suitable risk may be identified.
[0199] The auxiliary content may be retrieved according to the determined one or more risks or attributes. For example, the auxiliary content may be categorised according to risks or attributes.
[0200] The auxiliary content may be retrieved according to one or more analytes of the pathology data.
[0201] The auxiliary content may be retrieved according to one or more identified conditions of the pathology data.
[0202] The auxiliary content (170) usually includes detailed descriptions of analytes and medical conditions. These descriptions are tailored to meet the specific needs of healthcare practitioners and researchers, offering comprehensive insights into the properties, functions, and clinical significance of various analytes and conditions.
[0203] The auxiliary content (170) may be retrieved from any suitable source, including the Internet, as well as proprietary databases, systems within an organisation, to third-party services and data providers, e.g. using APIs.
[0204] Once the processing device (110) retrieves the auxiliary content (170) according to the corrected pathology data, the processing device (110) generates the healthcare report (180) from the corrected pathology data and the auxiliary content (170).
[0205] This may be performed using a template, or other standardised form, which may be selectable by the user.
[0206] The healthcare report (180) may take any suitable form, including a document or an interactive website. The user may select the form of the healthcare report (180) through user settings.
[0207] Generally, however, the healthcare report (180) is provided in a user-friendly format displaying the corrected pathology data with reference to the auxiliary content (170), which offers information to assist the user (190) in obtaining a better understanding of any health conditions.
[0208] The system (100) further includes a database (140) to store data. The database (140) may take any suitable form, but in this embodiment is in the form of a relational database, and is managed by a database management system (DBMS).
[0209] The pathology test report, the pathology data, the corrected pathology data, the auxiliary content and / or the healthcare report may be stored on the database (140).
[0210] In general, the data in the database (140) may be organised based on the data types and semantics, including patient information, specimen information, test details, results, interpretations / comments, recommendations, and other relevant data types. The system (100) uses different tables within the database (140) to represent different types of the data. Each table has well-defined columns that correspond to the attributes or characteristics of the data type.
[0211] Accordingly, the system (100) is capable of indexing the pathology data along with their corresponding patient information, specimen information, test details results, interpretations / comments and recommendations in the database (140) to facilitate efficient filtering and retrieval of information. By using natural language processing technologies, the system (100) generates embeddings for textual data in the database (140) and stores the embeddings within the database (140) alongside the original textual data for quick retrieval.
[0212] While the above description refers to the pathology test report, the patient, the pathology data, the corrected pathology data, the healthcare report and the user all in singular, the skilled addressee will readily appreciate that the system (100) is configured to process many such reports, and with many users. For example, the system (100) may be used with manyhealth care professionals.
[0213] Figure 2 illustrates a clinical pathology reporting method (200), according to the embodiment of the present invention. The method (200) may be similar or identical to the method performed by the system (100) as shown in Figure 1 to generate a healthcare report (180) from a pathology test report.
[0214] At step 210, pathology data is extracted from a pathology test report using machine learning. In particular, the pathology test report is provided to an already trained machine learning model to extract pathology data from the pathology test report. The machine learning model may comprise or include an Artificial Neural Network based Optical Character Recognition (ANN-based OCR) model.
[0215] At step 220, the extracted pathology data is provided on a data interface to a user and for review by the user. This may be done using a graphical user interface, enabling the user to compare the extracted pathology data to that of the original pathology test report to identify any errors. If any errors or inconsistencies in the extracted pathology data are identified, the user may correct the extracted pathology data using the user interface.
[0216] At step 230, the corrected pathology data is received on the data interface from the user.
[0217] At step 240, at least part of the corrected pathology data is provided to the machine learning model to further improve the machine learning model. The corrected data may be added into a training dataset and then the machine learning model is retrained using the updated data set. Alternatively, the corrected data may be used to additionally train the model. As a result, the machine learning model can learn from its errors and improve its accuracy over time.
[0218] At step 250, auxiliary content relevant to the corrected pathology data, which may include academic papers, textbooks, clinical guidelines, research articles, case studies, and educational modules, is retrieved. The auxiliary content may be retrieved from the Internet or any suitable data source.
[0219] The corrected pathology data may be analysed, and one or more attributes (e.g. health risks) may be determined. The auxiliary content may be retrieved according to those attributes, or any other criteria.
[0220] At step 260, a healthcare report is generated from the corrected pathology data and retrieved auxiliary content.
[0221] The auxiliary content may provide context, background information, and supplementary knowledge relevant to the interpretation and analysis of the corrected pathology data. As such, the auxiliary content assists in providing insights to the user.
[0222] Finally, the healthcare report is provided to the user.
[0223] Figures 3a, 3b and 3c together illustrate a pathology reporting method (300), according to an embodiment of the present invention. The method (300) may be similar or identical to the method performed by the system (100) of Figure 1.
[0224] Initially, the pathology test report is uploaded (or otherwise provided) at step 310.
[0225] At step 320, the pathology test report is pre-processed to enhance its quality and remove noise. This may be performed with various technologies, such as noise reduction, contrast enhancement, and / or image normalisation.
[0226] At step 330, a machine learning model comprising or including an ANN-based OCR model extracts pathology data from the pre-processed pathology test report.
[0227] The extracted pathology data from the pathology test report is then provided to the user for review and validation / correction in step 340. Analytes, ranges and units in medical standards and guidelines may also be provided to the user to review and make corrections. The corrected pathology data is stored in the database (140).
[0228] At step 350, at least part of the corrected pathology data is provided to the machine learning model to further train the machine learning model.
[0229] Auxiliary content (170) relating to the corrected pathology data is retrieved from one or more data sources, for use in generating a healthcare report. The auxiliary content 170 may be selected to provide context, background information, and supplementary knowledge relevant to the interpretation and analysis of the corrected pathology data.
[0230] The auxiliary content (170) may contain a wide range of materials, including academic papers, textbooks, clinical guidelines, research articles, case studies, and educational modules. The auxiliary content (170) may take various forms, including articles, audio, video, and e-books.
[0231] To ensure accessibility and usability, non-text auxiliary content (170) undergoes processing to extract or transcribe information therefrom into text format, for example, at step 360.
[0232] Once the auxiliary content (170) is converted into text format, it undergoes further processing at step 370. The resulting text content is then stored and indexed by encoding the text data into high-dimensional vector representations with a text embedding technology, enabling efficient storage and retrieval within the database (140).
[0233] At step 380, the corrected pathology data, including patient information, pathology results, and historical health information, and the auxiliary content are used to generate the healthcare report (180). In this process, Retrieval Augmented Generation (RAG) is used to augment a large language model (LLM) with the auxiliary content, external knowledge from content on the Internet and / or 3rdParty Data Providers, to generate content for the healthcare report (180). Reference links to webpages and / or websites on the internet as well as third party data may also be included in the healthcare report. The integration of auxiliary content with Al language models enables the healthcare report to include a wealth of contextual information.
[0234] Optionally, at step 385, the corrected pathology data is used to identify one or more conditions.
[0235] In particular, the corrected pathology data may be analysed with reference to a predefined list of medical conditions, each condition associated with at least one marker (e.g. corresponding to an analyte in the corrected pathology data) and an associated state. The state may comprise a range or threshold indicative of the condition.
[0236] Corrected pathology data aggregated from multiple pathology test reports may be analysed, to identify trends. Such analysis may be useful in detecting the emergence of potential conditions over time.
[0237] The analysis of the corrected pathology data may be performed using advanced analytical techniques, such as statistical analysis or machine learning algorithms, which can enhance the accuracy of condition identification. The identified medical conditions, as well as any potential future conditions, will be included in the healthcare report (180).
[0238] Lastly, the healthcare report (180) is presented to the user at step 390.
[0239] Advantageously, the systems and methods described above enable pathology data to be viewed by a user in a form that is accurate, easy to process, and with associated content, to facilitate informed decision-making. The systems and methods are able to continually improve the interpretation of pathology test reports over time, and to adapt to different types of reports over time through feedback to the machine learning model. The systems and methods are also able to maintain accuracy, even when interpretation errors occur, through user review.
[0240] In the present specification and claims (if any), the word ‘comprising’ and its derivatives including ‘comprises’ and ‘comprise’ include each of the stated integers but does not exclude the inclusion of one or more further integers.
[0241] Reference throughout this specification to ‘one embodiment’ or ‘an embodiment’ means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrases ‘in one embodiment’ or ‘in an embodiment’ in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more combinations.
[0242] In compliance with the statute, the invention has been described in language more or less specific to structural or methodical features. It is to be understood that the invention is not limited to specific features shown or described since the means herein described comprises preferred forms of putting the invention into effect. The invention is, therefore, claimed in any of its forms or modifications within the proper scope of the appended claims (if any) appropriately interpreted by those skilled in the art.
Claims
CLAIMS1 . A clinical pathology reporting system for generating a healthcare report from a pathology test report, including: at least one processing device configured to: extract pathology data from the pathology test report using a machine learning model; provide, on a data interface, the pathology data for review by a user; receive, on the data interface, and from the user, corrected pathology data; provide at least part of the corrected pathology data to the machine learning model to train the machine learning model; retrieve auxiliary content according to the corrected pathology data; and generate the healthcare report from the corrected pathology data and the auxiliary content.
2. The clinical pathology reporting system of claim 1 , wherein the healthcare report comprises an interactive webpage or document.
3. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is further configured to extract patient data from the pathology test report using the machine learning model, and wherein the healthcare report includes the patient data or a derivative thereof.
4. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is further configured to receive the pathology test report on the data interface.
5. The clinical pathology reporting system of claim 4, wherein the at least one processing device is configured to receive a plurality of pathology test reports on the data interface, including the pathology test report.
6. The clinical pathology reporting system of claim 5, wherein the plurality of pathology test reports relate to a patient, and wherein the at least one processing device is configured to generate the healthcare report for the patient from the corrected pathology data of the plurality of pathology test reports and the auxiliary content.
7. The clinical pathology reporting system of claim 6, wherein the plurality of pathology test reports comprise test reports relating to tests performed at different points in time, and wherein the healthcare report includes historical pathology data of the patient.
8. The clinical pathology reporting system of claim 1 , wherein the pathology test report comprises an electronic document or one or more electronic images representing a document.
9. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is configured to generate a user interface to display the extracted pathology data to the user.
10. The clinical pathology reporting system of claim 9, wherein the user interface is configurable to display the extracted pathology data and the pathology test report to enable comparison thereof.
11. The clinical pathology reporting system of claim 9, wherein the user interface is configured to enable the user to correct the pathology data, and thereby generate the corrected pathology data.
12. The clinical pathology reporting system of claim 1 , configured to quantify a difference between the extracted pathology data and the corrected pathology data, and flag data for further review according to a magnitude of the difference between the extracted pathology data and the corrected pathology data.
13. The clinical pathology reporting system of claim 1 , wherein the machine learning model is further trained using the pathology test report and at least part of the corrected pathology data.
14. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is configured to identify errors in the pathology data, and automatically correct the errors.
15. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is configured to generate audit data including the pathology data and the corrected pathology data.
16. The clinical pathology reporting system of claim 1 , wherein the at least one processing device is configured to select the machine learning model from a plurality of machine learning models.
17. The clinical pathology reporting system of claim 16, wherein the at least one processing device is configured to select the machine learning model according to a type of the pathology test report.
18. The clinical pathology reporting system of claim 1 , wherein the machine learning modelcomprises an Artificial Neural Network based Optical Character Recognition (ANN-based OCR) model configured to extract the pathology data in the form of text from the pathology test report.
19. The clinical pathology reporting system of claim 1 , wherein one or more risks or attributes are determined according to the corrected pathology data, wherein the auxiliary content is retrieved at least in part according to the determined one or more risks or attributes.
20. A clinical pathology reporting method including: extracting pathology data from a pathology test report using a machine learning model; providing, on a data interface, and to a user, the pathology data for review by the user; receiving, on the data interface, and from the user, corrected pathology data; providing at least part of the corrected pathology data to the machine learning model to further improve the machine learning model; retrieving auxiliary content according to the corrected pathology data; and generating a healthcare report from the corrected pathology data and the auxiliary content.
Citation Information
Patent Citations
Imaging related clinical context apparatus and associated methods
US20190156921A1
Mobile supplementation, extraction, and analysis of health records
US20200126663A1
Automated generation of structured patient data record
US20220044812A1
Oncology workflow for clinical decision support
US20240021280A1