Tool for interpreting and analyzing pulmonary nodule report based on large model

By automating the analysis of lung nodule reports through tools based on large models, the problems of low report analysis efficiency and unstable analysis results in existing technologies are solved, and efficient and accurate nodule diagnosis and data sharing are achieved.

CN120673432APending Publication Date: 2025-09-19SHANGHAI SHANTAI HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510762019.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have low report analysis efficiency and poor timeliness in underwriting and platform consulting, and the analysis results are limited by manual experience, resulting in unstable quality and high error rate, which cannot meet the needs of rapid response.

Method used

A large model-based tool is used to automatically analyze lung nodule reports, including modules such as electronic file upload, OCR recognition, text pre-analysis, key content extraction, report content analysis, and structured report output, to automatically extract and analyze patient information and nodule characteristics.

Benefits of technology

It improves the accuracy and automation level of nodule analysis, reduces the need for human intervention, ensures the reliability and consistency of diagnosis, and facilitates data sharing and system integration through structured data output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673432A_ABST
    Figure CN120673432A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical data processing and artificial intelligence, and discloses a tool for interpreting and analyzing a pulmonary nodule report based on a large model, and the tool comprises an electronic file uploading module which is used for receiving an electronic file of the pulmonary nodule report uploaded by a user; the OCR recognition module is used for performing optical character recognition on the uploaded electronic file and outputting preliminary text data; the text pre-analysis module is used for carrying out interference degree detection and cleaning on the text data subjected to OCR recognition; and a key content extraction module. Through automatic report analysis, structured data output and a multi-level image preprocessing technology based on a large model, the analysis accuracy and processing efficiency of the pulmonary nodule report are improved, misjudgment and missed judgment in manual analysis are avoided, data standardization and system compatibility are ensured, and the accuracy and processing efficiency of the pulmonary nodule report are improved. Meanwhile, the image quality problem in OCR recognition is effectively solved, and the reliability, timeliness and system adaptability of report analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical data processing and artificial intelligence technology, and specifically to a tool for interpreting and analyzing lung nodule reports based on a large model. Background Art

[0002] In the current underwriting process, most steps in insurance companies still rely on manual review. Typically, insurance companies require doctors or other professionals to review a patient's lung nodule report and assess their health risks based on the diagnosis. This process typically takes one to two days, resulting in low underwriting efficiency and failing to meet the demand for rapid response. When faced with a large number of cases, the efficiency and accuracy of manual review are significantly affected, further prolonging the entire underwriting process.

[0003] Furthermore, in platform consultation scenarios, doctors spend approximately five minutes carefully reading a report. This means that traditional manual interpretation methods cannot provide real-time service for cases requiring urgent diagnosis. This latency issue can lead to delayed diagnosis and treatment in some emergency situations. This is particularly true in insurance scenarios, where rapid decision-making is crucial, as manual interpretation does not meet industry requirements in terms of timeliness.

[0004] More seriously, the quality of manual analysis is affected by the experience of the reviewers. Inexperienced doctors may miss key information, leading to biased analysis results, while experienced doctors may make errors in judgment due to fatigue from long hours. This human-induced instability increases the error rate in report analysis, further reducing the quality of medical services and insurance underwriting.

[0005] Most existing manual or semi-automated analysis tools suffer from limitations in speed, accuracy, and consistency. Especially when dealing with large amounts of data, manual analysis is not only time-consuming but also prone to omissions, failing to meet the demands of modern insurance companies and healthcare platforms for real-time, efficient, and highly accurate diagnosis. Consequently, existing technologies fail to effectively address the timeliness issues in underwriting and platform consulting, and fail to provide a reliable automated report analysis solution to improve the efficiency and accuracy of report analysis.

[0006] In this situation, there is an urgent need for an efficient and automated technical solution that can extract key information from reports in a short period of time, conduct accurate analysis, and output structured diagnostic conclusions in real time to meet the growing needs of the medical and insurance industries. Summary of the Invention

[0007] In response to the shortcomings of existing technologies, the present invention provides a tool for interpreting and analyzing lung nodule reports based on a large model, solving the problems of low report analysis efficiency, poor timeliness, and analysis results being limited by experience in traditional manual underwriting and platform consultation.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: a tool for interpreting and analyzing lung nodule reports based on a large model, including: An electronic file upload module is used to receive electronic files of pulmonary nodule reports uploaded by users; The OCR recognition module is used to perform optical character recognition on uploaded electronic files and output preliminary text data; Text pre-analysis module, used to detect and clean the interference of text data after OCR recognition; Key content extraction module, used to extract the patient's name, gender, age, hospital, report date, and nodule size, density, location, and morphology from text data; Report content analysis module, used to analyze the extracted nodule information and infer the characteristics and risk level of the nodules; The structured report output module is used to convert nodule analysis results into structured reports and output standardized JSON format reports; The data storage and management module is used to store the generated structured reports.

[0009] Preferably, the electronic file upload module includes: The file receiving unit is used to receive electronic files uploaded by users and perform preliminary identification based on the file format; The format conversion unit is used to convert uploaded files of different formats into a unified format for use by the OCR recognition module.

[0010] Preferably, the OCR recognition module includes: Image preprocessing unit, used to preprocess the uploaded image files, eliminate noise, adjust image clarity and perform rotation correction; A character segmentation unit, used for segmenting the text in the image; The text recognition unit is used to recognize characters through convolutional neural networks and convert the text in the image into processable text data; The recognition result output unit is used to output the preliminary text data after OCR recognition for use by the subsequent text pre-analysis module.

[0011] Preferably, the text pre-analysis module includes: The interference detection unit is used to calculate the interference of the OCR text based on the clarity, font size and other characteristics of the text; Angle adjustment detection unit, used to detect text tilt or image deviation in the OCR recognition results and determine whether the OCR recognition angle needs to be adjusted; The pre-processing adjustment unit is used to guide the OCR module to adjust the angle or parameters of text recognition according to the interference detection result.

[0012] Preferably, the key content extraction module includes: Semantic analysis unit, used to perform semantic analysis on the text using a large model to identify key information in the report, including the patient's name, gender, age, hospital, report date, and the size, density, location, and morphology of the nodule; The prompt word matching unit is used to match relevant content in the text based on the set prompt word library and extract the size, density, location and morphology of the nodule; The information standardization unit is used to standardize the extracted information.

[0013] Preferably, the report content analysis module includes: Nodule feature inference unit, which is used to analyze the morphology and possible risk type of nodules based on the characteristics of nodule size, density, location, etc. through inference using a large model; Nodule risk assessment unit, which is used to evaluate the risk level of nodules and generate relevant diagnostic results by combining the pulmonary nodule knowledge base and pulmonary nodule guidelines; The analysis result output unit is used to pass the analysis results to the structured report output module to generate a detailed nodule diagnosis report.

[0014] Preferably, the structured report output module includes: The result conversion unit is used to convert the analysis results into a structured data format, JSON format, to ensure good data compatibility; Formatting output unit, used to format structured data into a final report, including the patient's basic information, nodule characteristics and diagnosis results; The result display unit is used to display the final report to the user for subsequent use by doctors or related personnel.

[0015] Preferably, the data storage and management module includes: A data storage unit, used for classifying and storing structured reports according to generation time; Data retrieval unit, used to support users to retrieve, query and download stored reports; The data security management unit is used to provide security protection for stored reports.

[0016] This invention provides a tool for interpreting and analyzing lung nodule reports based on a large model. It has the following beneficial effects: 1. The present invention adopts a technical solution based on large-scale model interpretation and analysis of lung nodule reports, which can automatically extract basic patient information and nodule characteristics from the report, and analyze the morphology and risk level of the nodule through large-scale model reasoning. This technical effect effectively improves the accuracy and automation level of nodule analysis. Compared with the manual analysis or traditional rule matching solutions in the existing technology, the present invention avoids misjudgments and missed judgments caused by differences in doctor experience or imperfect rule settings, greatly reduces the need for human intervention, and improves the reliability and consistency of diagnosis.

[0017] 2. The present invention adopts a structured data output method to convert the nodule analysis results into a standardized JSON format report, and formatted output and display. This technical solution ensures the uniformity and compatibility of the data, facilitates seamless connection with other systems, supports data sharing and further analysis. Compared with the non-uniform format or manual entry method in the existing technology, the present invention makes the data more standardized and normalized, reduces the complexity of information processing, and reduces the difficulty of system integration.

[0018] 3. The present invention improves the accuracy and image quality of OCR recognition by adopting multi-level optimization methods such as image preprocessing, interference detection, and angle adjustment. This technical solution ensures that even reports with shooting angle deviation or poor image quality can still obtain high-quality text output. Compared with the single OCR recognition solution in the prior art, the present invention effectively solves the negative impact of poor image quality on recognition results through multiple preprocessing and intervention measures, significantly improving the processing effect and system adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a system framework diagram of the present invention; Figure 2 This is a framework diagram of the electronic file upload module of the present invention; Figure 3 This is a framework diagram of the OCR recognition module of the present invention; Figure 4 This is a framework diagram of the text pre-analysis module of the present invention; Figure 5 This is a framework diagram of the key content extraction module of the present invention; Figure 6 This is a framework diagram of the report content analysis module of the present invention; Figure 7 This is a framework diagram of the structured report output module of the present invention; Figure 8 This is a framework diagram of the data storage and management module of the present invention; Figure 9 This is a report analysis flow chart of the system of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] Please see the attached Figure 1 and 9 , an embodiment of the present invention provides a tool for interpreting and analyzing lung nodule reports based on a large model, including: Please see the attached Figure 2 ,The electronic file upload module is used to receive the electronic file of the lung nodule report uploaded by the user.

[0022] This embodiment provides an electronic file upload module for receiving electronic files of pulmonary nodule reports uploaded by users. The main function of this module is to receive files of different formats and ensure that they can be seamlessly passed to the subsequent OCR recognition module for processing. Uploaded electronic files may include PDF files, scanned image files, pictures taken by mobile phones, etc. All of these file formats need to be uniformly formatted by this module to ensure the efficiency and accuracy of subsequent processing.

[0023] In some embodiments, the electronic file upload module can also detect and pre-process uploaded file formats, ensuring that uploaded files do not fail processing due to formatting issues. Through this module, the system can efficiently convert files of various formats into a format suitable for OCR processing, providing the necessary data support for subsequent report analysis and conclusion output.

[0024] In this embodiment, the electronic file upload module includes the following subunits: The file receiving unit is used to receive electronic files of pulmonary nodule reports uploaded by users. Uploaded files can be in various formats, including but not limited to PDF, scanned image files (such as JPEG and PNG), and mobile phone photos. Upon receiving a file, the file receiving unit first checks the file type and validity to ensure that the file format meets the requirements for subsequent processing.

[0025] In some embodiments, the file receiving unit uses the MIME type to perform preliminary identification of the uploaded file. By checking the MIME type, the system can determine whether the uploaded file is a PDF file, a scanned image, or another file type. The file receiving unit can also determine the file size to ensure that the uploaded file is within a specified size range. If the file does not meet the specified size range, the system will return an error message.

[0026] The file receiving unit can receive files through HTTP request, FTP upload or other compatible electronic file transmission methods.

[0027] The format conversion unit converts files of different formats into a unified format to ensure accurate processing by the OCR module. Since the OCR module typically works with image-based files (such as JPEG and PNG), the format conversion unit converts non-image-based files into a standard image format.

[0028] Specifically, in one possible implementation, the format conversion unit can convert each page of a PDF file into an image format, such as PNG or JPEG. This conversion process may include rendering the text content as an image and generating an image of each page that corresponds to the original document. Furthermore, the format conversion unit can also standardize the format of scanned or photographed image files to ensure that they meet OCR recognition requirements.

[0029] In some embodiments, the format conversion unit may also perform image enhancement processing on the image, such as increasing image clarity, adjusting contrast, or brightness, to improve OCR recognition accuracy. If the uploaded file is a low-resolution scanned image or photographed image, the format conversion unit may also perform image enlargement processing to ensure the clarity of image details.

[0030] In some embodiments, the format conversion unit may also perform rotation correction on the captured image to address tilt issues caused by improper shooting angles. In extreme cases where the image is excessively distorted or does not meet recognition requirements, the format conversion unit may also request the user to re-upload a clearer file.

[0031] Assume the input format of the file is , the file format type is , and the output format is , the output format is suitable for OCR recognition. The format conversion process can be expressed as: ; in: The original format of the input file, such as a PDF file, a scanned image, or a photographed image; The format type of the input file (such as PDF, JPG, PNG, etc.); The converted image is in a standard format (such as PNG, JPG, etc.), suitable for input into the OCR recognition module.

[0032] This formula shows how the format conversion unit converts the file into a standard format acceptable to the OCR module based on the format type of the input file, ensuring the smooth progress of subsequent processes.

[0033] In some embodiments, the file receiving unit and format conversion unit also have a file verification function. When a file is uploaded, the system can perform a file integrity check. By calculating the file's hash value (e.g., MD5, SHA-256), the system can determine whether the file has been tampered with or damaged. If the file's hash value does not match the expected hash value before uploading, the system will prompt the user to re-upload the file.

[0034] Furthermore, the electronic file upload module generates a unique identifier (UUID) for each uploaded file, which is used for subsequent file management and report generation. This unique identifier ensures that each report can be associated with a specific user or case, preventing data confusion.

[0035] The file receiving unit typically interacts with other modules using protocols such as HTTP and File Transfer Protocol (FTP). Users upload files via the web or mobile app. The file receiving unit receives and initially identifies the file format before passing it to the format conversion unit for processing. Once the format conversion is complete, the converted image data is passed to the OCR module for text recognition.

[0036] The electronic file upload module is responsible for standardizing the formatting of lung nodule report files in different formats, ensuring accurate processing by the OCR recognition module. Through the file receiving unit and format conversion unit, the system supports multiple file upload types, including PDF files, scanned image files, and mobile phone photos. During the file upload and format conversion process, the system also performs file verification, image enhancement, and rotation correction to ensure that the uploaded files meet OCR recognition requirements, thereby improving the system's flexibility and accuracy.

[0037] Please see the attached Figure 3 ,OCR recognition module is used to perform optical character recognition on the uploaded electronic files and output preliminary text data.

[0038] The image preprocessing unit is used to preprocess the uploaded image files to ensure that the image quality is suitable for subsequent character recognition. The goal of image preprocessing is to remove noise from the image, adjust the image clarity, and perform rotation correction, thereby improving the accuracy of OCR recognition.

[0039] In some embodiments, the image preprocessing unit employs an image denoising algorithm, such as a median filter or a mean filter, to remove any noise that may be present in the image, thereby preventing it from interfering with the character recognition process. For blurred or unclear images, the image preprocessing unit may also employ a sharpening algorithm to enhance image clarity, thereby improving character recognition.

[0040] The image preprocessing unit also performs rotation correction on the captured image to ensure the text is at the correct angle. Specifically, if the text in the image is tilted, the image preprocessing unit uses an image rotation correction algorithm to adjust the image angle so that the text orientation meets OCR recognition requirements. Image rotation correction can be automated using angle detection algorithms (such as the Hough transform), eliminating manual intervention.

[0041] Assume that the image input is , the output after image preprocessing is , then the image preprocessing process can be expressed as: ; in: is the original input image data; It is the image data after noise removal, sharpening and rotation correction.

[0042] The character segmentation unit segments text in an image into individual characters for easy recognition. In some cases, text may be written or printed compactly, with no clear spacing between characters, resulting in blurred character boundaries. The character segmentation unit uses a series of image processing algorithms to segment connected regions (usually individual characters) within an image.

[0043] Specifically, the character segmentation unit can use edge detection-based segmentation algorithms (such as the Canny edge detection algorithm) to extract character boundaries in the image. Alternatively, it can use methods based on connected component analysis to analyze the distance between characters to determine and segment them. For complex scenarios, the character segmentation unit can also use deep learning models to segment characters to accommodate different fonts and character spacing.

[0044] Assume that the image after image preprocessing is , the character set obtained after character segmentation is , then the character segmentation process can be expressed as: ; in: is the preprocessed image; It is a character set after character segmentation, and each element corresponds to an independent character.

[0045] The core task of the text recognition unit is to convert the segmented characters into machine-processable text. To this end, the text recognition unit uses a convolutional neural network (CNN) to extract and classify the features of each character and ultimately output the recognition result.

[0046] Specifically, the text recognition unit utilizes deep learning models, particularly convolutional neural networks, to identify each character. This network first extracts features from the character image, extracting high-dimensional feature vectors for each character. This is then classified through a fully connected layer, ultimately mapping the extracted features to the corresponding character. To improve accuracy, the text recognition unit may also employ a character-level attention mechanism, focusing on the most recognizable parts of the image and reducing interference from other parts.

[0047] In some embodiments, the text recognition unit can perform character recognition using a pre-trained model (such as Tesseract, EasyOCR, etc.). The output of the network is the recognized text data, that is, the text information in the original image.

[0048] Assume the input characters are , the recognized text is , then the text recognition process can be expressed as: ; in: is the character set after character segmentation; It is the text data output after recognition.

[0049] The recognition result output unit is responsible for passing the OCR-recognized text data to the subsequent text pre-analysis module. This unit's primary function is to ensure that the recognition results are accurate and meet the requirements of subsequent processing. The recognition result output unit converts the OCR results into a structured format (such as JSON) for further analysis.

[0050] The recognition result output unit is also responsible for verifying the completeness and accuracy of the recognition results and may further optimize the output text through certain error correction mechanisms (such as spelling checking or context analysis). Finally, the recognition result output unit passes the text data to the text pre-analysis module for further processing.

[0051] Assume that the text recognized by OCR is , the structured result output to the text pre-analysis module is , the output process can be expressed as: ; in: It is the preliminary text data obtained after text recognition; It is structured text data, ready to be input into subsequent analysis modules.

[0052] Please see the attached Figure 4 ,The text pre-analysis module is used to detect and clean the interference of text data after OCR recognition.

[0053] The interference detection unit calculates the interference level of the OCR text based on features such as clarity, font size, character spacing, and image noise. The OCR process is often affected by image quality, such as blur, noise, or low resolution, which can reduce the accuracy of the OCR results. Therefore, the interference detection unit is responsible for detecting the clarity of the OCR text and evaluating its interference level based on the feature values.

[0054] Specifically, the interference detection unit first extracts key text features, such as character edge information, font size, and font contrast. It then uses image processing algorithms (such as image gradient analysis, edge detection, and noise analysis) to analyze the text quality. If the unit detects blurry characters, noise interference, or small spacing between characters, it returns a higher interference score.

[0055] In some embodiments, the distractibility detection unit can also combine machine learning methods to determine the quality of the text using a training set. By comparing the OCR recognition results with the standard answer, the detection unit can use a similarity score (e.g., cosine similarity, Euclidean distance, etc.) to quantify the distractibility of the text.

[0056] Assume that the text recognized by OCR is , and its interference is , then the interference detection process can be expressed as: ; in: The text is recognized by OCR; The interference degree of the text indicates the quality of the text. The greater the interference degree, the worse the text quality.

[0057] The angle adjustment detection unit detects text tilt or image deviation in OCR results and determines whether the OCR angle needs to be adjusted. In practice, user-uploaded images may have tilted text due to factors such as the shooting angle, scanning device issues, or image rotation. This tilt or deviation can affect OCR accuracy, necessitating adjustment.

[0058] Specifically, the Angle Adjustment Detection Unit first analyzes the text orientation in the OCR recognition results. By calculating the text's primary orientation (e.g., through Hough transforms, projection analysis, etc.), the unit can identify the text's tilt angle. If the detected angle deviation exceeds a set threshold, the Angle Adjustment Detection Unit will issue a prompt requesting readjustment of the image's recognition angle.

[0059] In some embodiments, the angle adjustment detection unit may automatically detect the rotation angle in the image through an adaptive algorithm, and perform rotation correction on the image based on the angle to ensure the correctness of the text direction.

[0060] Assume that the tilt angle of the text is , then the angle adjustment detection process can be expressed as: ; in: is the tilt angle of the text; The text is recognized by OCR.

[0061] like Exceeding the set threshold , you need to adjust the recognition angle.

[0062] ; The pre-processing adjustment unit is used to guide the OCR module to adjust the angle or parameters for text recognition based on the interference detection results. By detecting interference and adjusting the angle, the pre-processing adjustment unit can optimize the input image for OCR recognition, ensuring more accurate subsequent OCR results.

[0063] In some embodiments, the pre-processing adjustment unit dynamically adjusts the parameters of the OCR recognition module based on the interference detection results. For example, when a high interference level is detected, the pre-processing adjustment unit may instruct the OCR module to use a higher resolution or a different text recognition model for processing; when the text is tilted at a large angle, the pre-processing adjustment unit may instruct the OCR module to rotate the text to ensure that the characters are recognized at the appropriate angle.

[0064] The preprocessing adjustment unit may also automatically update the recognition parameters of the OCR module, such as image enhancement parameters, character segmentation parameters, etc., to improve recognition accuracy.

[0065] Assuming the interference is , the recognition angle is , OCR recognition parameters are , then the preprocessing adjustment process can be expressed as: ; in: is the interference degree; is the tilt angle of the text; is the adjusted OCR recognition parameters.

[0066] Please see the attached Figure 5 ,The key content extraction module is used to extract the patient’s name, gender, age, hospital, ,report date and the size, density, location, and morphology of the nodule from the text ,data.

[0067] The semantic analysis unit uses a large model to perform semantic analysis on the text after OCR recognition, identifying and extracting key information from the report. Large models (such as BERT and GPT) can effectively understand the semantics of text based on deep learning technology, extracting key information related to the report, such as basic information such as the patient's name, gender, age, hospital, and report date; as well as imaging features such as nodule size, density, location, and morphology.

[0068] Specifically, the semantic analysis unit first performs word segmentation, part-of-speech tagging, and named entity recognition (NER) on the text. After word segmentation and part-of-speech tagging, the system identifies key information contained in the text, such as "name," "gender," "age," "hospital," and "nodule size." The large model accurately identifies the specific content of this information based on the context and marks it as structured data.

[0069] For example, when processing a text such as "Zhang San, male, 45 years old, Ruijin Hospital Affiliated to Shanghai Jiao Tong University, report date: January 18, 2025, nodule size: 1.2 cm, nodule location: right upper lobe of the lung, nodule morphology: lobed, density: high", the semantic analysis unit can identify the above key information and extract it.

[0070] Assume that the text recognized by OCR is , the key information output by the large model is , then the semantic analysis process can be expressed as: ; in: The text is recognized by OCR; Key information extracted from the text, such as patient name, gender, age, hospital, nodule size, density, location, etc.

[0071] The prompt word matching unit is used to match relevant content within the text based on a pre-defined prompt word library, further extracting features such as nodule size, density, location, and morphology. The prompt word matching unit relies on a pre-built prompt word library containing standard terms and keywords related to lung nodules (such as "nodule size," "nodule location," and "density"). By matching these terms with those found within the text, the system can effectively extract specific features of the nodules from the text.

[0072] For example, when processing nodule features, the prompt word matching unit extracts "1.2cm" as the nodule size feature by matching the "nodule size" keyword in the text; extracts "high density" by matching the "nodule density" keyword; and extracts "right upper lobe of the lung" by matching the "nodule location" keyword.

[0073] The prompt word matching unit goes beyond simple keyword matching and also considers contextual information to determine the accuracy of the extracted content. For example, the system will determine whether the term "nodule size" in the text corresponds to the size of the nodule, rather than other irrelevant content.

[0074] Assume the input text is , the prompt word library is , the extracted nodule features are , then the prompt word matching process can be expressed as: ; in: The text is recognized by OCR; To serve as a prompt for the vocabulary, nodule-related terms were included; Nodule features extracted from text, including size, density, position, morphology, etc.

[0075] The Information Standardization Unit is responsible for standardizing the extracted information, ensuring consistency and a unified format. Because terminology may vary between hospitals, physicians, or reports, the Standardization Unit unifies these terms into a standard terminology database, eliminating ambiguity caused by different expressions.

[0076] For example, the size of a nodule may appear as "1.2cm" or "1.2cm", and the information standardization unit will unify them into "1.2cm"; the morphology of a nodule may appear in different expressions such as "burr" and "lobed", and the standardization unit will map them to the unified standard term "lobed".

[0077] The standardization process also includes unit unification, date format unification, etc., to ensure that the final output structured data meets the standardization requirements.

[0078] Assume that the extracted nodule features are , the output after normalization is , then the information standardization process can be expressed as: ; in: The extracted nodule features include size, density, location, morphology, etc. To standardize the nodule characteristics, unify the format, terminology and units.

[0079] Please see the attached Figure 6 ,The report content analysis module is used to analyze the extracted nodule information and ,infer the characteristics and risk levels of the nodules.

[0080] The nodule feature inference unit is responsible for analyzing nodule morphology and potential risk types based on features such as size, density, and location, using a large model. Nodule morphology (such as lobulation and burrs) is a key indicator for determining whether a nodule is benign or malignant. This unit infers potential nodule features by combining imaging data with historical case information.

[0081] In some embodiments, the nodule feature inference unit uses a deep learning model (such as a convolutional neural network (CNN) or a Transformer model) to further analyze the nodule's imaging features. For example, nodule size and density are key indicators for assessing whether a nodule is malignant, and the feature inference unit can infer whether a nodule is high-risk based on these features. Nodule features such as density, morphology (e.g., lobulated or spiculated), and location are typically inferred by combining image data and text descriptions.

[0082] The core goal of nodule feature reasoning is to infer the possibility of malignancy by combining various features of the nodule, and output the morphological description and risk warning of the nodule based on the reasoning results.

[0083] Assume that the size, density, location and other characteristics of the nodule are , the result of nodule feature inference is , then the nodule feature inference process can be expressed as: ; in: The imaging characteristics of the nodule, including size, density, location, morphology, etc. The inference results for the nodule may include morphological analysis and risk assessment of the nodule.

[0084] The Nodule Risk Assessment Unit combines the pulmonary nodule knowledge base and pulmonary nodule guidelines to assess the risk level of nodules and generate relevant diagnostic results. This unit uses an expert system or large model to match nodule characteristics with the criteria in the pulmonary nodule guidelines to assess the risk level of nodules.

[0085] In some embodiments, the nodule risk assessment unit combines information from a pulmonary nodule knowledge base, such as nodule size, morphology, density, and other parameters, with known nodule risk criteria to assess the nodule's risk type (e.g., benign, malignant, or suspicious). The pulmonary nodule knowledge base may include multiple criteria, such as nodule size greater than 1 cm and irregular morphology, which may indicate a malignant tumor.

[0086] Specifically, the risk assessment unit evaluates nodules in combination with standardized clinical guidelines (such as the Boston University Pulmonary Nodule Risk Assessment System or the Fleischner Society Pulmonary Nodule Management Guidelines) and outputs the risk level of the nodules.

[0087] Assume that the risk assessment result of a nodule is , the nodule characteristics are , combining the information from the knowledge base and the guide for , then the risk assessment process can be expressed as: ; in: It is the imaging feature of the nodule; Provide standards for nodule risk assessment in the pulmonary nodule knowledge base or clinical guidelines; To evaluate the results, the risk level of the nodule is indicated (such as low risk, intermediate risk, high risk).

[0088] The analysis result output unit is responsible for transmitting the nodule analysis results to the structured report output module, generating a detailed nodule diagnosis report. The analysis results include the nodule's characteristic inference results, risk assessment level, and final diagnosis conclusion. This unit's mission is to ensure the accuracy and clarity of the analysis results and convert them into structured data for use in subsequent steps.

[0089] Specifically, the analysis result output unit first integrates the results of each analysis module based on the outputs of the nodule feature inference unit and the risk assessment unit to generate a complete nodule diagnosis report. The report includes patient information, a detailed description of the nodule characteristics, the risk assessment level, and the doctor's recommendations. The report is output in a structured format (such as JSON) to ensure ease of subsequent processing and presentation.

[0090] Assume that the nodule characteristics and risk assessment results are and , the output structured report is , then the report output process can be expressed as: ; in: The inference result of the nodule; The risk assessment results for nodules; Generate a structured nodule diagnosis report.

[0091] Please see the attached Figure 7 ,Structured report output module, used to convert nodule analysis results into structured reports and output standardized JSON format reports; In this embodiment, the result conversion unit primarily converts the analysis results generated by the nodule analysis module into a structured JSON data format. The analysis results typically include basic patient information, nodule characteristics (such as size, density, location, morphology, etc.), and the nodule's risk assessment level. The result conversion unit is responsible for converting this information into a unified format and ensuring data compatibility to support data exchange and sharing across different platforms.

[0092] Typically, the result conversion unit retrieves various analysis results from the nodule analysis module and populates these data into corresponding fields according to a pre-set JSON format template. The naming and format of each field strictly adhere to pre-defined standards to ensure that the generated JSON file meets the requirements of structured data.

[0093] Specifically, the patient's basic information (such as name, gender, age, and hospital) is mapped to the "patient_info" field in the JSON format. Nodule characteristics (such as size, density, location, and morphology) are populated in the "nodule_info" field, and the diagnosis result and risk level are populated in the "diagnosis" field. The result conversion unit ensures that all fields conform to the specified data types, such as numbers, strings, and dates.

[0094] Assume that the output of the nodule analysis module is , converted into the JSON format of the structured report: , then the result transformation process can be expressed as: ; in: It is the unstructured analysis result output by the nodule analysis module; The report is in the converted JSON format and has good compatibility and supports subsequent use.

[0095] The output formatter is responsible for further formatting the converted structured data, ensuring that the final report conforms to a standardized format and clearly and accurately conveys the nodule analysis results. The core task of the output formatter is to adjust and sort the fields according to the report requirements to conform to the standardized report format.

[0096] Specifically, the formatted output unit checks the generated structured data to ensure that each item is output according to the predetermined format. For example, the patient's basic information (such as name, gender, age, hospital, and report date) should be arranged in a predetermined order, and the date format should be consistent. Nodule characteristics (such as size, density, location, and morphology) should conform to a unified representation standard (for example, nodule size should be uniformly expressed as a number with a unit, such as "1.2 cm"). The diagnosis result should clearly express the nodule's risk level and possible conclusions.

[0097] In addition, the formatted output unit may also correct some fields, such as unifying the date format to "YYYY-MM-DD" and the nodule size to a number with a unit, etc., to ensure that the report complies with the standards of the medical industry.

[0098] Assume the structured report is The final report after formatting is , then the formatted output process can be expressed as: ; in: The converted structured data; It is the final report that has been formatted and complies with standardization requirements.

[0099] The results display unit displays the formatted final report to users, including physicians and other relevant personnel. It provides a user interface so users can view, download, print, and perform other actions on the report. The display unit ensures the accuracy and clarity of the report content and presents the report data through a graphical interface.

[0100] In some embodiments, the result display unit not only provides a static report display function but may also support dynamic interactive functions, such as filtering different nodule information in the report, viewing the risk assessment details of the nodule, or downloading the report for subsequent processing. In addition, the display unit can also provide a report printing function, making it convenient for doctors or other relevant personnel to obtain a paper report for further analysis.

[0101] Assume that the final report after formatting is , the report displayed to the user is , the display process can be expressed as: ; in: For the formatted final report; The report is usually presented to the user in a visual manner for doctors or relevant personnel to review.

[0102] Please see the attached Figure 8,Data storage and management module is used to store the generated structured reports.

[0103] The data storage unit is used to categorize and store generated structured reports based on generation time or other pre-defined criteria. Whenever nodule analysis results are converted into a JSON-formatted structured report through the structured report output module, the data storage unit stores the report and ensures its efficient and stable storage in the database.

[0104] Typically, structured reports are categorized and stored by generation date. For example, reports can be archived and stored in a "year-month-day" format, facilitating subsequent queries based on time ranges. Furthermore, the data storage unit can further categorize reports based on report type, patient information, nodule risk level, and other factors to ensure efficient data access during subsequent searches.

[0105] In some embodiments, the data storage unit uses a database management system (DBMS) for storage, such as a relational database (MySQL, PostgreSQL, etc.) or a non-relational database (MongoDB, Cassandra, etc.). By storing data in a structured manner, the data storage unit can conveniently store, manage, and access large amounts of data.

[0106] Assume the structured report is , the stored data is , then the data storage process can be expressed as: ; in: The final structured report after formatting; For data stored in the database; The time when the report was generated.

[0107] The data retrieval unit supports users in searching, querying, and downloading stored reports. Users can quickly retrieve the required reports based on multiple criteria, such as patient name, report generation time, and nodule risk level. This retrieval function improves data access efficiency and facilitates doctors and other personnel in obtaining the information they need during actual use.

[0108] In some embodiments, the data retrieval unit provides retrieval functions based on various criteria, such as keywords, time ranges, and report types. Retrieval functions can be flexibly combined based on user needs. For example, a user can enter a patient's name to retrieve all reports for a specific patient, or specify a time range to query all reports generated within a certain period. Furthermore, the data retrieval unit can support paging queries to address the query pressure brought about by large-scale data storage.

[0109] Data retrieval units typically use indexing technology to improve query efficiency. By indexing fields in the database, the system can significantly speed up queries and avoid query delays caused by large amounts of data.

[0110] Assume that the stored data is , the search conditions are The query report is , then the data retrieval process can be expressed as: ; in: For stored data; is the search criteria, such as name, time range, etc.; For search results, a structured report that meets the criteria is returned.

[0111] The data security management unit is used to securely protect stored reports, ensuring that data is not leaked, tampered with, or lost during storage, transmission, and access. Medical data often contains patients' personal information and sensitive health data, so ensuring data security is a core task of this module.

[0112] In some embodiments, the data security management unit implements various security measures to protect data, such as encrypted storage, permission control, and access logging. First, the data security management unit can encrypt report data stored in the database, ensuring that the data cannot be read even if illegally accessed while stored. Second, through fine-grained permission control, the system can ensure that only authorized users (such as doctors and administrators) can access specific data, preventing unauthorized access. To further enhance data security, the system can also record access logs to track who accessed which data and when, and promptly issue alarms when anomalies occur.

[0113] In addition, the data security management unit can also perform regular data backup to ensure that data can be restored in the event of system failure or data loss, thereby ensuring high system availability and data integrity.

[0114] Assume that the stored data is , the encrypted data is , then the data encryption process can be expressed as: ; in: Structured reports for storage; The report data is encrypted to ensure data security.

[0115] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A tool for interpreting and analyzing lung nodule reports based on a large model, characterized by: include: An electronic file upload module is used to receive electronic files of pulmonary nodule reports uploaded by users; The OCR recognition module is used to perform optical character recognition on uploaded electronic files and output preliminary text data; Text pre-analysis module, used to detect and clean the interference of text data after OCR recognition; Key content extraction module, used to extract the patient's name, gender, age, hospital, report date, and nodule size, density, location, and morphology from text data; Report content analysis module, used to analyze the extracted nodule information and infer the characteristics and risk level of the nodules; The structured report output module is used to convert nodule analysis results into structured reports and output standardized JSON format reports; The data storage and management module is used to store the generated structured reports.

2. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The electronic file upload module includes: The file receiving unit is used to receive electronic files uploaded by users and perform preliminary identification based on the file format; The format conversion unit is used to convert uploaded files of different formats into a unified format for use by the OCR recognition module.

3. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The OCR recognition module includes: Image preprocessing unit, used to preprocess the uploaded image files, eliminate noise, adjust image clarity and perform rotation correction; A character segmentation unit, used for segmenting the text in the image; The text recognition unit is used to recognize characters through convolutional neural networks and convert the text in the image into processable text data; The recognition result output unit is used to output the preliminary text data after OCR recognition for use by the subsequent text pre-analysis module.

4. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The text pre-analysis module includes: The interference detection unit is used to calculate the interference of the OCR text based on the clarity, font size and other characteristics of the text; Angle adjustment detection unit, used to detect text tilt or image deviation in the OCR recognition results and determine whether the OCR recognition angle needs to be adjusted; The pre-processing adjustment unit is used to guide the OCR module to adjust the angle or parameters of text recognition according to the interference detection result.

5. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The key content extraction module includes: Semantic analysis unit, used to perform semantic analysis on the text using a large model to identify key information in the report, including the patient's name, gender, age, hospital, report date, and the size, density, location, and morphology of the nodule; The prompt word matching unit is used to match relevant content in the text based on the set prompt word library and extract the size, density, location and morphology of the nodule; The information standardization unit is used to standardize the extracted information.

6. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The report content analysis module includes: Nodule feature inference unit, which is used to analyze the morphology and possible risk type of nodules based on the characteristics of nodule size, density, location, etc. through inference using a large model; Nodule risk assessment unit, which is used to evaluate the risk level of nodules and generate relevant diagnostic results by combining the pulmonary nodule knowledge base and pulmonary nodule guidelines; The analysis result output unit is used to pass the analysis results to the structured report output module to generate a detailed nodule diagnosis report.

7. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The structured report output module includes: The result conversion unit is used to convert the analysis results into a structured data format, JSON format, to ensure good data compatibility; Formatting output unit, used to format structured data into a final report, including the patient's basic information, nodule characteristics and diagnosis results; The result display unit is used to display the final report to the user for subsequent use by doctors or related personnel.

8. The tool for interpreting and analyzing lung nodule reports based on a large model according to claim 1, characterized in that: The data storage and management module includes: A data storage unit, used for classifying and storing structured reports according to generation time; Data retrieval unit, used to support users to retrieve, query and download stored reports; The data security management unit is used to provide security protection for stored reports.