Pulmonary nodule disease body dynamic underwriting method and device based on multi-modal data fusion

The dynamic underwriting method for lung nodule-bearing patients based on multimodal data fusion solves the problems of incomplete data analysis and static risk assessment in traditional underwriting, achieves more accurate risk assessment and flexible underwriting strategies, and improves underwriting efficiency and user satisfaction.

CN120673960APending Publication Date: 2025-09-19ZHEJIANG HUAFANG RUIBAO TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510577127.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The underwriting process of traditional critical illness insurance and medical insurance has problems such as incomplete data analysis, static risk assessment, inefficient process, and biased underwriting conclusions. In particular, the dynamic risk assessment of the sick population is insufficient, and there is a lack of data sharing and industry collaboration, resulting in inaccurate product pricing and low user satisfaction.

Method used

A dynamic underwriting method for lung nodules with diseased bodies based on multimodal data fusion is adopted. By obtaining user basic information and chest CT report images, OCR recognition and verification are performed to generate structured JSON data. The mapping rule library is used to standardize terminology. CT parameters, dynamic follow-up data and external risk factors are combined to input the lung cancer risk probability model, and step-by-step matching is performed to generate underwriting conclusions.

Benefits of technology

It improves underwriting efficiency and the accuracy of conclusions, can flexibly adjust underwriting strategies according to different insurance product types, enhances the model's adaptability and robustness to various input data, and improves customer satisfaction and trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673960A_ABST
    Figure CN120673960A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pulmonary nodule disease body dynamic underwriting method and device based on multi-modal data fusion, and the method comprises the steps: obtaining the basic information of a user and a chest CT report image, carrying out the OCR recognition and verification of the chest CT report image, and outputting structured JSON data; performing standardization processing and semantic extension on the medical terms extracted from the structured JSON data according to the mapping rule base to generate standardized report data; inputting the CT parameters, the dynamic follow-up visit data and the external risk factors in the report data into a lung cancer risk probability model, and calculating a lung cancer risk probability value; and step-by-step matching is carried out on the lung cancer risk probability value and other nodule characteristics with the underwriting rule, and a structured underwriting conclusion is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method and device for dynamic underwriting of diseased bodies with pulmonary nodules based on multimodal data fusion. Background Art

[0002] Traditional critical illness and medical insurance underwriting processes often use complex and obscure health disclosure questionnaires, requiring policyholders to self-report their health status. This process is time-consuming and cumbersome, leaving policyholders vulnerable to premium increases, exclusions, or even denials due to inadequate understanding or omissions. Furthermore, traditional underwriting, which primarily relies on health disclosures, lacks sufficient medical data support, making it difficult to accurately assess the complex risks of pre-existing conditions. This leads to inaccurate product pricing, terms, and benefits. Underwriting medicine focuses more on potential future risks rather than current health status, resulting in limited understanding and underwriting experience for pre-existing conditions and a lack of standardized risk stratification logic. This can lead to users being unable to purchase insurance due to underwriting restrictions or being deterred by high premiums. Insurance companies are collaborating with medical institutions to obtain de-identified medical data to build underwriting and pricing models for pre-existing conditions, thereby supporting the design of products that cover pre-existing conditions. While this approach can help insurers better assess pre-existing risk, the data barriers between the medical, pharmaceutical, and insurance industries have not been fully broken down, resulting in a lack of collaboration and hindering risk model optimization and product innovation. The underwriting process relies on manual review and requires the insured to provide physical examination reports, medical records and other materials. Although risks can be assessed in multiple dimensions, such as considering medical history and living habits, the process is inefficient and the subjective judgments of different underwriters may vary, affecting the consistency and fairness of the underwriting results.

[0003] With the application of intelligent underwriting technology, OCR (optical character recognition) and NLP (natural language processing) technologies are used to automatically analyze medical reports, and risk stratification models are built in combination with claims data, thereby improving underwriting efficiency. However, intelligent underwriting still relies on preset rules and has weak adaptability to new diseases or complex cases. For example, the conclusions of AI models are mostly standard coverage, premium increases, or rejections, lacking room for flexible adjustment, especially for the dynamic risk assessment of patients with chronic diseases. The AI ​​model is not sufficiently trained for new diseases and health risks and still needs to be supplemented by manual underwriting. Therefore, it is necessary to combine data sharing, algorithm optimization, and industry collaboration to promote the standardization of risk stratification, dynamic underwriting decision-making, and personalized product design to meet the urgent medical insurance needs of the sick population. Summary of the Invention

[0004] In response to the four core problems existing in the existing underwriting process for patients with diseases, namely "incomplete data analysis, static risk assessment, low process efficiency, and biased underwriting conclusions", the embodiments described in this article provide a dynamic underwriting method and device for patients with lung nodules based on multimodal data fusion, as well as a computer-readable storage medium storing a computer program.

[0005] According to the first aspect of the present disclosure, a dynamic underwriting method for lung nodule-carrying patients based on multimodal data fusion is provided, comprising: obtaining basic information of the user and chest CT report images, performing OCR recognition and verification on the chest CT report images, and outputting structured JSON data; standardizing and semantically expanding the medical terms extracted from the structured JSON data according to a mapping rule library to generate standardized report data; inputting CT parameters, dynamic follow-up data, and external risk factors in the report data into a lung cancer risk probability model to calculate the lung cancer risk probability value; and matching the lung cancer risk probability value and other nodule characteristics with the underwriting rules step by step to generate a structured underwriting conclusion.

[0006] In some embodiments of the present disclosure, obtaining a user's basic information and chest CT report image, performing OCR recognition and verification on the chest CT report image, and outputting structured JSON data include: obtaining the insured's name, date of birth, gender, smoking history, and family history of lung cancer in immediate relatives entered by the user in the client interface; obtaining a chest CT report image uploaded by the user, the chest CT report image containing complete patient information, examination institution name, examination date, nodule description, and doctor's signature information; checking whether the user input information is complete, whether the data format is correct, whether the image file format and size meet the requirements, and whether the content of the image report is true; performing report integrity verification, identity consistency verification, and follow-up time verification on the user's basic information and chest CT report image according to a preset verification rule library, and outputting the verification results; performing OCR recognition on the chest CT report image that passes the verification, extracting text data, and converting the text data into structured JSON data.

[0007] In some embodiments of the present disclosure, report integrity verification, identity consistency verification and follow-up time verification are performed on the user's basic information and chest CT report image according to a preset verification rule library, and the output verification results include: checking whether the chest CT report image contains the CT examination conclusion, doctor's signature, examination agency name, examination date information, and verifying whether the examination date meets the validity period requirements; comparing whether the patient's name, gender, and age in the chest CT report image are consistent with the identity information entered by the user; verifying the timeliness of the report based on the number and type of nodules, and judging whether the follow-up time requirements are met by calculating the time difference between the examination dates.

[0008] In some embodiments of the present disclosure, performing OCR recognition on a verified chest CT report image, extracting text data, and converting the text data into structured JSON data includes: preprocessing the chest CT report image for denoising and contrast enhancement; using a layout analysis model to locate key areas in the report of the preprocessed chest CT report image, the key areas including the patient information area, the nodule description area, and the doctor's signature area; using a deep learning OCR model to extract text data in the key areas, and converting the text data into structured JSON data.

[0009] In some embodiments of the present disclosure, medical terms extracted from structured JSON data are standardized and semantically expanded according to a mapping rule library to generate standardized report data, including: establishing a standardized mapping rule library for medical terminology of pulmonary nodules, and the mapping rule library stores mapping rules including original keywords, mapping targets, weights, and applicable scenarios; based on the mapping rules, non-standard terms in the report are converted into unified medical terms, synonyms and negative descriptions are identified and processed uniformly, and when there are multiple versions of the description of the same nodule, the most accurate description is selected according to priority to obtain a mapping result; and the mapping result is generated into report data in a standardized JSON format.

[0010] In some embodiments of the present disclosure, CT parameters, dynamic follow-up data, and external risk factors in the report data are input into a lung cancer risk probability model, and calculating the lung cancer risk probability value includes: inputting the CT parameters, dynamic follow-up data, and external risk factors in the report data into a pre-trained lung cancer risk probability model for multimodal data fusion analysis to calculate the patient's risk probability value for lung cancer, the CT parameters include the size of the nodule's long axis / short axis, malignant signs and their combination, the number of nodules, and the location of the nodules; the dynamic follow-up data include the annual growth rate of nodule volume; and the external risk factors include the patient's age, gender, smoking history, and family history data; the patient's lung cancer risk is divided into different levels according to the risk probability value, and the key factors affecting the risk probability value are output and listed.

[0011] In some embodiments of the present disclosure, CT parameters, dynamic follow-up data, and external risk factors in the report data are input into a pre-trained lung cancer risk probability model for multimodal data fusion analysis to calculate the patient's risk probability value for lung cancer, including: quantifying the synergistic effect between malignant signs based on the combination of malignant signs in the CT parameters; introducing a time attenuation factor into the dynamic follow-up data to dynamically adjust the weight of the historical dynamic follow-up data.

[0012] In some embodiments of the present disclosure, lung cancer risk probability values ​​and other nodule features are matched with underwriting rules step by step to generate a structured underwriting conclusion, including: matching nodule features and risk probability values ​​with underwriting rules step by step according to the order of priority from high to low, such as rejection, exclusion, surcharge, and standard body, and outputting a structured underwriting conclusion, which includes whether the surcharge conditions, surcharge ratio, exclusions, special agreements, and push objects are met; attaching a risk quantification basis based on the risk probability value to the underwriting conclusion, and providing an explanation compared with the clinical guidelines; conducting closed-loop feedback on user data collected during the underwriting process and subsequent claims data to continuously optimize the risk assessment model; triggering manual review if any of the following situations occurs: the risk probability value is within the boundary value range, the historical imaging data provided by the user is inconsistent with the current report, the output conclusions of multiple rules for the same case are contradictory, and the confidence level of the OCR recognition result is lower than the preset threshold.

[0013] According to the second aspect of the present disclosure, a dynamic underwriting device for lung nodule-bearing diseased bodies based on multimodal data fusion is provided. The device includes at least one processor; and at least one memory storing a computer program. When the computer program is executed by at least one processor, the device can obtain the user's basic information and chest CT report image, perform OCR recognition and verification on the chest CT report image, and output structured JSON data; perform standardization and semantic expansion on the medical terms extracted from the structured JSON data according to the mapping rule library to generate standardized report data; input the CT parameters, dynamic follow-up data and external risk factors in the report data into the lung cancer risk probability model to calculate the lung cancer risk probability value; and match the lung cancer risk probability value and other nodule characteristics with the underwriting rules step by step to generate a structured underwriting conclusion.

[0014] According to a third aspect of the present disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program implements the steps of the method according to the first aspect of the present disclosure when executed by a processor.

[0015] According to the dynamic underwriting method and device for lung nodule-bearing diseased bodies based on multimodal data fusion provided by the embodiment of the present disclosure, by using OCR technology to automatically extract information from chest CT reports and perform verification to ensure data accuracy and consistency, and by standardizing the extracted medical terms through a mapping rule base, it can solve the problem of inconsistent terminology that may exist between different institutions and different reports, and can enhance the adaptability and robustness of the model to various input data. The risk assessment model can comprehensively consider different types of factors, and by fusing multimodal data, it can provide a more comprehensive lung cancer risk assessment, which is more accurate than a single-factor model. According to the lung cancer risk probability value and other nodule characteristics and different underwriting rules, the risk threshold can be automatically adjusted according to different insurance product types, so that insurance companies can flexibly adjust underwriting strategies according to the needs of specific products, thereby improving customer satisfaction and trust. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be noted that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure.

[0017] Figure 1 is an exemplary flow chart of a dynamic underwriting method 100 for lung nodules with disease based on multimodal data fusion according to an embodiment of the present disclosure;

[0018] Figure 2 It is a schematic block diagram of a dynamic underwriting device 200 for diseased bodies with lung nodules based on multimodal data fusion according to an embodiment of the present disclosure.

[0019] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative work also fall within the scope of protection of the present disclosure.

[0021] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the subject matter of the present disclosure belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal manner unless otherwise explicitly defined herein. In addition, terms such as "first" and "second" are only used to distinguish one component (or a portion of a component) from another component (or another portion of a component).

[0022] Figure 1 An exemplary flow chart of a dynamic underwriting method 100 for lung nodules with diseased bodies based on multimodal data fusion according to an embodiment of the present disclosure is shown. Figure 1 In step S102, the user's basic information and chest CT report image are obtained, the chest CT report image is OCR recognized and verified, and structured JSON data is output.

[0023] According to one embodiment of the present disclosure, the user fills in the basic information form through the client, such as a mini program or H5 page, and uploads the chest CT report image. The basic information form includes basic information such as the insured's name, date of birth, gender, smoking history, whether there are immediate relatives (such as parents, siblings, etc.) with lung cancer, etc. Smoking records include data such as whether smoking, years of smoking, average number of cigarettes per day, etc. The user needs to upload 2-4 chest CT report images. The report image file is in JPG or PNG format. The size of a single image file shall not exceed 5MB. The image content contains complete patient information (such as name, gender, age), the name of the examination institution, the examination date, nodule description (such as the long axis / short axis length of the nodule, the location and nature of the nodule, whether there are malignant features such as burr signs, lobulation signs, pleural traction, etc.), doctor's signature, etc.

[0024] After receiving this information, the system checks the completeness of the user's input, the correct data format, the image file format and size, and the authenticity of the image report. First, the system verifies that the required fields on the basic information form are complete and the data format is correct. For example, the date of birth should be in the standard format (YYYY-MM-DD) and the number of years smoked should be a positive integer. The uploaded image file is checked for the required format (JPG or PNG) and the file size is no larger than 5MB. If there are any file format errors or the file size exceeds the limit, the system prompts the user to re-upload a file that meets the requirements. The system checks the digital watermark in the image file to ensure that it is legitimate and has not been tampered with. The authenticity of the image file is verified by checking the image metadata, including the capture device information and file creation time, to ensure that the image creation time matches the inspection date in the report to prevent forgery. The image file undergoes digital watermark detection and metadata verification (such as device information and capture time) to ensure the authenticity of the report. If the report information is incomplete or questionable, the system prompts the user to modify or supplement it.

[0025] The system then verifies the user's basic information and chest CT report images for report integrity, identity consistency, and follow-up time based on a pre-set validation rule base, outputting the verification results. The verification results determine whether the report can proceed to the next steps. If the verification fails, an error code and prompt message are generated to help users quickly identify and correct the problem.

[0026] According to one embodiment of the present disclosure, the verification rule base is shown in the following table:

[0027]

[0028] As shown in the table above, report integrity verification includes checking whether the examination report contains key information such as the CT examination conclusion, physician signature, examination institution name, and examination date. Missing key fields such as "Institution Name" or "Examination Date" triggers an "Incomplete Data" warning. Verify that the examination date in the report meets validity requirements, for example, whether the examination date is within the valid insurance period or meets the latest examination requirements.

[0029] Compare the patient's name, gender, and age in the report to the patient information provided by the user to ensure identity matching. Cross-validation through multi-source data ensures the authenticity of the information. Use an accurate matching algorithm to ensure that name differences (such as "Zhang San" vs. "Zhang San") do not lead to false matches.

[0030] According to the number and type of nodules, verify the timeliness of the report. Determine whether the follow-up time requirements are met by calculating the time difference between the inspection dates. For example, if it is a single nodule, the interval between the last two inspections is required to be ≥1 year. That is, when a patient has a single nodule, the time interval between follow-up inspections must be at least 1 year. If it is multiple nodules, the interval between the last two inspections is required to be ≥2 years. For patients with multiple nodules, the inspection interval is required to be longer. Whether it is a single nodule or multiple nodules, check whether the date of the last inspection report is within 6 months to ensure that the follow-up requirements are met. In order to avoid misjudgment due to different date formats (such as "August 20, 2023" and "8 / 20 / 2023"), the date format needs to be unified and normalized. By converting the date format, ensure that all dates can be compared uniformly during verification to avoid verification failure due to format differences.

[0031] If any rule fails the verification, the system will output an error code and prompt information to guide the user to upload a correct or updated report. If all verification items meet expectations, subsequent processing will be carried out.

[0032] Finally, the chest CT report images that have passed the verification are subjected to OCR recognition to extract text data and convert the text data into structured JSON data. Image preprocessing is the first step in OCR recognition. First, the chest CT report images are subjected to image preprocessing, including data enhancement processing such as denoising and contrast adjustment to improve image clarity and the contrast between text and background, making the text more prominent and easier to identify. The layout analysis model is used to locate key areas in the preprocessed chest CT report images, such as the patient information area, nodule description area, and doctor signature area, for subsequent recognition. The patient information area contains basic information such as name, gender, and age. The nodule description area contains a detailed description of the nodules in the CT image, such as the location, size, and nature of the nodules. The doctor signature area is the signature or stamp area of ​​the report.

[0033] A deep learning OCR model is used to extract text information in key areas. In medical images, special challenges are often encountered, such as medical abbreviations (such as "GGO"), handwritten signatures, fuzzy seals, etc. According to one embodiment of the present disclosure, an OCR model is trained using training data to identify medical abbreviations, handwritten signatures, and fuzzy seals. Among them, the training data is enhanced (such as rotation, scaling, blurring, etc.) to improve the robustness of the OCR model in practical applications, so that the trained OCR model can not only extract printed text, but also recognize handwriting, seals, and medical-specific terms. After the OCR model is recognized, errors can be further reduced through secondary proofreading and manual review.

[0034] The extracted text is output as structured JSON data to facilitate subsequent processing and analysis. For example, the output JSON data structure contains the following information:

[0035]

[0036] Then, in step S104 , the medical terms extracted from the structured JSON data are standardized and semantically expanded according to the mapping rule library to generate standardized report data.

[0037] According to one embodiment of the present disclosure, a standardized mapping rule library is established for medical terminology related to pulmonary nodules. The mapping rule library stores mapping rules including original keywords, mapping targets, weights, and applicable scenarios. Original keywords refer to the terms or descriptions in the report, mapping targets refer to the standardized or unified medical terms, weights refer to the credibility or priority of the mapping, with higher weights indicating higher accuracy or priority of the mapping, and applicable scenarios refer to the specific situations to which the mapping rule applies, such as describing nodule morphology or determining malignant signs.

[0038] The mapping rule base is in the form of a database table, which stores all the term mapping rules. For example, the database table structure is as follows:

[0039] Original keywords Mapping Target Weight Applicable Scenarios "glitch" "Burr sign" 0.9 Nodule morphology description "Pleural indentation" "Pleural sign" 0.85 Judgment of malignant signs "Partial reality" "mixedGGO" 1.0 International Standard Terminology Mapping

[0040] The library helps convert various non-standard terminologies into unified standard terminologies and prioritizes them according to weights.

[0041] Specific mapping rules include converting non-standard terms in reports into standardized medical terminology, such as converting "rough margins" to "spicule sign" and "GGO" to "ground-glass opacity." This standardization avoids data inconsistencies caused by different expressions used by different physicians or systems. In some cases, a term may have different expressions or its meaning may change when co-occurring with other terms. Synonyms and negated descriptions can be identified and handled uniformly. For example, "pleural indentation" may have different expressions in reports; it should be uniformly mapped to "pleural indentation," while "absent lobulation sign" should be marked as a negated result. When multiple descriptions of the same nodule exist, the most accurate description is selected based on priority. To ensure consistency and accuracy in report content, the system uses a conflict resolution algorithm to automatically resolve conflicts between different descriptions. For example, the nodule morphology description in the report may differ from the physician's diagnosis. In this case, the system will prioritize the description in the report because it better aligns with the imaging evidence. When the user's historical report uses the term "burr sign" but the new report uses "rough edge", the system will give priority to "burr sign" based on the weight of the historical data to avoid misunderstandings caused by inconsistent expressions. The unit conversion of numerical data such as nodule size is unified to avoid errors, such as automatically converting "0.8cm" to "8mm". The mapping results are generated into report data in a standardized JSON format. The mapping results include the patient's basic information, nodule characteristics, risk warnings, and review records. The example output is as follows:

[0042]

[0043]

[0044] In the above report data, nodule characteristics are mapped to standardized terms such as "MIXED." Spiculation, lobulation, and pleural signs are marked as true, indicating that these features are present in the report. Review records include historical nodule size and the time interval between two examinations (in days).

[0045] You can set up mandatory extraction of 25 key parameters, including but not limited to the nodule's long / short axis, the combination of malignant signs, and follow-up time. These are core data that influence lung nodule risk assessment. By forcing the extraction of these key parameters, we ensure that all key data are taken into account, thus avoiding risk assessment bias caused by missing data or neglect of certain features.

[0046] Then, in step S106, the CT parameters, dynamic follow-up data, and external risk factors in the report data are input into the lung cancer risk probability model to calculate the lung cancer risk probability value.

[0047] The disclosed embodiment comprehensively considers more factors that affect the risk of lung cancer by expanding the input dimensions. Input parameters include CT parameters, dynamic follow-up data and external risk factors. CT parameters include the size of the long axis / short axis of the nodule, malignant signs (such as burr sign, lobulation sign, etc.) and their combination, the number of nodules, the location of the nodules (such as upper lung / lower lung), etc. These parameters help describe the basic characteristics of the nodules and are directly related to the probability of lung cancer. Dynamic follow-up data include time series data such as the annual growth rate of nodule volume. The annual growth rate of nodule volume is calculated based on historical data and reflects the growth trend of the nodule. External risk factors include known external factors such as the patient's age, gender, smoking history (such as years, average number of cigarettes per day), family history, etc. The risk of lung cancer increases with age. The incidence of lung cancer varies by gender. Men are generally more likely to develop lung cancer than women. Smoking is one of the most important risk factors for lung cancer. Years of smoking and average number of cigarettes per day can help assess the risk. The above parameters are input into the pre-trained lung cancer risk probability model for multimodal data fusion analysis, and a risk value (ie, RB value) of 0-100% is output, which represents the risk probability of the patient suffering from lung cancer.

[0048] Since the combination of multiple malignant signs may be far more dangerous than a single sign, this solution introduces combined feature enhancement to quantify the synergistic effect between malignant signs based on the combination of malignant signs. For example, when the two features "lobulation sign" and "pleural traction" appear simultaneously in the same nodule, its malignancy probability is 2.3 times that of a single feature. By quantifying the synergistic effect between these features, the complex characteristics and interrelationships of lung nodules can be effectively captured, thereby improving the accuracy of risk prediction. In addition, the impact of historical examination data is usually static and fails to fully consider the risks that change over time. In the embodiment of the present disclosure, a time decay factor is introduced into the dynamic follow-up data (such as nodule volume growth rate, review interval) to dynamically adjust the weight of historical data. For example, if the review interval exceeds 2 years, the weight of this parameter will be reduced (for example, halved, i.e., ×0.5), and the follow-up data 2 years ago will only have a weight of 30% of the latest data when assessing the current nodule risk. In this way, the model can more accurately reflect the impact of the latest data on risk assessment and avoid the adverse effects of outdated data. The introduction of time series modeling makes assessments more dynamic, enabling real-time updates of risk assessment results to reflect changes in a patient's condition. Multimodal data fusion makes risk assessments more dynamic and accurate, avoiding missed assessments of high-risk users due to ignoring dynamic factors and enabling real-time adjustments to risk assessments based on the latest data.

[0049] After calculating the risk probability value, the patient's lung cancer risk is categorized into different levels based on the risk probability value, and the key factors influencing the risk probability value are listed. For example, a risk probability value (RB value) of 6.1356 indicates a 6.14% risk of lung cancer. Based on the RB value, the patient's lung cancer risk is categorized as low, medium, or high. Risk levels are determined based on pre-set risk thresholds. For example, an RB value between 0% and 5% might be considered "low risk," a RB value between 5% and 20% might be considered "medium risk," and a RB value exceeding 20% ​​might be considered "high risk." To help doctors and patients understand the main sources of risk, key factors influencing the RB value are listed. For example, if "family history of lung cancer," "spicule sign + pleural traction," and "annual volume growth rate of 15%" are major risk factors, these factors will be listed and used as a reference for further decision-making by the user and doctor.

[0050] For example, the model output is as follows:

[0051]

[0052] The RB value represents the probability risk of lung cancer. A value of 6.1356 indicates a 6.14% risk for this patient. The risk level is determined based on the range of RB values, for example, "medium risk" in this case. The primary risk factors indicate key factors influencing a patient's risk, including: "family history of lung cancer," which indicates a higher risk for patients with a family history. "spicule sign + pleural traction," a combination of these malignant signs, further increases the risk of lung cancer. "Annual volume growth rate of 15%," also indicates a rapid increase in nodule volume, which is also a high-risk indicator.

[0053] Finally, in step S108, the lung cancer risk probability value and other nodule characteristics are matched with the underwriting rules step by step to generate a structured underwriting conclusion.

[0054] The underwriting rules contain a series of judgment conditions based on patient data and characteristics (such as nodule size, RB value, malignant signs, etc.), and assign a specific underwriting conclusion to each condition. These rules can be dynamically adjusted according to different product requirements to ensure compliance with the requirements of various insurance products. The underwriting rule library supports online editing and updating, which means that underwriting rules can be quickly added, modified or deleted without stopping the system. For example, for new health risks (such as "post-COVID-19 pulmonary fibrosis"), the rules can be quickly adjusted and effective without downtime, ensuring the flexibility and adaptability of the system. The exemplary underwriting rules are as follows:

[0055]

[0056]

[0057] The rule matching engine matches nodule information and RB values ​​against underwriting rules in descending order of priority: denial, exclusion, premium surcharge, and standard case. It then outputs a structured underwriting conclusion, which includes whether the premium surcharge conditions are met, the premium surcharge percentage, exclusions, special agreements, and the recipients of the policy. Denial has the highest priority within the underwriting rules. If the denial conditions are met, the policy is immediately rejected. If the conclusion is an exclusion, lung cancer coverage is excluded. Next is premium surcharge coverage. Finally, if none of the above conditions are met, the policy is considered standard case coverage. For example, if the RB value is 2.5% and the nodule's long axis is 12mm, the system first checks to see if the "Malignant Sign Combination" or "Nodule Long Axis Threshold" rules are triggered. For risk values ​​(RB values) between 1% and 2%, a 10% premium surcharge or the addition of an annual CT scan clause can be selected. This flexible underwriting approach can reduce customer disputes caused by perceived differences and improve user acceptance and satisfaction. If the long axis of the nodule exceeds the threshold of 10mm, it will trigger the "exclusion of liability" and the system will output "exclusion of lung cancer liability" and terminate further judgment.

[0058] In order to make the underwriting process more flexible and targeted, the embodiment of the present disclosure uses a dynamic underwriting rule engine to automatically match risk thresholds according to different product types. For example, for inclusive insurance products, the underwriting standards are relatively loose, and a higher RB value (lung cancer risk value) is allowed. For example, it may be allowed to be underwritten when the RB value is ≤5%. For high-end medical insurance products, due to the higher underwriting risk, the risk value requirements are more stringent, requiring a RB value of ≤0.8% to be underwritten. The dynamic underwriting rule engine can automatically adjust the underwriting risk threshold according to the characteristics of the product, thereby optimizing the underwriting strategy, which can not only improve the flexibility of insurance products, but also ensure the risk control capabilities of different types of insurance.

[0059] Based on the judgment results of the rule matching engine, the system generates an underwriting conclusion and pushes the structured results to the insurance company's core system. The conclusion is not just "underwriting" or "rejection" but also includes additional conditions (such as surcharge percentages, special agreements, etc.) for the insurance company to refer to during the underwriting process. The underwriting output example is as follows:

[0060]

[0061] In order to bridge the cognitive gap between underwriting medicine and clinical medicine, the underwriting conclusion is supplemented with a risk quantification basis based on the RB value, and an explanation is provided in comparison with clinical guidelines. For example, if the user's lung nodules have lobulation signs and a family history of lung cancer, an additional explanation can be provided as "RB value = 4% due to lobulation signs + family history of lung cancer." This quantitative data helps clinicians and users better understand the basis of the conclusion and improves the transparency of the underwriting conclusion. The underwriting conclusion not only provides quantitative data, but can also be compared with recognized clinical guidelines (such as the Fleischner Society guidelines) to enhance the credibility of the conclusion. For example, the underwriting conclusion can state that "according to the Fleischner Society guidelines, nodules that have not changed for 3 years are generally considered low risk, but in certain specific cases further follow-up examinations may still be required."

[0062] Furthermore, user data collected during the underwriting process (such as RB values ​​and underwriting conclusions) is integrated with subsequent claims data in a closed-loop feedback loop to continuously optimize the risk assessment model and verify whether combinations of imaging features, such as the "spicule sign + cavitation sign," can effectively predict the malignancy rate of nodules. As data accumulates, the model becomes more precise, thereby improving the accuracy of underwriting conclusions. Structured underwriting data is also used to train AI models and optimize CT scanning protocols. For example, scanning protocols can be adjusted based on specific features such as the nodule's location and size, thereby improving image clarity and accuracy and better supporting underwriting decisions. Premium gradients are dynamically adjusted based on the RB value distribution derived by the model. For example, for every 0.5% increase in RB value, the premium increases by 3%. This dynamic pricing approach based on risk-based values ​​more accurately reflects the user's actual risk and avoids the "extensive premium increase" model common in traditional insurance products. For insured patients with lung nodules, personalized health management services, such as CT review reminders and smoking cessation programs, are provided to help them reduce long-term health risks, thereby enhancing the product's added value.

[0063] If the risk probability value is within the boundary value range, or there is inconsistency between the historical image data provided by the user and the current report, or the output conclusions of multiple rules for the same case are contradictory, or the confidence of the OCR recognition result is lower than the preset threshold, manual review will be triggered.

[0064] According to one embodiment of the present disclosure, the determination rules for automatic transfer to manual verification are shown in the following table:

[0065]

[0066] As shown in the table above, when the RB value is 0.9%, the system will automatically mark it as requiring manual review. If the confidence level of the OCR recognition result is lower than 85%, or if certain key fields (such as nodule size) have multiple versions of descriptions (for example, the "long axis" in the CT report is described as 8mm and 0.8cm in two different units), manual review will be triggered. For example, a surcharge is triggered when the RB value is 1.1%, but the long axis of the nodule is 9.9mm, which is close to the threshold of 10mm. The exclusion rule may also be triggered. This situation requires manual judgment and confirmation. For example, if the 2022 report shows that the long axis of the nodule is 12mm, and the 2023 report shows it is 8mm, and this reduction exceeds the medically reasonable range, the system will mark it as an abnormality and require manual review.

[0067] During the manual review stage, the reviewer needs to manually verify the pixel-level alignment of the OCR recognition results with the original image to ensure that key data such as the nodule size in the report is accurate. If there is an inconsistency between the OCR recognition results and the original report content, the reviewer needs to make further confirmation. The manual review interface displays a detailed judgment path for rule conflicts, allowing reviewers to understand how the system reaches the current conclusion and check whether there are conflicting rules that need to be resolved. It also provides multi-dimensional data comparison views, such as historical nodule size change curves, CT images at different time points, etc., so that manual underwriters can comprehensively assess the patient's condition. During the review process, manual reviewers also need to view the model's contribution analysis data for different features to determine whether the model has reasonably assessed high-risk parameters, especially the impact of risk factors such as family history of lung cancer and malignant signs.

[0068] The system automatically determines underwriting conclusions for lung cancer insurance through an underwriting rule library and rule-matching engine, generating detailed underwriting or rejection decisions based on the patient's specific circumstances (such as nodule size and RB value). To address complex situations such as boundary values, data inconsistencies, and rule conflicts, the system also incorporates a manual review mechanism to ensure the accuracy and rationality of the conclusions. Ultimately, the structured underwriting conclusions are pushed to the insurance company's core systems to support further underwriting decisions.

[0069] Figure 2 This is a schematic block diagram of a dynamic underwriting device for lung nodules with diseased bodies based on multimodal data fusion according to an embodiment of the present disclosure. Figure 2 As shown, the apparatus 200 may include a processor 210 and a memory 220 storing a computer program. When the computer program is executed by the processor 210, the apparatus 200 may perform the following operations: Figure 1The steps of method 100 are shown. In one example, device 200 can be a computer device or a cloud computing node. Device 200 can obtain the user's basic information and chest CT report image, perform OCR recognition and verification on the chest CT report image, and output structured JSON data; perform standardization and semantic expansion on the medical terms extracted from the structured JSON data according to the mapping rule library to generate standardized report data; input the CT parameters, dynamic follow-up data and external risk factors in the report data into the lung cancer risk probability model to calculate the lung cancer risk probability value; and match the lung cancer risk probability value and other nodule characteristics with the underwriting rules step by step to generate a structured underwriting conclusion.

[0070] According to one embodiment of the present disclosure, the device 200 can obtain the insured's name, date of birth, gender, smoking history, and family history of lung cancer in immediate relatives entered by the user in the client interface; obtain the chest CT report image uploaded by the user, which contains complete patient information, examination institution name, examination date, nodule description and doctor's signature information; check whether the user-input information is complete, whether the data format is correct, whether the image file format and size meet the requirements, and whether the content of the image report is true; perform report integrity verification, identity consistency verification and follow-up time verification on the user's basic information and chest CT report image according to a preset verification rule library, and output the verification results; and perform OCR recognition on the chest CT report image that passes the verification, extract text data, and convert the text data into structured JSON data.

[0071] According to one embodiment of the present disclosure, the device 200 can check whether the chest CT report image contains the CT examination conclusion, doctor's signature, examination institution name, and examination date information, and verify whether the examination date meets the validity requirements; compare whether the patient's name, gender, and age in the chest CT report image are consistent with the identity information entered by the user; and verify the timeliness of the report based on the number and type of nodules, and determine whether it meets the follow-up time requirements by calculating the time difference between the examination dates.

[0072] According to one embodiment of the present disclosure, the device 200 can pre-process the chest CT report image for denoising and contrast enhancement; use the layout analysis model to locate the key areas in the report of the pre-processed chest CT report image, the key areas including the patient information area, the nodule description area and the doctor's signature area; use the deep learning OCR model to extract the text data in the key areas, and convert the text data into structured JSON data.

[0073] According to one embodiment of the present disclosure, the device 200 can establish a standardized mapping rule library for medical terminology of pulmonary nodules, and the mapping rule library stores mapping rules including original keywords, mapping targets, weights, and applicable scenarios; based on the mapping rules, non-standard terms in the report are converted into unified medical terms, synonyms and negative descriptions are identified, and they are processed uniformly. When there are multiple versions of the description of the same nodule, the most accurate description is selected according to priority to obtain the mapping result; and the mapping result is generated into report data in a standardized JSON format.

[0074] According to one embodiment of the present disclosure, the device 200 can input the CT parameters, dynamic follow-up data and external risk factors in the report data into a pre-trained lung cancer risk probability model for multimodal data fusion analysis to calculate the patient's risk probability value for lung cancer. The CT parameters include the size of the nodule's long axis / short axis, malignant signs and their combination, the number of nodules, and the location of the nodules. The dynamic follow-up data includes the annual growth rate of the nodule volume. The external risk factors include the patient's age, gender, smoking history, and family history data. The patient's lung cancer risk is divided into different levels according to the risk probability value, and the key factors affecting the risk probability value are output and listed.

[0075] According to one embodiment of the present disclosure, the device 200 can quantify the synergistic effect between malignant signs based on the combination of malignant signs in CT parameters; introduce a time attenuation factor into the dynamic follow-up data, and dynamically adjust the weight of the historical dynamic follow-up data.

[0076] According to one embodiment of the present disclosure, the device 200 can match the nodule characteristics and risk probability values ​​with the underwriting rules step by step according to the order of priority of rejection, exclusion, surcharge, and standard body from high to low, and output a structured underwriting conclusion, which includes whether the surcharge conditions, surcharge ratio, exclusion, special agreement and push object are met; attach a risk quantification basis based on the risk probability value to the underwriting conclusion, and provide an explanation compared with the clinical guidelines; conduct closed-loop feedback on the user data collected during the underwriting process and the subsequent claims data to continuously optimize the risk assessment model; and, if there is any of the following situations: the risk probability value is within the boundary value range, the historical imaging data provided by the user is inconsistent with the current report, the output conclusions of multiple rules for the same case are contradictory, and the confidence of the OCR recognition result is lower than the preset threshold.

[0077] In an embodiment of the present disclosure, the processor 210 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 220 may be any type of memory implemented using data storage technology, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0078] Furthermore, in the embodiments of the present disclosure, the apparatus 200 may also include an input device 230, such as a keyboard or mouse, for inputting the user's basic information and chest CT report images. Furthermore, the apparatus 200 may also include an output device 240, such as a display, for displaying the structured underwriting conclusion.

[0079] In other embodiments of the present disclosure, a computer-readable storage medium storing a computer program is further provided, wherein the computer program can achieve the following when executed by a processor: Figure 1 The steps of the method are shown.

[0080] In summary, according to the embodiment of the present disclosure, the dynamic underwriting method and device for lung nodule-bearing diseased bodies based on multimodal data fusion, by using OCR technology to automatically extract information from chest CT reports and perform verification to ensure data accuracy and consistency, and standardize the extracted medical terms through a mapping rule base, it can solve the problem of inconsistent terminology that may exist between different institutions and different reports, and can enhance the adaptability and robustness of the model to various input data. The risk assessment model can comprehensively consider different types of factors, and by fusing multimodal data, it can provide a more comprehensive lung cancer risk assessment, which is more accurate than a single-factor model. According to the lung cancer risk probability value and other nodule characteristics and different underwriting rules, the risk threshold can be automatically adjusted according to different insurance product types, so that insurance companies can flexibly adjust underwriting strategies according to the needs of specific products, thereby improving customer satisfaction and trust.

[0081] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the apparatus and method according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0082] Unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular includes the plural, and vice versa. Thus, when referring to the singular, the plural of the corresponding term is generally included. Similarly, the words "include" and "comprising" are to be interpreted as inclusive rather than exclusive. Likewise, the terms "include" and "or" should be interpreted as inclusive unless such interpretation is expressly prohibited herein. Where the term "example" is used herein, particularly when it follows a group of terms, the "example" is merely exemplary and illustrative and should not be considered exclusive or comprehensive.

[0083] Further aspects and scope of adaptability become apparent from the description provided herein. It should be understood that various aspects of the present application can be implemented individually or in combination with one or more other aspects. It should also be understood that the description and specific embodiments herein are intended to be illustrative only and are not intended to limit the scope of the present application.

[0084] Several embodiments of the present disclosure have been described in detail above, but it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The scope of protection of the present disclosure is defined by the appended claims.

Claims

1. A dynamic underwriting method for lung nodule-bearing patients based on multimodal data fusion, characterized by: The method comprises: Obtain the user's basic information and chest CT report image, perform OCR recognition and verification on the chest CT report image, and output structured JSON data; performing standardization and semantic expansion on the medical terms extracted from the structured JSON data according to a mapping rule library to generate standardized report data; Inputting the CT parameters, dynamic follow-up data, and external risk factors in the report data into a lung cancer risk probability model to calculate a lung cancer risk probability value; and The lung cancer risk probability value and other nodule characteristics are matched with the underwriting rules step by step to generate a structured underwriting conclusion.

2. The dynamic underwriting method for lung nodules with disease based on multimodal data fusion according to claim 1 is characterized in that: The steps of obtaining the user's basic information and chest CT report image, performing OCR recognition and verification on the chest CT report image, and outputting structured JSON data include: Obtain the insured's name, date of birth, gender, smoking history, and family history of lung cancer among immediate family members entered by the user on the client interface; Obtain chest CT report images uploaded by users, which contain complete patient information, examination institution name, examination date, nodule description, and doctor's signature information; Check whether the user input information is complete, the data format is correct, the image file format and size meet the requirements, and whether the content of the image report is true; Perform report integrity verification, identity consistency verification, and follow-up time verification on the user's basic information and chest CT report images according to the preset verification rule library, and output the verification results; and Perform OCR recognition on verified chest CT report images, extract text data, and convert the text data into structured JSON data.

3. The dynamic underwriting method for lung nodules with disease based on multimodal data fusion according to claim 2 is characterized in that: The user's basic information and chest CT report image are verified for report integrity, identity consistency, and follow-up time based on a preset verification rule library, and the output verification results include: Check whether the chest CT report image contains the CT examination conclusion, doctor's signature, examination institution name, examination date information, and verify whether the examination date meets the validity period requirements; Compare the patient's name, gender, and age in the chest CT report image to see if they are consistent with the identity information entered by the user; and Based on the number and type of nodules, verify the timeliness of the report and determine whether the follow-up time requirements are met by calculating the time difference between the examination dates.

4. The dynamic underwriting method for lung nodules with disease based on multimodal data fusion according to claim 2 is characterized in that: The performing OCR recognition on the verified chest CT report image, extracting text data, and converting the text data into structured JSON data includes: Preprocess chest CT report images for denoising and contrast enhancement; Using a layout analysis model to locate key areas in the pre-processed chest CT report image, wherein the key areas include a patient information area, a nodule description area, and a doctor's signature area; and A deep learning OCR model is used to extract text data in the key area and convert the text data into structured JSON data.

5. The dynamic underwriting method for lung nodule-bearing patients based on multimodal data fusion according to claim 1 is characterized in that: The step of performing standardization and semantic expansion on the medical terms extracted from the structured JSON data according to the mapping rule library to generate standardized report data includes: Establishing a standardized mapping rule library for pulmonary nodule medical terminology, wherein the mapping rule library stores mapping rules including original keywords, mapping targets, weights, and applicable scenarios; Based on the mapping rules, non-standard terms in the report are converted into unified medical terms, synonyms and negative descriptions are identified and processed uniformly, and when there are multiple versions of the description of the same nodule, the most accurate description is selected according to priority to obtain the mapping result; and The mapping results are used to generate report data in a standardized JSON format.

6. The dynamic underwriting method for lung nodule-bearing patients based on multimodal data fusion according to claim 1 is characterized in that: Inputting the CT parameters, dynamic follow-up data, and external risk factors in the report data into the lung cancer risk probability model to calculate the lung cancer risk probability value includes: Inputting CT parameters, dynamic follow-up data, and external risk factors in the report data into a pre-trained lung cancer risk probability model for multimodal data fusion analysis to calculate the patient's risk probability value for lung cancer, wherein the CT parameters include the size of the nodule's long axis / short axis, malignant signs and their combination, the number of nodules, and the location of the nodules; the dynamic follow-up data includes the annual growth rate of nodule volume; and the external risk factors include the patient's age, gender, smoking history, and family history data; The patient's lung cancer risk is divided into different levels according to the risk probability value, and key factors affecting the risk probability value are output and listed.

7. The dynamic underwriting method for lung nodules with disease based on multimodal data fusion according to claim 6 is characterized in that: Inputting the CT parameters, dynamic follow-up data, and external risk factors in the report data into a pre-trained lung cancer risk probability model for multimodal data fusion analysis to calculate the patient's risk probability value for lung cancer includes: Quantify the synergistic effect between malignant features based on their combination in CT parameters; and A time decay factor is introduced into the dynamic follow-up data to dynamically adjust the weight of the historical dynamic follow-up data.

8. The dynamic underwriting method for lung nodule-bearing patients based on multimodal data fusion according to claim 1 is characterized in that: The stepwise matching of the lung cancer risk probability value and other nodule characteristics with the underwriting rules to generate a structured underwriting conclusion includes: Based on the order of priority (from high to low) for rejection, exclusion, surcharge, and standard body, the nodule characteristics and risk probability values ​​are matched against the underwriting rules step by step, and a structured underwriting conclusion is output. The underwriting conclusion includes whether the surcharge conditions are met, the surcharge ratio, exclusions, special agreements, and the recipients of the notification. Include risk quantification based on risk probability values ​​in the underwriting conclusion and provide explanations comparing with clinical guidelines; Close the loop by integrating user data collected during the underwriting process with subsequent claims data to continuously optimize the risk assessment model; and Manual review will be triggered if any of the following situations occur: the risk probability value is within the boundary value range, the historical image data provided by the user is inconsistent with the current report, multiple rules have contradictory output conclusions for the same case, or the confidence level of the OCR recognition result is lower than the preset threshold.

9. A dynamic underwriting device for lung nodule diseased body based on multimodal data fusion, characterized by: The device comprises: at least one processor; and at least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the device is caused to perform the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: The computer program implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Method and device for determining insurance data

    CN111275558A

  • Iconography examination report structuring method and device

    CN111814478A

  • Intelligent underwriting method and system based on pulmonary nodule risk quantification

    CN114638715A

  • General physical examination report OCR (Optical Character Recognition) method and data processing system

    CN115273084A

  • Lung cancer high-risk group identification device

    CN115966305A