Cognitive calculation correction method and system for screening questionnaires for high-risk groups of lung cancer
By analyzing questionnaire semantics using the BERT-QA model and clinical terminology knowledge graph, and combining topological entropy weights and fuzzy logic, the weights of risk factors are dynamically adjusted and regional pollution data is integrated. This addresses the shortcomings of existing screening technologies for high-risk groups of lung cancer, enabling more accurate and personalized risk assessment and supporting precision medicine decision-making.
Patent Information
- Application Number
- CN202511568016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-03-17
AI Technical Summary
Current lung cancer high-risk screening technologies lack a deep understanding of questionnaire content, cannot adapt to the risk characteristics of different regions and populations, ignore the complex nonlinear relationships between risk factors, and fail to effectively integrate environmental factors, resulting in insufficient accuracy and personalization of assessment results.
The questionnaire semantics were analyzed using a BERT-QA model and a clinical terminology knowledge graph. Combined with topological entropy weights and fuzzy logic, the weights of risk factors were dynamically adjusted, and regional pollution data were integrated to generate a personalized lung cancer risk assessment.
It improves the accuracy of questionnaire comprehension, enables dynamic adjustment of risk factor weights, captures higher-order nonlinear relationships, effectively integrates environmental factors, provides personalized risk assessment, supports precision medicine decision-making, and enhances the accuracy and coverage of screening.
Smart Images

Figure CN121687344A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, specifically to a cognitive calculation correction method and system for screening questionnaires for high-risk groups of lung cancer, and in particular, an intelligent parsing and risk assessment technology for screening questionnaires for high-risk groups of lung cancer based on topological entropy weight fuzzy cognition. Background Technology
[0002] Lung cancer is one of the leading causes of cancer-related deaths worldwide, with squamous cell carcinoma accounting for approximately 25-30% of all lung cancer cases. Early screening and intervention are crucial for improving the survival rate of lung cancer patients. Currently, lung cancer screening for high-risk groups primarily relies on questionnaires to collect basic information about participants, including smoking history and occupational exposure risk factors, and then assesses lung cancer risk through manual or simple calculation methods.
[0003] However, existing screening technologies for high-risk groups of lung cancer have the following shortcomings: First, traditional questionnaire analysis methods lack a deep understanding of natural language, making it difficult to accurately extract key risk factors from questionnaires. In particular, they lack understanding of ambiguous expressions and non-standard terminology, which limits the accuracy of risk assessment.
[0004] Secondly, existing risk assessment models typically employ fixed-weight linear models, which cannot adapt to the differences in risk characteristics among different regions and populations, lack dynamic adjustment capabilities, and result in insufficient universality of assessment results.
[0005] Third, traditional models fail to fully consider the complex nonlinear relationships between risk factors, especially ignoring higher-order interactions, which limits the accuracy of risk assessment.
[0006] Fourth, existing methods lack effective integration of regional environmental factors, while environmental pollution is an important influencing factor in lung cancer incidence. Ignoring this factor can lead to biased risk assessment.
[0007] Finally, existing risk assessment results are usually simple high / low risk classifications, lacking personalized risk probability assessments and confidence intervals, making it difficult to support precision medicine decisions.
[0008] Therefore, there is an urgent need to develop an intelligent cognitive computation correction method and system for screening questionnaires for high-risk groups of lung cancer, in order to improve the accuracy and personalization of screening. Summary of the Invention
[0009] The purpose of this invention is to provide a cognitive computation correction method and system for screening questionnaires for high-risk groups of lung cancer, aiming to solve the problems existing in the prior art. By integrating advanced technologies such as deep learning, topology and fuzzy logic, it achieves intelligent parsing and accurate risk assessment of screening questionnaires for high-risk groups of lung cancer.
[0010] This invention proposes a cognitive calculation correction method for lung cancer high-risk population screening questionnaires, including: Obtain the text data of the subject's questionnaire; A BERT-QA model was constructed based on a clinical terminology knowledge graph to analyze the semantics of the subject's questionnaire text data and extract structured risk factor data. The structured risk factor data is mapped to a topological space to construct a risk topological representation; Based on the principle of information entropy, the dynamic weight coefficients of each risk factor in the risk topology representation are calculated to generate an entropy-weighted optimized risk assessment value. The risk assessment value optimized by the entropy weight is calibrated by combining regional pollution map data; Based on fuzzy logic and the calibrated risk assessment value, the high-risk probability of lung squamous cell carcinoma for the subject was determined.
[0011] Preferably, the step of obtaining the subject questionnaire text data specifically includes: Content words in the subject questionnaire text are labeled "BIO" and function words are labeled "B". The annotated text is converted into XML format for semantic tagging. The city epidemiology dictionary is trained based on the word2vec algorithm, and a questionnaire vector representation containing regional characteristics is generated by combining the epidemiological characteristics of the city where the subject is located.
[0012] Preferably, the steps for constructing the BERT-QA model based on the clinical terminology knowledge graph specifically include: Construct an entity set containing diseases, symptoms, risk factors, treatment methods, and examination methods, as well as a relation set containing causes, correlations, aggravations, relief, and detections, forming an "entity-relationship-entity" triplet knowledge representation; The triplet knowledge representation is converted into an embedded representation, which is then fused with the BERT model output through an attention mechanism to generate a semantically enhanced representation. Based on the semantically enhanced representation, the smoking age assessment value and occupational exposure assessment value are extracted from the subject's questionnaire.
[0013] Preferably, the step of mapping structured risk factor data to a topological space specifically includes: Construct a multidimensional space with risk factors as coordinate axes, and define an open set family containing risk levels of {low, low-medium, medium, medium-high, high}. Design a topological distance function based on the characteristics of risk factors to ensure the continuity of risk assessment; A three-level topology mapping process is constructed, including mapping from raw data to standardized topology space, mapping from standard space to fuzzy membership space, and mapping from membership space to risk assessment space.
[0014] Preferably, the step of calculating the dynamic weight coefficients of each risk factor based on the information entropy principle specifically includes: The risk factors are discretized into multiple intervals, and the probability density distribution of each interval is calculated. The information entropy value of each risk factor is calculated based on the probability density distribution. Weighting coefficients are generated based on the inverse relationship between information entropy and weighting coefficients; the smaller the information entropy, the larger the weighting coefficient. The original risk assessment value is weighted using the weighting coefficients to obtain the entropy-weighted optimized risk assessment value.
[0015] Preferably, the steps based on fuzzy logic specifically include: Define fuzzy sets and membership functions based on topological continuity for risk factors; Establish a three-level fuzzy rule base, including first-level rules for single risk factors, second-level rules for multi-factor interactions, and third-level rules that consider the impact of environmental factors; The fuzzy inference process includes converting precise values into fuzzy membership degrees, triggering corresponding fuzzy rules based on input, aggregating multi-rule results using the centroid method, and converting fuzzy results into precise risk values.
[0016] Preferably, the step of calibrating by combining regional pollution map data specifically includes: Collect regional environmental pollution data including fine particulate matter (PM2.5), sulfur dioxide, and nitrogen oxides; Establish association models between pollutants and lung cancer risk, including direct impact models, synergistic effect models, and long-term cumulative models; The risk assessment value was calibrated based on the pollution level in the region where the subjects were located, using a combination of linear adjustment, threshold adjustment, and individual difference adjustment strategies.
[0017] Preferably, the step of determining the high-risk probability of a subject's squamous cell carcinoma of the lung specifically includes: The evaluation criteria of multiple international lung cancer screening guidelines were integrated, and weights were assigned based on the level of evidence and applicability. The risk threshold was adjusted based on the subject's gender, age, and body mass index. Generate 0-100% risk probability values for squamous cell carcinoma of the lung and their confidence intervals; When the risk probability value of lung squamous cell carcinoma is higher than a certain threshold, the subject is determined to be a high-risk group for lung cancer.
[0018] Preferably, the following steps are also included: A multi-level feedback optimization mechanism is set up, including parameter self-optimization within modules, result verification between adjacent modules, and system-level global optimization; Validate intermediate results based on the internal validation set and estimate the uncertainty of the results; Receive expert feedback and dynamically adjust system parameters to improve the accuracy of risk assessment.
[0019] A cognitive computation correction system for lung cancer high-risk population screening questionnaires includes: The questionnaire input and preprocessing module is used to acquire the questionnaire text data of the respondents and perform standardization processing; The knowledge graph-enhanced BERT-QA module is used to analyze questionnaire semantics based on clinical terminology knowledge graphs and extract structured risk factor data. The topological entropy weight calculation module is used to construct a risk topological representation and calculate the dynamic weight coefficients of risk factors. The fuzzy cognitive reasoning module is used to perform fuzzy logic reasoning based on topological structures. The regional pollution fusion module is used to integrate regional environmental pollution data and calibrate risk assessment values. The individualized risk assessment module is used to determine the high-risk probability of lung squamous cell carcinoma in subjects; The modules interact with each other through standardized data interfaces to form a complete risk assessment pipeline.
[0020] The present invention has the following beneficial effects: 1. Improve the accuracy of questionnaire comprehension: By combining clinical terminology knowledge graphs and BERT-QA models, the semantic understanding of questionnaire texts is significantly improved, enabling accurate extraction of key risk factors and effective handling of ambiguous expressions and non-standard terminology.
[0021] 2. Achieve dynamic weight calculation of risk factors: Based on the topological entropy weight method, the weights of each factor can be automatically adjusted according to the risk characteristics of different regions and populations, making the assessment results more adaptable and accurate.
[0022] 3. Capturing complex relationships between risk factors: Through topological characterization and fuzzy logic reasoning, it is possible to discover and utilize high-order nonlinear relationships between risk factors, thereby improving the accuracy of risk assessment.
[0023] 4. Effectively integrate environmental factors: Combine regional pollution map data to incorporate environmental factors into the risk assessment system, making the assessment results more comprehensive and more in line with the actual situation.
[0024] 5. Provides personalized risk assessment: Generates individualized lung cancer risk probabilities and confidence intervals, rather than simple high / low risk classifications, providing more accurate support for clinical decision-making.
[0025] 6. Optimize the allocation of medical resources: Through precise screening, reduce unnecessary high-cost examinations, increase the screening coverage of high-risk groups, and reduce the total screening cost.
[0026] 7. Improve early detection rate: Accurate identification of high-risk groups helps improve the early detection rate of lung cancer, thereby reducing treatment costs and improving patient survival rates. Attached Figure Description
[0027] Figure 1 This is a flowchart of the cognitive calculation correction method for the lung cancer high-risk population screening questionnaire of the present invention; Figure 2 This is a schematic diagram of the questionnaire input and preprocessing module in this invention; Figure 3 This is an architecture diagram of the knowledge graph-enhanced BERT-QA module in this invention; Figure 4 This is a flowchart illustrating the workflow of the topological entropy weight calculation module in this invention. Figure 5 This is a schematic diagram of the fuzzy cognitive reasoning module in this invention; Figure 6 This is a data processing flowchart of the regional pollution fusion module in this invention; Figure 7 This is a diagram illustrating the overall architecture of the cognitive computation correction system for lung cancer high-risk population screening questionnaires of the present invention. Detailed Implementation
[0028] Please refer to the attached document. Figure 1-7 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the following embodiments are only used to illustrate the present invention and are not intended to limit the present invention.
[0029] Reference Figure 1 This invention provides a cognitive computation correction method for screening questionnaires for high-risk groups of lung cancer, comprising the following steps: acquiring the subject's questionnaire text data; constructing a BERT-QA model based on a clinical terminology knowledge graph, parsing the semantics of the subject's questionnaire text data, and extracting structured risk factor data; mapping the structured risk factor data to a topological space to construct a risk topological representation; calculating the dynamic weight coefficients of each risk factor in the risk topological representation based on the principle of information entropy, and generating an entropy-weighted optimized risk assessment value; calibrating the entropy-weighted optimized risk assessment value in conjunction with regional pollution map data; and determining the high-risk probability of the subject's squamous cell carcinoma of the lung based on fuzzy logic and the calibrated risk assessment value.
[0030] In a preferred embodiment of the present invention, the method further includes setting a multi-level feedback optimization mechanism, including parameter self-optimization within the module, result verification between adjacent modules, and system-level global optimization, so as to further improve the accuracy of risk assessment.
[0031] This invention also provides a corresponding cognitive computation correction system for a lung cancer high-risk population screening questionnaire. This system includes a questionnaire input and preprocessing module 1, a knowledge graph-enhanced BERT-QA module 2, a topological entropy weight calculation module 3, a fuzzy cognitive reasoning module 4, a regional pollution fusion module 5, and an individualized risk assessment module 6. These modules interact through standardized data interfaces, forming a complete risk assessment pipeline.
[0032] The various technical solutions of the present invention will be described in detail below with reference to specific embodiments.
[0033] Reference Figure 2 In one embodiment of the present invention, the questionnaire input and preprocessing module 1 is responsible for acquiring the subject's questionnaire text data and performing standardization processing. This step specifically includes: labeling content words in the subject's questionnaire text with "BIO" tags and function words with "B" tags; converting the labeled text into XML format and performing semantic annotation; training an urban epidemiology dictionary based on the word2vec algorithm, and generating a questionnaire vector representation containing regional characteristics by combining the epidemiological characteristics of the subject's city.
[0034] In practical applications, the "BIO" annotation system is a commonly used sequence annotation method, where "B" represents the beginning of an entity, "I" represents the inside of an entity, and "O" represents the outside of an entity. For example, for the text "The patient smokes 20 cigarettes a day for 15 years" in a questionnaire, the annotation result might be "Patient / B daily / B smokes / B20 / B cigarettes / I, / O has / B continued / B15 / B years / I".
[0035] Next, the annotated text will be converted into XML format for easier subsequent processing.
[0036] For training the urban epidemiological dictionary, this invention employs the word2vec algorithm, a neural network-based word embedding method that maps words to a high-dimensional vector space, making semantically similar words closer together in that space. In this embodiment, the training corpus includes epidemiological data, environmental data, and lung cancer incidence data from various cities.
[0037] Preferably, the training parameters for the word2vec model are set as follows: vector dimension of 300, window size of 5, minimum word frequency of 2, and number of training iterations of 100. In this way, each city can be represented as a 300-dimensional vector containing its epidemiological characteristics. For example, for cities with severe air pollution, the dimension related to pollution will have a higher value in its vector representation.
[0038] Through the above preprocessing steps, the original questionnaire text is converted into a structured data representation, laying the foundation for subsequent semantic analysis and risk assessment.
[0039] Reference Figure 3 In one embodiment of the present invention, the knowledge graph-enhanced BERT-QA module 2 is responsible for parsing the semantics of the questionnaire based on the clinical terminology knowledge graph and extracting structured risk factor data. This step specifically includes: constructing an entity set containing diseases, symptoms, risk factors, treatment methods, and examination methods, and a relation set containing causes, correlations, aggravations, relief, and detections, forming an "entity-relationship-entity" triplet knowledge representation; converting the triplet knowledge representation into an embedded representation, and fusing it with the BERT model output through an attention mechanism to generate a semantically enhanced representation; and based on the semantically enhanced representation, extracting smoking age assessment values and occupational exposure assessment values from the subject's questionnaire.
[0040] First, a knowledge graph of clinical terminology in the field of lung cancer is constructed. The knowledge graph adopts a "entity-relationship-entity" triple structure, such as "smoking-cause-lung cancer" and "asbestos-aggravation-lung cancer risk." In this embodiment, the entity set includes diseases (such as lung cancer, squamous cell carcinoma of the lung, adenocarcinoma of the lung, etc.), symptoms (such as cough, hemoptysis, chest pain, etc.), risk factors (such as smoking, occupational exposure, family history, etc.), treatment methods (such as surgery, radiotherapy, chemotherapy, etc.), and diagnostic methods (such as CT scans, biopsies, tumor markers, etc.). The relationship set includes causing, related, aggravating, relieving, and detecting.
[0041] To combine knowledge graphs with the BERT model, this invention employs knowledge embedding technology.
[0042] Specifically, the TransE model is used to embed entities and relations from a knowledge graph into a low-dimensional vector space. The basic idea of the TransE model is that for a valid triple... The relationship between entity embedding vectors should satisfy: , Where: h is the embedding vector of the head entity with dimension d; r is the embedding vector of the relation with dimension d; t is the embedding vector of the tail entity with dimension d; d is the dimension of the embedding space, which is set to 100 in this embodiment.
[0043] The model is trained by minimizing the following objective function: , Where: L is the loss function; S is the set of positive triples, which includes triples existing in the knowledge graph; The set of negative triples is generated by randomly replacing the head or tail entity in the positive triples. is a marginal parameter used to control the distance between positive and negative examples, set to 1.0; d is the vector distance function, which can use either the L1 norm or the L2 norm. This embodiment uses the L2 norm. This indicates taking the positive value, when hour, .
[0044] In this embodiment, the knowledge graph embedding dimension is set to 100, the number of training rounds is 1000, the learning rate is 0.01, and the batch size is 128. After training, each entity and relation corresponds to a 100-dimensional vector representation.
[0045] Next, the knowledge graph embedding is fused with the BERT model. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model with powerful semantic understanding capabilities. This invention uses a BERT-Base model fine-tuned with a medical corpus. This model contains a 12-layer Transformer structure with a hidden layer dimension of 768 and 12 attention heads.
[0046] Knowledge fusion is achieved through an attention mechanism, and the specific calculation formula is as follows: , Where: A is the attention output matrix, with dimensions... Q is the query matrix, generated by BERT output, with dimensions of [missing information]. K is the key matrix, generated by embedding from the knowledge graph, with dimensions of ; V is a value matrix generated by embedding from a knowledge graph, with dimensions of [missing information]. n is the sequence length; m is the number of triples in the knowledge graph. Let be the dimension of the key vector, set to 100; Let 100 be the dimension of the value vector; The normalization function ensures that the sum of the weights of each row is 1; This is a scaling factor used to stabilize the gradient.
[0047] In this way, BERT's output is enhanced with knowledge graph information, forming a semantically enhanced representation.
[0048] Finally, based on semantically enhanced representation, a linear classifier was used to extract key risk factors from the questionnaire, including smoking age assessment (RA) and occupational exposure assessment (PA). In practical applications, the smoking age assessment is calculated as follows: RA = Years of smoking × Weighting of smoking age Where: RA is the smoking age assessment value, which is dimensionless; smoking years is the number of years the subject has continuously smoked, in years; smoking age weight is an adjustment coefficient, which is dimensionless and reflects the degree of influence of different tobacco types and smoking amounts.
[0049] The formula for calculating the smoking age weight is: , Wherein: the unit for standard cigarette tar content is milligrams (mg); the unit for standard cigarette smoke carbon monoxide emission is milligrams (mg); 30 and 40 are normalization coefficients, in milligrams (mg), used to standardize indicators with different units of measurement to a similar numerical range. For ordinary cigarettes, the tar content is approximately 10-15 mg, and the carbon monoxide emission is approximately 10-20 mg, therefore the weight of smoking age is usually between 0.6 and 1.0.
[0050] Occupational Exposure Assessment (PA) values range from 0 to 100, and are dimensionless, based on a comprehensive assessment of the type, intensity, and duration of occupational exposure. For example, a subject who has long been engaged in asbestos-related work may have a PA value between 80 and 100; while a subject who is occasionally exposed to a weak carcinogen may have a PA value between 10 and 30.
[0051] Reference Figure 4 In one embodiment of the present invention, the topological entropy weight calculation module 3 is responsible for constructing a risk topological representation and calculating the dynamic weight coefficients of risk factors. This step specifically includes: constructing a multi-dimensional space with risk factors as coordinate axes, defining an open set family containing risk levels {low, low-medium, medium, medium-high, high}; designing a topological distance function based on the characteristics of risk factors to ensure the continuity of risk assessment; and constructing a three-level topological mapping process, including mapping from raw data to a standardized topological space, mapping from the standard space to a fuzzy membership space, and mapping from the membership space to the risk assessment space.
[0052] In the process of constructing the topological space, each risk factor is first treated as a coordinate axis, forming a multidimensional risk space. For example, for the case containing two main risk factors, smoking history (RA) and occupational exposure (PA), a two-dimensional topological space is constructed. Within this space, a family of open sets is defined to represent different risk levels. For each risk level, the corresponding open set is defined as follows: , , , , , in: , , , and , respectively, represent the open sets of low, low-medium, medium, medium-high, and high risk levels; x is the smoking age assessment value (RA), dimensionless; y is the occupational exposure assessment value (PA), dimensionless; , , , , is a threshold parameter for smoking age, and is dimensionless; , , , The threshold parameter for occupational exposure factors is dimensionless.
[0053] In this example, based on clinical experience and statistical analysis, the following settings were established. , , , (dimensionless) , , , (Dimensionless). These threshold parameters reflect the dividing points of different risk levels and can be adjusted according to the specific characteristics of the population.
[0054] To ensure the continuity of risk assessment, this invention designs a topological distance function suitable for risk assessment. In two-dimensional space, a weighted Euclidean distance is used: , Where: d is the distance between the two points, which is dimensionless; and Two points in space represent two different risk states; and The weights for the coordinate axes are dimensionless and reflect the relative importance of different risk factors. In the initial stage, they can be set... The value will be dynamically adjusted subsequently through entropy weight calculation.
[0055] Next, a three-level topology mapping process is constructed. The first-level mapping maps the raw data to a normalized topology space using a min-max normalization method: , , in: and The coordinates are normalized, dimensionless, and range from [0,1]; x and y are the original smoking age assessment value and occupational exposure assessment value. and The minimum and maximum values for the smoking age assessment are set to 0 and 60 (dimensionless), respectively. and The minimum and maximum values for occupational exposure assessment are set to 0 and 100 (dimensionless), respectively.
[0056] The second-level mapping maps the standardized space to the fuzzy membership space, using an appropriate membership function (see the subsequent fuzzy logic section for details). The third-level mapping maps the membership space to the risk assessment space, generating the final risk assessment value.
[0057] After the topological space is constructed, this invention calculates the dynamic weight coefficients of each risk factor based on the principle of information entropy. First, the risk factors are discretized into multiple intervals, and the probability density distribution of each interval is calculated. For example, for the smoking age factor, it can be divided into five intervals: [0,10), [10,20), [20,30), [30,40), and [40,60]. Then, the frequency of the sample in each interval is calculated to obtain the probability distribution. , ,..., ,in Let represent the probability of the i-th interval, and satisfy . .
[0058] Based on the probability distribution, calculate the information entropy of smoking age as a factor: , in: The information entropy of smoking age is dimensionless and its value range is [value range missing]. (For the case of 5 intervals); Let be the probability of the i-th interval; The logarithmic function with base 2. The entropy is maximized when all intervals have equal probabilities, indicating the least amount of information; the entropy is 0 when the probability of a certain interval is 1 and the others are 0, indicating the maximum amount of information.
[0059] Similarly, the information entropy of occupational exposure factors can be calculated. According to the inverse principle of information entropy, the lower the information entropy, the stronger the discriminative ability of the factor, and therefore it should be assigned a higher weight. Therefore, the weight calculation formula is: , , in: and These are the weighting coefficients for smoking duration and occupational exposure, respectively. They are dimensionless, range from [0,1], and satisfy the following conditions: ; denominator This is a normalization factor, ensuring the sum of the weights is 1. When When the two factors have equal weight, then the two factors have equal weight; when At that time, the duration of smoking had a relatively large weight.
[0060] Finally, the original risk assessment value is weighted using the above weights to obtain the entropy-weighted optimized risk assessment value: , , in: and The risk assessment value after entropy weight optimization is dimensionless. and This is the original risk assessment value, dimensionless.
[0061] In this way, the present invention realizes dynamic weight calculation of risk factors, making the assessment results more adaptable and accurate.
[0062] Reference Figure 5 In one embodiment of the present invention, the fuzzy cognitive reasoning module 4 is responsible for performing fuzzy logic reasoning based on topological structure. This step specifically includes: defining fuzzy sets and membership functions based on topological continuity for risk factors; establishing a three-level fuzzy rule base, including first-level rules for single risk factors, second-level rules for multi-factor interactions, and third-level rules considering the influence of environmental factors; and executing the fuzzy reasoning process, including converting precise values into fuzzy membership degrees, triggering corresponding fuzzy rules based on input, aggregating multi-rule results using the centroid method, and converting the fuzzy results into precise risk values.
[0063] First, fuzzy sets and membership functions are defined for the risk factors. In this embodiment, the risk levels are divided into five levels of fuzzy sets: {low, low-medium, medium, high-medium, high}. For the smoking age factor (RA), its membership function can be expressed as: , , , , , in: , , , and RA represents the membership degree of the smoking age factor RA to low, low-medium, medium, medium-high, and high risk levels, respectively. It is dimensionless and ranges from [0,1]. RA is the smoking age assessment value, which is dimensionless.
[0064] Similarly, membership functions can be defined for occupational exposure factors (PAs). These membership functions are designed based on topological continuity, ensuring a smooth transition in risk assessment.
[0065] Next, a three-level fuzzy rule base will be established. Level 1 rules target a single risk factor, for example: If RA is low, then the risk of lung cancer is low; If RA is moderate, then the risk of lung cancer is moderate; If PA is high, then the risk of lung cancer is high; Second-level rules consider the interaction of multiple factors, such as: If both RA and PA are moderate, then the risk of lung cancer is moderate to high. If RA is high and PA is low, the risk of lung cancer is moderate to high. If RA is low and PA is high, the risk of lung cancer is moderate to high. If both RA and PA are high, then the risk of lung cancer is high. Level 3 rules further consider the impact of environmental factors, such as: If RA is moderate, PA is moderate, and pollution level is high, then the risk of lung cancer is high. If RA is low, PA is low, and pollution levels are low, then the risk of lung cancer is low. In the process of fuzzy inference, the precise risk factor values are first converted into fuzzy membership degrees.
[0066] For example, for a smoking history RA = 25 years, its membership degree is: , , , , .
[0067] Then, the corresponding fuzzy rules are triggered based on the input membership degrees. For each rule, its trigger strength is calculated. For example, for the rule "If RA is moderate and PA is moderate, then the risk of lung cancer is moderate to high", assume... , Then the trigger strength of this rule is The min function returns the minimum value among the parameters and is used to implement the AND operation in fuzzy logic.
[0068] Next, the results of multiple rules are aggregated using the centroid method. First, for each point in the output fuzzy set, the maximum membership degree assigned to that point by all triggering rules is taken, forming a truncated fuzzy set. Then, the centroid of this truncated set is calculated as the defuzzification result. The centroid calculation formula is: , Where: Risk is the final lung cancer risk value, which is dimensionless and usually between 0 and 100; n is the number of discretization points; Let i be the membership value of the i-th point; Let be the risk value corresponding to the i-th point.
[0069] Assuming the final lung cancer risk value is Risk, which typically ranges from 0 to 100, reflecting the subject's relative risk of developing lung cancer. In this example, a risk value greater than 70 is considered high risk, a risk value between 30 and 70 is considered intermediate risk, and a risk value less than 30 is considered low risk.
[0070] Reference Figure 6 In one embodiment of the present invention, the regional pollution fusion module 5 is responsible for integrating regional environmental pollution data and calibrating risk assessment values. This step specifically includes: collecting regional environmental pollution data containing fine particulate matter (PM2.5), sulfur dioxide, and nitrogen oxides; establishing a correlation model between pollutants and lung cancer risk, including a direct impact model, a synergistic effect model, and a long-term cumulative model; and calibrating the risk assessment values based on the pollution level of the subject's region, combined with linear adjustment, threshold adjustment, and individual difference adjustment strategies.
[0071] First, regional environmental pollution data is collected. This invention focuses on three types of air pollutants: fine particulate matter (PM2.5), sulfur dioxide (SO2), and nitrogen oxides (NOx), which have been confirmed by multiple studies to be significantly associated with lung cancer risk. Pollution data can be obtained from national environmental monitoring stations or publicly available environmental databases and organized according to annual averages, seasonal variations, and regional distribution.
[0072] In this embodiment, PM2.5 concentrations are divided into five levels: low (<35 μg / m³), low-medium (35-50 μg / m³), medium (50-75 μg / m³), medium-high (75-100 μg / m³), and high (>100 μg / m³). Similarly, SO2 and NOx are also classified into levels.
[0073] Next, a model linking pollutants and lung cancer risk is established. The direct impact model considers the carcinogenic effects of pollutants, and based on epidemiological studies, a functional relationship is established between pollution concentration and increased lung cancer risk. For example, studies show that for every 10 μg / m³ increase in PM2.5 concentration, the risk of lung cancer increases by approximately 8-14%. Based on this data, the following relationship can be established: , in: The risk value considering the impact of PM2.5 is dimensionless. The baseline risk value is dimensionless and does not consider pollution factors; PM2.5 is the local PM2.5 concentration, in units of... ; The value represents the reference concentration level, indicating the baseline pollution level; 10 represents the unit increment, with units of... 0.01 is the risk increase coefficient, representing the percentage increase in risk per unit area. The risk increases by 1%, dimensionless.
[0074] Synergistic models consider the interaction between pollutants and other risk factors, such as smoking. Studies have shown that smokers in heavily polluted areas have a more significant increased risk of lung cancer. This synergistic effect can be expressed by the following formula: , in: The risk value after considering synergistic effects is dimensionless; RA is the smoking age assessment value, dimensionless; 0.005 is the synergistic effect coefficient, representing a 0.5% increase in pollution impact per unit smoking age, dimensionless. This formula shows that the longer the smoking age, the more significant the increase in risk due to pollution.
[0075] The long-term cumulative model considers the cumulative effect of long-term low-dose exposure. This model assumes that the effect of pollutants on lung cancer has a certain lag and cumulative nature, which can be expressed by the following formula: , in: The risk value is dimensionless and takes into account the long-term cumulative effect; n is the number of years considered, usually 5-10 years. The PM2.5 concentration in year i is given by [value]. ; It is time-weighted, reflecting the degree of pollution impact at different times, dimensionless, and satisfies... In this embodiment, an exponentially decaying weight is used: ,in It is a decay parameter, dimensionless, representing the weight of the most recent year. This represents the relative weight decay in year i.
[0076] Finally, the risk assessment values were calibrated based on the pollution levels in the participants' areas and the aforementioned model. The calibration process employed multiple strategies: 1. Linear adjustment: The risk value is adjusted linearly based on the pollutant concentration, which is suitable for areas with light pollution.
[0077] 2. Threshold adjustment: Set a pollutant concentration threshold. When the concentration exceeds the threshold, a non-linear adjustment is adopted, which is suitable for heavily polluted areas.
[0078] 3. Individual Difference Adjustment: Considering the differences in individual sensitivity to pollutants, the degree of influence is adjusted according to factors such as age and gender.
[0079] In this embodiment, taking into account the above strategies, the final risk calibration formula is: , in: The calibrated risk value is dimensionless. The pollution impact factor is dimensionless and is calculated by combining three models: direct impact, synergistic effect, and long-term accumulation. The individual variability factor is dimensionless and reflects an individual's sensitivity to pollutants. It is determined based on the subject's age, sex, and underlying health condition, and is typically between 0.8 and 1.2.
[0080] In one embodiment of the present invention, the individualized risk assessment module 6 is responsible for determining the high-risk probability of a subject for squamous cell carcinoma of the lung. This step specifically includes: integrating assessment criteria from multiple international lung cancer screening guidelines and assigning weights based on the level of evidence and applicability; adjusting the risk threshold according to the subject's gender, age, and body mass index; generating a 0-100% risk probability value for squamous cell carcinoma of the lung and its confidence interval; and determining the subject as a high-risk group for lung cancer when the risk probability value for squamous cell carcinoma of the lung is higher than a specific threshold.
[0081] First, we need to integrate the evaluation criteria of international lung cancer screening guidelines. Currently, the main international lung cancer screening guidelines include the National Comprehensive Cancer Network (NCCN) guidelines, the American Thoracic Society (ACCP) guidelines, and the US Preventive Services Task Force (USPSTF) guidelines. These guidelines differ to some extent in terms of the definition of high-risk groups, the age at which screening should begin, and the frequency of screening.
[0082] This invention assigns weights to guidelines based on their level of evidence and applicability. The level of evidence reflects the strength of the scientific basis for the guideline's recommendations, while applicability reflects the guideline's suitability for the local population. For example, a guideline with a level of evidence of 1A (strong recommendation based on high-quality randomized controlled trials) and high applicability can be assigned a weight of 0.5; a guideline with a level of evidence of 2B (weak recommendation based on lower-quality studies) and moderate applicability can be assigned a weight of 0.2.
[0083] In this embodiment, the following weighting is used: NCCN guidelines 0.4, ACCP guidelines 0.3, and USPSTF guidelines 0.3. Then, risk factors commonly recognized by each guideline are extracted to form a comprehensive evaluation standard.
[0084] Next, the risk thresholds were adjusted based on the participants' gender, age, and body mass index (BMI). Studies have shown that men, the elderly, and those with lower BMIs have a relatively higher risk of lung cancer. Therefore, the high-risk thresholds should be appropriately lowered for these groups. The specific adjustment methods are as follows: 1. Gender adjustment: The threshold for males is the standard threshold, and the threshold for females is 1.1 times the standard threshold.
[0085] 2. Age adjustment: The threshold for those under 50 years old is 1.2 times the standard threshold, the threshold for those aged 50-65 years old is the standard threshold, and the threshold for those over 65 years old is 0.9 times the standard threshold.
[0086] 3. BMI Adjustment: The threshold for BMI < 18.5 is 0.9 times the standard threshold, the threshold for 18.5 ≤ BMI < 25 is the standard threshold, and the threshold for BMI ≥ 25 is 1.1 times the standard threshold.
[0087] Based on the above adjustments, the final formula for calculating the individualized threshold is: , in: Individualized risk thresholds, dimensionless, in percentage form; The standard risk threshold is set at 65%, representing the baseline threshold without individual adjustments. The sex adjustment factor is dimensionless, with 1.0 for males and 1.1 for females. The age adjustment factor is dimensionless, with 1.2 for those under 50, 1.0 for those between 50 and 65, and 0.9 for those over 65. This is a BMI adjustment factor, dimensionless, with a value of 0.9 for BMI < 18.5, 1.0 for BMI ≤ BMI < 25, and 1.1 for BMI ≥ 25.
[0088] Based on the calibrated risk assessment value and individualized threshold, a final risk probability value for squamous cell carcinoma of the lung is generated. The risk probability value is calculated using a logistic regression model: , Where: P(lung squamous cell carcinoma) is the risk probability of lung squamous cell carcinoma, dimensionless, with a value range of [0,1]; e is the base of the natural logarithm, approximately equal to 2.71828; z is the linear combination term of the Logistic model, dimensionless, calculated using the following formula: ; The model parameters are derived through training with large-scale clinical data; RA calibration and PA calibration are calibrated smoking age assessment values and occupational exposure assessment values, which are dimensionless; Age is age in years; Sex is gender, with 1 for males and 0 for females; BMI is body mass index in kg / m². In this embodiment, , , , , , .
[0089] In addition to risk probability values, this invention also provides confidence interval estimates, reflecting the uncertainty of risk assessment. The confidence interval is calculated based on the Delta method or the Bootstrap method, typically providing a 95% confidence interval. The formula for the Delta method is: , in: The 95% confidence interval is given; P(lung squamous cell carcinoma) is the risk probability of lung squamous cell carcinoma; 1.96 is the 97.5th percentile of the standard normal distribution; SE(P) is the standard error of the risk probability, calculated using the following formula: , in For parameters variance For the corresponding eigenvalues (for , ).
[0090] Finally, the risk probability value is compared with an individualized threshold. When the risk probability value is higher than the threshold, the subject is identified as a high-risk group for lung cancer and further screening (such as low-dose CT) is recommended. For example, for a 60-year-old male subject with a BMI of 22, the individualized threshold is 65% × 1.0 × 1.0 × 1.0 = 65%. If his risk probability value is 72% and the 95% confidence interval is [65%, 79%], he is identified as a high-risk individual.
[0091] Reference Figure 7 This invention also provides a cognitive computation correction system for a lung cancer high-risk population screening questionnaire, used to implement the above method. The system includes a questionnaire input and preprocessing module 1, a knowledge graph-enhanced BERT-QA module 2, a topological entropy weight calculation module 3, a fuzzy cognitive reasoning module 4, a regional pollution fusion module 5, and an individualized risk assessment module 6.
[0092] These modules interact through standardized data interfaces, forming a complete risk assessment pipeline. Specifically, the output of the questionnaire input and preprocessing module 1 is standardized questionnaire text data; the output of the knowledge graph-enhanced BERT-QA module 2 is structured risk factor data; the output of the topological entropy weight calculation module 3 is the entropy weight-optimized risk assessment value; the output of the fuzzy cognitive reasoning module 4 is the fuzzy reasoning result; the output of the regional pollution fusion module 5 is the calibrated risk assessment value; and the output of the individualized risk assessment module 6 is the final risk probability of squamous cell carcinoma of the lung.
[0093] In actual deployment, the system can adopt a modular, microservice architecture, including a front-end module (questionnaire input interface, result display interface), semantic parsing service (microservice for deploying BERT-QA model), risk calculation service (computation engine including topological entropy weight algorithm), data storage module (knowledge graph library, pollution map database, patient archive library) and API interface layer (providing standardized interfaces for easy integration with hospital HIS system).
[0094] The system's implementation requirements are reasonable and suitable for deployment in medical institutions. In terms of computing resources, standard servers are sufficient, eliminating the need for dedicated GPUs; in terms of storage resources, the knowledge graph and patient data require approximately 10GB of storage space; in terms of network resources, the hospital's local area network is sufficient for system operation; and in terms of human resources, only basic IT maintenance personnel are needed, and doctors can use it without special training.
[0095] The system implementation follows the principle of "pilot first, then rollout," and includes four phases: prototype validation, data accumulation, effectiveness evaluation, and expansion. In the prototype validation phase, a prototype system was established in the pulmonology department of selected hospitals; in the data accumulation phase, initial screening data was collected, and model parameters were optimized; in the effectiveness evaluation phase, accuracy was verified by comparing the system with traditional screening methods; and in the expansion phase, the system was promoted and applied to more medical institutions.
[0096] In one embodiment of the invention, a multi-level feedback optimization mechanism is further included to improve the accuracy of risk assessment. This mechanism includes parameter self-optimization within modules, result verification between adjacent modules, and system-level global optimization.
[0097] Internal parameter self-optimization refers to each module automatically adjusting its internal parameters based on the characteristics of the input data. For example, the BERT-QA module can dynamically adjust the attention parameter based on the complexity of the questionnaire text; the topological entropy weight module can automatically adjust the parameters of the topological distance function based on the data distribution characteristics; and the fuzzy inference module can dynamically adjust the rule weights based on the rule triggering frequency.
[0098] Result verification between adjacent modules refers to the mutual verification of the reasonableness of results between adjacent modules. For example, the output of the knowledge graph enhancement BERT-QA module can be compared with the preliminary analysis results of the preprocessing module to detect possible anomalies; the output of the topological entropy weight calculation module can be compared with historical data to ensure the consistency of results.
[0099] System-level global optimization refers to system-level adjustments based on the final evaluation results and external feedback. Specific implementations include: 1. Validate intermediate results based on the internal validation set, and calculate model accuracy, sensitivity, specificity, and other metrics; 2. The uncertainty of the estimation results; for cases with high uncertainty, recalculation or request for additional information is triggered. 3. Receive expert feedback, including approval or revision opinions on high-risk assessments, and adjust system parameters accordingly; 4. Regularly update the knowledge graph and model parameters to reflect the latest medical research progress and clinical practice experience.
[0100] Through the aforementioned feedback optimization mechanism, this system can continuously learn and improve, adapt to the characteristics of different groups and regions, and improve the accuracy and applicability of risk assessment.
[0101] The embodiments described above are merely illustrative of specific implementations of the present invention, and while the descriptions are detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A cognitive computing correction method for a lung cancer high-risk population screening questionnaire, characterized in that, The method comprises the following steps: acquiring subject questionnaire text data; constructing a BERT-QA model based on a clinical term knowledge graph, analyzing the semantics of the subject questionnaire text data, and extracting structured risk factor data; mapping the structured risk factor data to a topological space and constructing a risk topological representation; calculating the dynamic weight coefficients of each risk factor in the risk topological representation based on the principle of information entropy to generate an entropy-optimized risk assessment value; calibrating the entropy-optimized risk assessment value in combination with regional pollution map data; determining the high-risk probability of lung squamous cell carcinoma of the subject based on fuzzy logic and the calibrated risk assessment value.
2. The method of claim 1, wherein, The step of acquiring subject questionnaire text data specifically comprises: annotating real words in the subject questionnaire text as "BIO" tags and virtual words as "B" tags; converting the annotated text into XML format for word sense annotation; training a city epidemiology dictionary based on a word2vec algorithm and generating a questionnaire vector representation containing regional characteristics in combination with the epidemiological characteristics of the city where the subject is located.
3. The method of claim 1, wherein, The step of constructing a BERT-QA model based on a clinical term knowledge graph specifically comprises: constructing an entity set containing diseases, symptoms, risk factors, treatment methods, and examination methods, and a relationship set containing causes, correlations, aggravations, alleviations, and detections to form a "entity-relation-entity" triple knowledge representation; converting the triple knowledge representation into an embedding representation, fusing it with the BERT model output through an attention mechanism, and generating a semantic enhanced representation; extracting the smoking age assessment value and the occupational exposure assessment value in the subject questionnaire based on the semantic enhanced representation.
4. The method of claim 1, wherein, The step of mapping the structured risk factor data to a topological space specifically comprises: constructing a multi-dimensional space with risk factors as coordinate axes and defining an open set family containing {low, low-medium, medium, high-medium, high} risk levels; designing a topological distance function based on the characteristics of risk factors to ensure the continuity of risk assessment; constructing a three-level topological mapping process, including mapping of raw data to standardized topological space, mapping of standard space to fuzzy membership space, and mapping of membership space to risk assessment space.
5. The method of claim 1, wherein, The step of calculating the dynamic weight coefficients of each risk factor based on the principle of information entropy specifically comprises: discretizing risk factors into multiple intervals and calculating the probability density distribution of each interval; calculating the information entropy value of each risk factor based on the probability density distribution; generating weight coefficients according to the inverse proportion principle of information entropy, i.e., the smaller the information entropy, the larger the weight coefficient; applying the weight coefficients to the original risk assessment value for weighted calculation to obtain the entropy-optimized risk assessment value.
6. The method of claim 1, wherein, The step based on fuzzy logic specifically comprises: defining fuzzy sets and membership functions based on topological continuity for risk factors; establishing a three-level fuzzy rule base, including a first-level rule for a single risk factor, a second-level rule for multi-factor interaction, and a third-level rule considering environmental factors; performing a fuzzy reasoning process, including converting precise values to fuzzy membership degrees, triggering corresponding fuzzy rules according to inputs, aggregating multi-rule results through the barycenter method, and converting fuzzy results to precise risk values.
7. The method of claim 1, wherein, The step of calibrating the regional pollution map data specifically includes: Collecting regional environmental pollution data containing fine particulate matter (PM2.5), sulfur dioxide, and nitrogen oxides; Establishing a correlation model between pollutants and lung cancer risk, including a direct impact model, a synergistic model, and a long-term cumulative model; According to the pollution level of the subject's region, combined with linear adjustment, threshold adjustment and individual difference adjustment strategy, the risk assessment value is calibrated.
8. The method of claim 1, wherein, The step of determining the high-risk probability of lung squamous cell carcinoma of the subject specifically includes: Integrating the evaluation criteria of multiple international lung cancer screening guidelines, assigning weights based on evidence levels and applicability; Adjusting the risk threshold according to the subject's gender, age, and body mass index; Generating a lung squamous cell carcinoma risk probability value of 0-100% and its confidence interval; When the lung squamous cell carcinoma risk probability value is higher than a certain threshold, the subject is determined to be a high-risk group for lung cancer.
9. The method of claim 1, wherein, Further comprising the following steps: Setting up a multi-level feedback optimization mechanism, including parameter self-optimization within the module, result verification between adjacent modules, and global optimization at the system level; Based on the internal validation set, verify the intermediate results and estimate the uncertainty of the results; Receive expert feedback, dynamically adjust system parameters, and improve the accuracy of risk assessment.
10. A cognitive computing correction system for lung cancer high-risk population screening questionnaires, characterized by, It includes: Questionnaire input and preprocessing module, used to obtain subject questionnaire text data and perform standardized processing; Knowledge graph enhanced BERT-QA module, used to parse questionnaire semantics based on clinical terminology knowledge graph and extract structured risk factor data; Topological entropy weight calculation module, used to construct risk topological representation and calculate dynamic weight coefficients of risk factors; Fuzzy cognitive reasoning module, used to perform fuzzy logic reasoning based on topological structure; Regional pollution fusion module, used to integrate regional environmental pollution data and calibrate risk assessment value; Individualized risk assessment module, used to determine the high-risk probability of lung squamous cell carcinoma of the subject; Among them, the modules interact through standardized data interfaces to form a complete risk assessment pipeline.