Method for evaluating exposure risk of indoor dust mite allergen
By constructing a multi-dimensional interaction effect model, combining genetic factors, objective environment and individual behavior data, and using deep learning and logistic regression models, the high cost and inaccuracy of traditional dust mite allergen evaluation are solved, and low-cost, real-time and accurate dust mite allergen exposure risk assessment and personalized prevention and control are achieved.
Patent Information
- Application Number
- CN202510378834.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-19
AI Technical Summary
In the prior art, dust mite allergen exposure risk assessment relies on laboratory testing, is costly and cannot be applied on a large scale, and lacks considerations for genetic susceptibility, individual behavior and environmental spatiotemporal heterogeneity, resulting in inaccurate assessment.
Multi-dimensional interaction effect modeling is constructed, and by collecting genetic factors, objective environmental data and individual behavior data, using MLP deep learning neural network and multi-factor logistic regression model, a multi-classification prediction model is established, the risk of dust mites is evaluated, and the severity of symptoms is predicted.
It realizes low-cost, real-time and accurate dust mite allergen exposure risk assessment, supports home self-testing, provides personalized prevention and control suggestions, and improves the prospective and convenient evaluation.
Smart Images

Figure CN120511040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of risk assessment of allergic diseases caused by dust mites (allergens), and in particular relates to a method for assessing exposure risk of indoor dust mite allergens. Background Art
[0002] The incidence of allergic diseases has been increasing year by year, among which dust mite allergy is the core environmental factor that induces asthma, allergic rhinitis and atopic dermatitis.
[0003] In existing technologies, dust mite exposure risk assessment mainly relies on laboratory tests (such as ELISA or PCR technology) to directly measure the concentration of dust mite allergens, but such methods have significant defects: A. The test requires high technical requirements in the laboratory environment, and the process is complex and the reagent and labor costs are high, making it impossible to use it on a large scale in clinical practice; B. Currently, doctors / patients and users are unable to conduct self-testing for environmental dust mites at home, and there is a lack of big data accumulation, which seriously affects the diagnosis and treatment of allergic diseases.
[0004] Furthermore, traditional assessment models are often based on a linear correlation between a single environmental indicator (such as temperature, humidity, or dust mite concentration) and symptoms. They lack consideration of genetic susceptibility, another major factor in treatment, as well as real-world differences in individual behavior (such as cleaning habits) and environmental temporal and spatial heterogeneity (such as regional differences in exposure risk). In real life, however, the risk level of dust mite allergen exposure is determined by the interaction of multiple factors.
[0005] Therefore, there is an urgent need for a low-cost, high-precision method to assess the risk of exposure to indoor dust mite allergens that can support real-time intervention. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] A method for assessing the risk of exposure to house dust mite allergens, including:
[0008] Collect genetic factor data and objective environmental data related to allergic factors, as well as individual behavioral data related to allergic factors, to construct the allergic factor-related dataset A;
[0009] Detect dust mite allergen concentrations at different detection points in different environmental areas of the home. Based on the set dust mite allergen concentration standard classification, determine the dust mite exposure risk category to which the dust mite allergen concentration belongs. Add the dust mite exposure risk category as a new dimension to the allergy factor-related dataset A to form the household dust mite exposure risk dataset B.
[0010] The objective environmental data and individual behavior data, as well as the dust mite exposure risk classification, were extracted from the household dust mite exposure risk dataset B to form the original dataset. After encoding the original dataset, a sample dataset was obtained. The sample dataset was input into the MLP deep learning neural network for training, and finally a multi-classification prediction model was obtained. The multi-classification prediction model was used to predict the probability that different detection points belonged to different dust mite exposure risk classifications, and to determine the dust mite exposure risk classification to which they belonged.
[0011] Furthermore, the method further includes collecting data on the diagnosis of allergic diseases and the severity of symptoms of the diagnosed allergic diseases, establishing a diagnostic data set C of allergic disease conditions and symptom severity, and encoding the diagnostic data set C;
[0012] Establish an association algorithm between genetic factors, dust mite exposure risk, allergic disease conditions and symptom severity: encode genetic factor data; construct a multivariate logistic regression model, input the dust mite exposure risk classification code in the household dust mite exposure risk dataset B, the allergic disease code in the diagnostic dataset C, and the genetic factor code into the multivariate logistic regression model, output the probability of the symptom severity of the allergic disease, and predict the symptom severity of the allergic disease.
[0013] Furthermore, the method further includes integrating the genetic factor data, the dust mite exposure risk classification predicted by the multi-classification prediction model, and the diagnostic dataset C, and then quantizing the integration to output a joint feature vector;
[0014] The joint feature vector was input into the multivariate logistic regression model to output the probability of symptom severity.
[0015] Furthermore, a multi-classification prediction model is obtained, including:
[0016] The sample data set is divided into training set and test set using K-fold cross-validation method, and the sample data set is divided into k mutually exclusive subsets. K-1 subsets are selected as training sets each time, and the remaining 1 subset is used as the test set.
[0017] Input the sample data set into the MLP deep learning neural network for training and testing, repeating k times;
[0018] The input layer is used to receive the sample data set; the hidden layer uses the ReLU activation function, which performs nonlinear processing through the piecewise linear function F(x)=Max(0,x);
[0019] The output layer is divided into three nodes, corresponding to three dust mite exposure risk categories: low risk, medium risk, and high risk. The output layer directly outputs the original scores logits of different dust mite exposure risk categories;
[0020] The Softmax function is used to convert the original score logits into a probability distribution and output the probability of different dust mite exposure risk categories.
[0021] Furthermore, it also includes: selecting a loss function and an optimizer to compile, train and evaluate the multi-classification prediction model.
[0022] Furthermore, the Softmax function formula is:
[0023]
[0024] Among them, input: vector Z=[z1,z2...z K ], K is the total number of dust mite exposure risk categories, K = 3, Z i Represents the raw score logits of the i-th dust mite exposure risk category; Output: probability distribution p = [p1, p2...p K ], p i represents the probability of the i-th dust mite exposure risk category; i=1...K, j=1...K.
[0025] Furthermore, the loss function uses the cross entropy loss function, and the formula is:
[0026]
[0027] Among them, y i is the one-hot code of the dust mite exposure risk classification based on the detected dust mite allergen concentration, L is the loss value, p i represents the probability of the i-th dust mite exposure risk category.
[0028] Furthermore, the multivariate logistic regression model is formulated as follows:
[0029]
[0030] Among them, P is the probability of an event occurring; k is the threshold; Y is the severity level of the symptoms, which is divided into three levels: low, moderate, and severe; β0 is the intercept term; β1, β2,…, βp1 are the regression coefficients of the independent variables; X1,…, Xp1 are independent variables, representing the code values of whether the parents have allergic diseases, disease type, dust mite exposure risk classification, and diagnosis type; p1 is the number of independent variables.
[0031] Furthermore, the OR value of dust mite exposure risk to symptom severity was calculated, where the OR value was the corresponding value after indexation of β0…βp1.
[0032] Furthermore, a standard classification of dust mite allergen concentration is set: according to the color development results of the dust mite instant detection reagent, the results indicate that when the dust mite allergen concentration is <2μg / g, it is a low concentration; 2μg / g≤dust mite allergen concentration≤10μg / g is a medium concentration; and dust mite allergen concentration>10μg / g is a high concentration. The results are correspondingly classified as low risk, medium risk and high risk.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. The present invention provides a method for assessing the risk of exposure to indoor dust mite allergens, which improves the foresight, convenience, practicality and accuracy of indoor dust mite exposure risk assessment, is conducive to guiding the prevention and control of indoor dust mite exposure risks, and can help the medical industry provide more intelligent prevention and control services for patients with dust mite allergies.
[0035] 2. This invention establishes complete objective environmental data and individual behavior data for the first time, and simultaneously collects related genetic factor data. It also establishes, for the first time, an association algorithm and prediction model based on large sample sizes and real-world data, linking environmental data, individual behavior, and dust mite allergen concentrations in different indoor areas. This enables dynamic prediction without laboratory testing, replacing traditional laboratory testing, reducing costs and supporting user self-testing and real-time evaluation at home. It offers low cost, fast response, and greater accuracy.
[0036] 3. The present invention's multi-dimensional interaction effect modeling: It pioneered a four-dimensional data collaborative analysis framework of "genetics-environment-behavior-diagnosis", using a multi-factor logistic regression model to quantify the interaction effect between dust mite exposure risk and disease type, breaking through the limitations of single environmental variable analysis.
[0037] 4. This invention uses contribution penetration technology to trace the key driving factors of symptom risk and provide a quantitative basis for personalized intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flow chart of the evaluation method of the present invention. DETAILED DESCRIPTION
[0039] The technical solution of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0040] Example
[0041] like Figure 1 As shown, the present invention provides a method for assessing exposure risk of indoor dust mite allergens, comprising:
[0042] 1. Collect genetic factor data and objective environmental data related to allergic factors, as well as individual behavioral data related to allergic factors, store them in a database, and construct an allergic factor-related dataset A;
[0043] Allergy factor related dataset A includes: genetic factor dataset A1, objective environment dataset A2, and individual behavior dataset A3;
[0044] 1.1. Genetic Factor Dataset A1 Contents, including:
[0045] a. Do parents have allergic diseases: Yes; No;
[0046] b. What kind of allergic disease do the parents have: allergic rhinitis, asthma, atopic dermatitis, others.
[0047] 1.2. The objective environment dataset A2 includes three aspects: the macro environment (i.e., the user's location), the micro environment (the user's home location), and the individual indoor environment (the user's indoor home).
[0048] 1.2.1. Macro Environment: The user's geographical environment, measured using the following dimensions:
[0049] Seasons: spring, summer, autumn, winter;
[0050] Information collection date: specific year, month, and day;
[0051] Region: defined as district-level administrative unit;
[0052] Urban temperature and humidity: Data source: China Weather Network.
[0053] 1.2.2. Microenvironment: This refers to the home's geographical location, including the type and floor of the house, measured using the following dimensions:
[0054] Type of residential housing: record the type of residential housing, such as villa, flat, low-rise, high-rise and other different housing types;
[0055] Residential floor: records the highest floor and residential floor information of a residential building.
[0056] 1.2.3. Individual indoor environment: Indoor home, measured using the following:
[0057] Home indoor temperature and humidity: based on data from indoor temperature and humidity monitors;
[0058] House age: record the construction and use age of the house you live in;
[0059] Lighting and ventilation conditions: record the direction of home windows and daily ventilation time;
[0060] Materials of household items: Record the distribution and material selection of household items, with a focus on sofas, carpets, bedding, mattresses, pillows, plush toys and other items.
[0061] 1.3. Individual behavior dataset A3, whose dimensions include household cleaning methods, cleaning frequency, selection and use of household anti-mite products, etc.
[0062] Cleaning methods for household items: Record the cleaning methods for different household items, including the types of household items, specific cleaning methods, equipment and supplies used in the cleaning process, and cleaning personnel, etc.
[0063] Cleaning frequency of household items: Based on the choice of household items and the cleaning method, record the cleaning frequency of household items.
[0064] Selection of household anti-mite products: Record the selection and use of anti-mite related products, including the type of products, brand, usage methods, and subjective feelings.
[0065] For example, the data dimensions in the allergy factor-related data set A finally collected include: parents' allergy history, season, information collection date, region, city temperature and humidity, family geographical location (including the type of house and floor), family indoor temperature and humidity, house age, orientation, ventilation conditions, distribution of household items, material, cleaning frequency, cleaning method, brand of anti-mite products, usage of anti-mite products, etc.
[0066] The data in the allergy factor related data set A collected in step 1 are cleaned. Data cleaning includes removing data items with null values and outliers in the allergy factor related data set A in step 1; checking whether each data item has missing values, and using statistical methods (such as mean interpolation, nearest neighbor interpolation, etc.) to fill the missing values; using statistical methods (such as Z-score or IQR method) to detect outliers and eliminate outliers.
[0067] 2. Based on the detection of instant dust mite detection reagents, the dust mite allergen concentrations at different detection points in different environmental areas of the home are tested, and based on the set standard classification of dust mite allergen concentrations, the dust mite exposure risk category to which the dust mite allergen concentration belongs is determined; the set standard classification of dust mite allergen concentrations: according to the color development results of the instant dust mite detection reagents, the results indicate that when the dust mite allergen concentration is <2μg / g, it is a low concentration; 2μg / g≤dust mite allergen concentration≤10μg / g is a medium concentration; and dust mite allergen concentration>10μg / g is a high concentration. The results are accordingly classified as low risk, medium risk and high risk.
[0068] The dust mite exposure risk classification is added as a new dimension to the allergy factor related dataset A to establish the household dust mite exposure risk dataset B; specifically,
[0069] Detect dust mite allergen concentrations at different test points in the home, including:
[0070] 1) Use a vacuum cleaner equipped with a high-efficiency air filter to vacuum different test points in different environmental areas of the home for 1 minute and collect ash samples.
[0071] 2) Use the instant dust mite detection reagent to test the dust mite allergen concentration in the ash samples at different detection points. The test results are divided into three types: low concentration, medium concentration, and high concentration.
[0072] 3) Dust mite allergen concentrations (low, medium, and high) were mapped to dust mite exposure risk categories (low, medium, and high). This category was added as a new dimension to the allergy-related dataset A, associated with the objective environment and individual behavior dimensions, and stored in the database to form the household dust mite exposure risk dataset B. The objective environment data and individual behavior data, as well as the dust mite exposure risk categories, were extracted from the household dust mite exposure risk dataset B to form the original dataset. The data in the original dataset is shown in Table 1 below (genetic factor data will be used later in this article, so the data example only contains objective environment data, individual behavior data, and dust mite exposure risk categories, and does not include genetic factors):
[0073] Table 1
[0074] season area temperature humidity floor House use period Home Point Material Individual behavior Dust mite exposure risk spring Shanghai 23 54 6 5 Quilt fabric Clean once a month High risk
[0075] 3. Collect data on the diagnosis of allergic diseases and the severity of symptoms of allergic diseases, and establish a diagnostic dataset C of allergic disease status and symptom severity; specifically,
[0076] Do you have allergic diseases? a. Yes, b. No;
[0077] Diagnosis: a. Allergic rhinitis; b. Asthma; c. Atopic dermatitis; d. Comorbidity,
[0078] Symptom severity:
[0079] a. Allergic rhinitis:
[0080] i. Use a visual assessment scale (VAS) to assess symptom severity, with 0 being "no symptoms" and 10 being "very severe." Different diseases correspond to different VAS scales.
[0081] ii. Nasal congestion, runny nose, sneezing, itchy nose, itchy eyes, tearing, red and swollen eyes, and eye pain.
[0082] iii. For each item, 1 to 3 points indicate mild; 4 to 7 points indicate moderate; and 8 to 10 points indicate severe.
[0083] b. Asthma:
[0084] i. Asthma symptom severity was scored using the Asthma Control Test score, or ACT score. The ACT score is shown in Table 2 below.
[0085] Table 2
[0086] ACT Questionnaire and Scoring Criteria
[0087]
[0088] Note: Scoring method: Step 1: Record the score for each question; Step 2: Add up the scores for each question to get the total score; Step 3: The meaning of ACT score: A score of 20-25 indicates good asthma control; 16-19 indicates poor asthma control; 5-15 indicates very poor asthma control
[0089] c. Atopic dermatitis: The AD score is used to score the severity of atopic dermatitis symptoms. The AD score is shown in Table 3 (consisting of Tables 3.1-3.4):
[0090] Table 3.1
[0091] AD score (SCORAD)
[0092] Scoring is based on three aspects: first, the area of lesions (i.e., the percentage of each site on the body surface); second, the severity of lesions (the corresponding condition of the skin as observed by the user); and finally, the subjective symptom score (including the degree of itching and sleep loss).
[0093] Percentage of body surface area occupied by each body part (adults and children over 2 years old)
[0094] neck 9% (4.5% each before and after) upper limbs 9% on one side (4.5% each on the front and back) lower limbs 18% on one side (9% each on the front and back) chest and abdomen 18% back 18% perineum 1%
[0095] Table 3.2
[0096] Percentage of each body part on the body surface (children under 2 years old)
[0097] neck 18% (9% each before and after) upper limbs 9% on one side (4.5% each on the front and back) lower limbs 14% on one side (7% each on the front and back) chest and abdomen 18% back 18%
[0098] Table 3.3
[0099] Severity score
[0100] erythema edema Oozing / crusting Exfoliation Thickening of the skin dry skin 0 none none none none none none 1 Mild Mild Mild Mild Mild Mild 2 Moderate Moderate Moderate Moderate Moderate Moderate 3 severe severe severe severe severe severe
[0101] Table 3.4
[0102] Subjective symptom score
[0103] itching 0 (no itching) to 10 (extreme itching) sleep loss 0 (no sleep loss) - 10 (unable to fall asleep)
[0104] SCORAD score = skin lesion area score / 5 + 7 × severity score / 2 + subjective symptom score
[0105] SCORAD mild: 0-24 points, moderate: 25-50 points, severe: >50 points
[0106] Finally, the data in the diagnostic dataset C are obtained as shown in Table 4 below:
[0107] Table 4
[0108]
[0109] 4. Establish a multi-classification prediction model combining objective environmental data, individual behavioral data, and dust mite exposure risk classification. The multi-classification prediction model has two independent variables: objective environmental data and individual behavioral data; and the dependent variable is dust mite exposure risk classification.
[0110] The specific construction process of the multi-classification prediction model includes:
[0111] 4.1. Build a sample dataset for the multi-classification prediction model and divide it into training and test sets;
[0112] Obtaining a sample data set, including:
[0113] The original data set consisting of objective environment data and individual behavior data is feature encoded and converted into a combined feature vector and a classification vector. The combined feature vector and the classification vector are concatenated to form a feature vector as a sample data set.
[0114] Example of raw data (Table 5):
[0115] Table 5
[0116] season area temperature humidity floor House use life Home Point Material Individual behavior Dust mite exposure risk spring Shanghai 23 54 6 5 Quilt fabric Clean once a month High risk
[0117] Construction of feature vector:
[0118] The variables related to the influencing factors of the objective environment and individual behavior are called features. In the original data set, features are stored as discrete textual data or numerical data. When constructing feature vectors, it is necessary to convert discrete textual data into numerical data, and the numerical data in the original data set is retained.
[0119] Use One-Hot encoding to encode textual data into numerical data.
[0120] Taking the seasonal characteristics in Table 5 as an example, one-hot encoding is performed, and the encoded results are shown in Table 6:
[0121] Table 6
[0122] season spring summer Autumn winter spring 1 0 0 0 summer 0 1 0 0 Autumn 0 0 1 0 winter 0 0 0 1
[0123] Each season is converted into a vector of length 4, such as the feature vector of spring is [1,0,0,0].
[0124] Perform One-Hot encoding on other textual data and sequentially concatenate all single feature vectors to form a combined feature vector: E = {e1, e2, ...e n},
[0125] Examples of combined feature vectors in Table 6:
[0126] [1,0,0,0,0,0,1,23,54,6,5,0,0,1,0,0,1,0,1,0],
[0127] Dust mite exposure risk classification is text data, which is also encoded using One-Hot encoding to construct the classification vector, as shown in Table 7.
[0128] Table 7
[0129] Dust mite exposure risk Low risk Medium risk High risk Low risk 1 0 0 Medium risk 0 1 0 High risk 0 0 1
[0130] The dust mite exposure risk classification is converted into a vector of length 3. For example, the classification vector for "high risk" is [0, 0, 1]. The feature vector and classification vector are called a piece of sample data. The entire original dataset is converted into vector form and used as the sample dataset for the multi-classification prediction model and as the data input for the MLP deep learning neural network.
[0131] Training set and test set division:
[0132] Use the K-fold cross-validation method to divide the sample data set into a training set and a test set; divide the sample data set into k mutually exclusive subsets, select k-1 subsets as the training set each time, and the remaining 1 subset as the test set; repeat k times and take the average result.
[0133] 4.2. Input the sample data set into the MLP deep learning neural network for training to obtain a multi-classification prediction model. Specifically,
[0134] The MLP deep learning neural network is divided into input layer, hidden layer and output layer.
[0135] The input layer is used to receive sample data sets. The hidden layer is a multi-layer fully connected layer used to learn state representation based on the sample data sets.
[0136] The hidden layer uses the ReLU activation function. ReLU introduces nonlinearity into the model through the piecewise linear function F(x)=Max(0,x), enabling the MLP deep learning neural network to learn complex nonlinear mapping relationships.
[0137] The output layer is used to output the predicted probability of dust mite exposure risk classification.
[0138] The output layer is divided into 3 nodes, corresponding to 3 categories, and directly outputs logits, that is, the original scores.
[0139] Use the Softmax function to convert the original score logits into a probability distribution
[0140] The Softmax function formula is:
[0141]
[0142] Among them, input: vector Z=[z1,z2...z K ], K is the total number of dust mite exposure risk categories, K = 3, Z i Represents the raw score logits of the i-th dust mite exposure risk category; Output: probability distribution p = [p1, p2...p K ], p i represents the probability of the i-th dust mite exposure risk category; i=1...K, j=1...K.
[0143] Function: The Softmax function amplifies the influence of high scores through exponential operation, and then normalizes it to probability so that high scores correspond to high probabilities.
[0144] Assuming that the logits output of the MLP deep learning neural network for the three categories is: Z = [3.0, 1.0, 0.2], corresponding to the categories [low risk, medium risk, high risk], the specific calculation process of the probability of mite exposure risk classification is given:
[0145] Step 1: Calculate the index value:
[0146] e 3.0 ≈20.085,e 1.0 ≈2.7183,e 0.2 ≈1.2214,
[0147] Step 2: Calculate the denominator, the sum of all exponents:
[0148] 20.0855+2.7183+1.2214≈24.0252,
[0149] Step 3: Calculate the probability of each classification:
[0150]
[0151] This means that there is an 83.6% probability that the output is low risk.
[0152] Select loss function and optimizer to compile, train and evaluate the multi-classification prediction model:
[0153] Use the cross entropy loss function:
[0154]
[0155] where y i is the one-hot code of the dust mite exposure risk classification based on the detected dust mite allergen concentration, L is the loss value, p i represents the probability of the i-th dust mite exposure risk category;
[0156] For example, if the dust mite exposure risk classification vector is high risk [0, 0, 1], the predicted probability is [0.1, 0.2, 0.7].
[0157] According to the cross entropy loss function formula, the loss is calculated as:
[0158] L=-(0*log(0.1)+0*log(0.2)+1*log(0.7)=-log(0.7)≈0.3567,
[0159] The Adam optimizer is used to update the parameters. The data loader is iterated in each epoch, the gradients are calculated, and the weights are updated. The sample dataset is fed into the input layer of the MLP deep learning neural network. The multi-classification prediction model is trained and tested on the test set. The following are example results of the multi-classification prediction model training, as shown in Table 8.
[0160] Table 8
[0161] Epoch[10 / 100], Loss:0.6872
[0162] Epoch[20 / 100], Loss:0.5251 ...
[0164] Epoch[100 / 100], Loss: 0.2138
[0165] TeSt ACCuraCy: 0.93
[0166] The results show that the multi-classification prediction model achieved an accuracy of 93% on the test set.
[0167] By training the original data with the MLP deep learning neural network, a multi-classification prediction model for dust mite exposure risk classification based on objective environmental data and individual behavioral data was constructed.
[0168] For newly generated data, by constructing the feature vector of the new data, the multi-classification prediction model can output the classification prediction.
[0169] The output data of this multi-class prediction model can be used as input for predicting the subsequent occurrence of allergic diseases and the severity of their symptoms.
[0170] During the implementation process, each feature will be continuously modified according to the real-world situation. This method can be repeated each time the feature vector is adjusted and the model is retrained with the newly generated data to achieve continuous evolution of the model.
[0171] 5. Develop an algorithm to correlate genetic factors and dust mite exposure risk with allergic disease prevalence and symptom severity, including:
[0172] A multivariate logistic regression model was used to analyze the impact of genetic factors, household dust mite exposure risk, and dust mite exposure risk at different detection points on disease status and symptom severity. This includes:
[0173] Encode the genetic factor data; construct a multivariate logistic regression model, input the dust mite exposure risk classification code in the household dust mite exposure risk dataset B, the allergic disease code in the diagnostic dataset C, and the genetic factor code into the multivariate logistic regression model, and output the severity of the symptoms of the allergic disease.
[0174] 5.1. Classification and coding of variables in the multivariate logistic regression model:
[0175] Its independent variables include:
[0176] Environmental factors: household dust mite exposure risk classification, number of high-risk locations (e.g. bedroom high risk is scored as 1, otherwise 0);
[0177] Genetic factors: whether parents have allergic diseases and disease diagnosis;
[0178] Diagnostic factors: disease type (categorical variable coding: 0 = rhinitis, 1 = asthma, 2 = dermatitis), comorbidity status (0 / 1).
[0179] Its dependent variable (target variable):
[0180] Symptom severity (ordered categorical variable): The three allergic diseases each have their own severity levels.
[0181] 5.2 Construction of multi-factor logistic regression model:
[0182] Model formula:
[0183]
[0184] Among them, P is the probability of an event occurring; k is the threshold; Y is the severity level of the symptoms, which is divided into three levels: low, moderate, and severe; β0 is the intercept term; β1, β2,…, βp1 are the regression coefficients of the independent variables; X1,…, Xp1 are independent variables, representing the code values of whether the parents have allergic diseases, disease type, dust mite exposure risk classification, and diagnosis type; p1 is the number of independent variables.
[0185] Based on the multivariate logistic regression model, the key analysis contents are:
[0186] Calculate the OR value for whether the parents have allergic diseases, disease type, and dust mite exposure risk for symptom severity. The OR value is the indexed value of β0…βp1 calculated in the multivariate logistic regression model. Output the probability of specific diseases at high-risk sites and identify the association between high-risk sites and specific diseases.
[0187] The multivariate logistic regression model inputs the patient's household dust mite exposure data and diagnostic information, outputs the probability of symptom severity, and predicts the severity of symptoms of the allergic disease.
[0188] 6. Associate the multi-classification prediction model with the multi-factor logistic regression model; collect genetic factor data, objective environmental data, individual behavior data, and diagnostic data, derive the severity of symptoms, and quickly predict the association between the environment and allergic symptoms without the need for dust mite testing. Specifically,
[0189] The multi-classification prediction model is associated with the multi-factor logistic regression model to construct an association model, which includes an input layer, an intermediate layer, and an output layer;
[0190] Input layer: objective environmental data, individual behavior data, and disease types in diagnostic data.
[0191] Data preprocessing: A multi-classification prediction model is used to convert objective environmental and individual behavior data into predicted dust mite exposure risk classification (low / medium / high).
[0192] Output format: probability vector of dust mite exposure risk classification (e.g. [0.85, 0.12, 0.03] means 85% probability of dust mite exposure risk classification being low risk).
[0193] Middle layer: Integrate genetic factor data, dust mite exposure risk classification predicted by the multi-classification prediction model, and diagnostic data (disease type, comorbidity status) into a joint feature vector (for example: [0.85, 0.12, 0.03, 1, 0, 0], where the last 3 digits encode asthma diagnosis).
[0194] Output layer: Use a multivariate logistic regression model, input the joint feature vector, and output the probability of symptom severity (probability distribution of mild / moderate / severe).
[0195] Contribution penetration analysis:
[0196] Through the size of the OR value, the contribution of each factor can be traced, and the key driving factors of the symptom prediction results can be traced (such as "the probability of severe symptoms is 70% contributed by the high-risk exposure predicted by the multi-classification prediction model, of which the weight of bedroom humidity accounts for 55%).
[0197] The above technical features constitute the best embodiment of the present invention, which has strong adaptability and best implementation effect. Non-essential technical features can be added or removed according to actual needs to meet the needs of different situations.
[0198] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions of the technical solution of the present invention by ordinary technicians in this field do not deviate from the essence and scope of the technical solution of the present invention.
Claims
1. A method for assessing exposure risk of indoor dust mite allergens, characterized in that: include: Collect genetic factor data and objective environmental data related to allergic factors, as well as individual behavioral data related to allergic factors, to construct the allergic factor-related dataset A; Detect dust mite allergen concentrations at different detection points in different environmental areas of the home. Based on the set dust mite allergen concentration standard classification, determine the dust mite exposure risk category to which the dust mite allergen concentration belongs. Add the dust mite exposure risk category as a new dimension to the allergy factor-related dataset A to form the household dust mite exposure risk dataset B. The objective environmental data and individual behavior data, as well as the dust mite exposure risk classification, were extracted from the household dust mite exposure risk dataset B to form the original dataset. After encoding the original dataset, a sample dataset was obtained. The sample dataset was input into the MLP deep learning neural network for training, and finally a multi-classification prediction model was obtained. The multi-classification prediction model was used to predict the probability that different detection points belonged to different dust mite exposure risk classifications, and to determine the dust mite exposure risk classification to which they belonged.
2. The evaluation method according to claim 1, wherein: It also includes collecting data on the diagnosis of allergic diseases and the severity of symptoms of diagnosed allergic diseases, establishing a diagnostic data set C of allergic disease conditions and symptom severity, and encoding the diagnostic data set C; Establish an algorithm to associate genetic factors, dust mite exposure risk, and allergic disease status and symptom severity: Encode genetic factor data; A multivariate logistic regression model was constructed. The dust mite exposure risk classification code in the household dust mite exposure risk dataset B, the allergic disease code in the diagnosis dataset C, and the genetic factor code were input into the multivariate logistic regression model. The probability of the symptom severity of the allergic disease was output to predict the symptom severity of the allergic disease.
3. The evaluation method according to claim 2, wherein: It also includes integrating the genetic factor data, the dust mite exposure risk classification predicted by the multi-classification prediction model, and the diagnostic dataset C, and then quantizing them to output a joint feature vector; The joint feature vector was input into the multivariate logistic regression model to output the probability of symptom severity.
4. The evaluation method according to claim 1, wherein: Obtain a multi-classification prediction model, including: The sample data set is divided into training set and test set using K-fold cross-validation method, and the sample data set is divided into k mutually exclusive subsets. K-1 subsets are selected as training sets each time, and the remaining 1 subset is used as the test set. Input the sample data set into the MLP deep learning neural network for training and testing, repeating k times; The input layer is used to receive the sample data set; the hidden layer uses the ReLU activation function, which performs nonlinear processing through the piecewise linear function F(x)=Max(0,x); The output layer is divided into three nodes, corresponding to three dust mite exposure risk categories: low risk, medium risk, and high risk. The output layer directly outputs the original scores logits of different dust mite exposure risk categories; The Softmax function is used to convert the original score logits into a probability distribution and output the probability of different dust mite exposure risk classifications.
5. The evaluation method according to claim 4, characterized in that Also includes: Select loss functions and optimizers to compile, train, and evaluate multi-class prediction models.
6. The evaluation method according to claim 4, characterized in that The Softmax function formula is: Among them, input: vector Z=[z1,z2...z K ], K is the total number of dust mite exposure risk categories, K = 3, Z i Represents the raw score logits of the i-th dust mite exposure risk category; Output: probability distribution p = [p1, p2...p K ], p i represents the probability of the i-th dust mite exposure risk category; i=1...K, j=1...K.
7. The evaluation method according to claim 6, characterized in that Loss function, using the cross entropy loss function, the formula is: Among them, y i is the one-hot code of the dust mite exposure risk classification based on the detected dust mite allergen concentration, L is the loss value, p i represents the probability of the i-th dust mite exposure risk category.
8. The evaluation method according to claim 2, wherein: The multivariate logistic regression model is as follows: Among them, P is the probability of an event occurring; k is the threshold; Y is the severity level of the symptoms, which is divided into three levels: low, moderate, and severe; β0 is the intercept term; β1, β2,…, βp1 are the regression coefficients of the independent variables; X1,…, Xp1 are independent variables, representing the code values of whether the parents have allergic diseases, disease type, dust mite exposure risk classification, and diagnosis type; p1 is the number of independent variables.
9. The evaluation method according to claim 8, characterized in that The OR value of the risk of dust mite exposure to symptom severity was calculated, where the OR value was the corresponding value after indexation of β0…βp1.
10. The evaluation method according to claim 1, wherein: The set dust mite allergen concentration standard classification: According to the color development results of the dust mite instant detection reagent, the results indicate that when the dust mite allergen concentration is <2μg / g, it is a low concentration; 2μg / g≤dust mite allergen concentration≤10μg / g is a medium concentration; dust mite allergen concentration>10μg / g is a high concentration. The results are correspondingly classified as low risk, medium risk and high risk.
Citation Information
Patent Citations
Incremental neural network model-based mite dermatitis prediction method and prediction system
CN106384011A
Method and biomedical electronic device for predicting target exceeding of allergen content
CN110069022A
Alopecia type identification and detection method and system based on comparative learning
CN119480087A
Allergy onset risk prediction system, method and program
JP2016177678A
Allergic disease determination method and determination system
US20230324405A1