Disease early warning method based on odor pattern recognition, electronic device, and storage medium

CN122814835APending Publication Date: 2026-09-25SHANGHAI FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610900466.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于气味模式识别的疾病预警方法、电子设备及存储介质,该疾病预警方法克服了现有仅以主含分子单一映射进行疑似判定,且难以合理解释共性成分重叠的缺陷,实现了对多疾病并发情形的联合推断与误警抑制,最终提升了可靠性与临床可用性

Benefits of technology

[0016]本发明提供的基于气味模式识别的疾病预警方法,通过构建无量纲的多成分气体响应动力学特征并进行跨通道结构化建模,解决了过度依赖主含分子的单向关联、忽略多成分相关性及时变演化特性,而导致的多病理并存或共性分子重叠场景下鉴别分辨率不足的缺陷;通过非负稀疏映射生成非负强度向量与残差评分因子,并构建组合概率分布模型匹配疾病权重矩阵,解决了病理特征与气味特征强度关联缺失的可信度量化缺陷,实现了以概率分布置信区间量化的预警可信度输出;通过建立面向多疾病共存的组合概率判别模型,克服了现有仅以主含分子单一映射进行疑似判定,且难以合理解释共性成分重叠的缺陷,实现了对多疾病并发情形的联合推断与误警抑制,最终提升了可靠性与临床可用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122814835A_ABST
    Figure CN122814835A_ABST
Patent Text Reader

Abstract

The present application relates to a disease early warning method based on odor pattern recognition, electronic equipment and storage medium, relate to odor detection and early warning technical field, the disease early warning method includes synchronous acquisition target area in multichannel gas sensor array response signal sequence, and uniform time stamp environment parameter sequence, through sliding time window segmentation response signal sequence, combine pre-constructed interference feature library, the baseline drift of response signal driven by environment parameter is corrected, output and time window alignment response signal segment, response signal segment extracts gas response dynamics feature set containing steady state and transient information. The advantage lies in: through the establishment of combination probability discriminant model for multiple diseases coexistence, overcome the defect that only the main containing molecule single mapping is used for suspected determination, and it is difficult to reasonably explain the overlap of common components, realizes the joint inference and false alarm suppression of multiple disease concurrent situation, and finally improves the reliability and clinical availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of odor detection and early warning technology, and more specifically, to a disease early warning method, electronic device and storage medium based on odor pattern recognition. Background Technology

[0002] Currently, odor-based disease early warning technology using electronic noses has been integrated into digital healthcare recommendation frameworks in multiple countries. Its system architecture consists of a multi-channel chemical sensor array, signal conditioning and data acquisition modules, feature extraction and discrimination engines, and alarm linkage devices. The basic workflow is as follows: Within the monitored space, response signals from each channel to specific volatile organic compounds or inorganic gases are continuously acquired. Through a calibration model, the response values ​​are converted into target molecule concentrations or equivalent intensity values. Early warnings are triggered based on preset thresholds or rule bases, and a feature molecule database or predefined rule base is used to identify suspected disease states. In engineering deployment, environmental parameters such as temperature, humidity, and airflow are simultaneously incorporated for basic calibration and quality control. Combined with data cleaning, baseline correction, and sliding time window statistical calculations, the timeliness and operational feasibility of alarms are ensured. While the current solution has achieved continuous monitoring and rapid response of key molecules in multiple scenarios in laboratory environments, its environmental robustness in threshold triggering mechanisms, the limited pathological specificity of odor fingerprint characterization, and the lack of interpretability in clinical decision support are gradually becoming apparent during migration to complex real-world spaces.

[0003] Existing electronic nose odor warning systems employ multi-channel sensor arrays to estimate the concentration of target molecules and use fixed concentration thresholds or their static combinations as alarm trigger conditions. They then rely on the main molecules of individual diseases for suspected odor assessment. This approach has the following drawbacks: the fixed concentration threshold strategy lacks adaptive adjustment capabilities to fluctuations in environmental temperature and humidity and time-varying sensor drift, leading to increased false alarm and false negative rates; reliance on unidirectional correlations of main molecules ignores the structured correlations and temporal dynamic response characteristics between multiple components, resulting in insufficient resolution for scenarios involving multiple coexisting diseases or overlapping common molecules; and the lack of interpretable odor feature intensity characterization and uncertainty quantification models that match disease pathological characteristics makes it difficult to quantify the reliability of warnings, ultimately limiting the reliability and clinical usability of odor warnings in complex scenarios.

[0004] The statements herein provide only background information in relation to this invention and do not necessarily constitute prior art. Summary of the Invention

[0005] The purpose of this invention is to provide a disease early warning method, electronic device and storage medium based on odor pattern recognition. This disease early warning method overcomes the shortcomings of existing methods that only use a single mapping of the main constituent molecules for suspected cases and have difficulty in reasonably explaining the overlap of common components. It realizes joint inference and false alarm suppression for multiple concurrent diseases, and ultimately improves reliability and clinical usability.

[0006] This invention provides a disease early warning method based on odor pattern recognition, comprising the following steps: S1: Synchronously acquire the response signal sequence of the multi-channel gas sensor array within the target area, as well as the environmental parameter sequence with a unified timestamp; S2: The response signal sequence is segmented by a sliding time window, combined with a pre-built interference feature library, and the baseline drift of the response signal driven by environmental parameters is corrected by the environmental parameter sequence, and the response signal segment aligned with the time window is output. S3: Perform the following operations for each response signal segment: Step S31: Extract the gas response dynamics feature set containing steady-state and transient information from the response signal segment, and perform inter-channel normalization and cross-channel ratio construction on the feature set to generate a dimensionless feature vector; Step S32: Perform non-negative sparse mapping between the dimensionless feature vector and the pre-constructed disease odor component dictionary. By minimizing the reconstruction residual and sparsity constraints, output the non-negative intensity vector and its residual scoring factor. Step S33: Match the intensity values ​​of each component in the non-negative intensity vector with the weight distribution of the corresponding disease in the preset odor component-disease association weight matrix using a probability model, and output the independent matching degree of each disease. Step S34: Based on the independent matching degree, different candidate disease sets are screened, and the expected intensity value and variance of the pre-set intensity value of each disease-related odor component are extracted. A combined probability distribution model is constructed by weighted cumulative intensity value expectation and combined variance, and the combined matching degree of the non-negative intensity vector under the combined probability distribution model is calculated. S4: The calculated combination matching degree is weighted and fused according to the time window and compared with the adaptive threshold. The weight is the residual scoring factor, and the output includes the warning signal containing the disease combination and the intensity value of the odor component.

[0007] Furthermore, step S2, which combines a pre-built interference feature library and uses an environmental parameter sequence to correct the baseline drift of the response signal driven by environmental parameters, includes the following steps: S21: Extract the environmental parameters within the current time window from the environmental parameter sequence, retrieve the interference feature library based on the time window index, and obtain candidate interference entries that match the current environmental parameters; S22: Calculate the cosine similarity between the response signal segment and each candidate interference item, and select the candidate interference items that exceed the preset similarity threshold as interference sources; S23: Input the temperature, humidity and airflow velocity in the environmental parameter sequence into the piecewise linear regression model, output the dynamic compensation coefficient and apply it to the response signal segment for correction; S24: When the interference source is not empty, project the eigenvector component of the interference source from the response signal segment and then output it.

[0008] Further, step S31 includes the following steps: S311: Calculate the steady-state mean, peak amplitude, and signal integral area of ​​each channel in the response signal segment to form a steady-state feature subset; S312: Extract the extreme points of the rise time, recovery time and first derivative sequence of each channel in the response signal segment to form a transient feature subset; S313: Concatenate the steady-state feature subset and the transient feature subset according to the channel number to form the original feature vector; S314: Obtain the median baseline library of the multi-channel gas sensor array, perform inter-channel median normalization on the original feature vector, and generate gain-invariant feature vectors; S315: Calculate the peak value ratio and mean value ratio of adjacent channels and combine them into a ratio vector; S316: Concatenate the ratio vector and the gain-invariant eigenvector into a dimensionless eigenvector.

[0009] Further, step S32 involves performing a non-negative sparse mapping between the dimensionless feature vector and the pre-constructed dictionary of disease odor components, including the following steps: S321: Obtain the historical background residual library of the multi-channel gas sensor array; S322: Using the dictionary of disease odor components as the basis matrix, minimize the reconstruction residual of the dimensionless eigenvector, and apply L1 norm regularization constraints to solve for the non-negative intensity vector. S323: Extract the reconstructed residual vector corresponding to the non-negative intensity vector, calculate its KL divergence value with the historical background residual library, and generate residual scoring factors.

[0010] Furthermore, the output of the non-negative intensity vector and its residual scoring factor in step S32 includes exponential smoothing and a consistency verification mechanism, wherein the exponential smoothing and consistency verification mechanism includes: Before step S4 is executed, the number of consecutive time windows for which the combined matching degree has been calculated is counted. When the count reaches a preset number, the non-negative intensity vectors of the consecutive preset number of time windows are exponentially weighted and smoothed before being output. When the residual scoring factor is lower than the preset consistency threshold, steps S2 to S3 are executed again to recalibrate the response signal sequence to generate a new response signal segment. Based on the new response signal segment, the non-negative intensity vector is resolved and the residual scoring factor is generated again until the preset consistency threshold is met or the maximum number of iterations is reached.

[0011] Furthermore, the recalibration includes a parameter adjustment mechanism, which includes: increasing the preset similarity threshold and expanding the order of piecewise linear regression.

[0012] Furthermore, step S34, which involves filtering different candidate disease sets based on independent matching degree, includes the following steps: S341: Based on the comparison between independent matching degree and preset matching degree interval, the diseases are divided into a main set and a sub-set, and the preset intensity value expectation of all disease-related odor components is extracted. S342: Construct different initial subsets from one or more diseases in the main set, and combine them with the intensity value expectation weighted by the odor component-disease association weight matrix to generate an intensity value expectation vector; S343: With the minimum number as a constraint, independently select diseases or combinations thereof from the subset and include them in the initial subset. Iterate and update the expected intensity vector until its Euclidean distance with the non-negative intensity vector meets the preset deviation range or reaches the maximum number of inclusions. S344: Denote the initial subset that meets the preset deviation range as the candidate disease set, calculate the sum of the independent matching degree of each disease in it, and output the candidate disease set in descending order accordingly.

[0013] Furthermore, the generation of the adaptive threshold in step S4 includes the following steps: S41: Set a default value for the adaptive threshold; S42: When there are more than m valid warning signals, a fixed number of time windows before the warning signal is triggered are taken as the warning period, and the combined matching degree values ​​of all time windows within it are extracted to form the total matching degree sequence. S43: Concatenate the total matching degree sequence corresponding to the most recent n warning periods into a global matching degree sequence according to the time window order, input the global matching degree sequence into the moving average model, and output the baseline threshold curve; S44: Collect all residual score factor values ​​from the most recent n warning periods to form a global residual sequence, calculate its standard deviation and mean, and use the ratio of the standard deviation to the mean as the adjustment coefficient; S45: Generate a smoothly adjusted adaptive threshold by scaling the baseline threshold curve by adjusting the coefficients.

[0014] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described disease early warning method based on odor pattern recognition.

[0015] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described disease early warning method based on odor pattern recognition.

[0016] The disease early warning method based on odor pattern recognition provided by this invention addresses the shortcomings of insufficient discrimination resolution in scenarios with multiple coexisting pathologies or overlapping common molecules by constructing dimensionless multi-component gas response dynamics features and performing cross-channel structured modeling. This is achieved by over-relying on the unidirectional correlation of main molecules and ignoring the correlation and time-varying evolution characteristics of multiple components. By generating non-negative intensity vectors and residual scoring factors through non-negative sparse mapping and constructing a combined probability distribution model to match the disease weight matrix, the method solves the quantification defect of the lack of reliable correlation between pathological features and odor feature intensity, and realizes the output of early warning credibility quantified by probability distribution confidence interval. By establishing a combined probability discrimination model for multiple coexisting diseases, the method overcomes the shortcomings of existing methods that only use a single mapping of main molecules for suspected cases and are difficult to reasonably explain the overlap of common components. This enables joint inference and false alarm suppression in the case of multiple coexisting diseases, ultimately improving reliability and clinical usability. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a disease early warning method based on odor pattern recognition, provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of a disease early warning system based on odor pattern recognition provided in an embodiment of the present invention. Detailed Implementation

[0019] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, used only to facilitate and clearly illustrate the embodiments of the present invention. Please refer to the drawings to make the objectives, features, and advantages of the present invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the implementation conditions of the present invention. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to the size, without affecting the effects and objectives achieved by the present invention, should still fall within the scope of the technical content disclosed in the present invention.

[0020] Figure 1 This is a flowchart illustrating a disease early warning method based on odor pattern recognition according to an embodiment of the present invention. Please refer to... Figure 1 The disease early warning method based on odor pattern recognition provided in this embodiment of the invention includes the following steps: S1: Synchronously acquire the response signal sequence of the multi-channel gas sensor array within the target area, as well as the environmental parameter sequence with a unified timestamp.

[0021] S2: The response signal sequence is segmented by a sliding time window, combined with a pre-built interference feature library, and the baseline drift of the response signal driven by environmental parameters is corrected by the environmental parameter sequence, and the response signal segment aligned with the time window is output; for example, the sliding time window length can be 90 seconds and the step size is 30 seconds.

[0022] Specifically, step S2 of this invention combines a pre-constructed interference feature library and uses an environmental parameter sequence to correct the baseline drift of the response signal driven by environmental parameters, including the following steps: S21: Extract the environmental parameters within the current time window from the environmental parameter sequence, retrieve the interference feature library based on the time window index, and obtain candidate interference entries that match the current environmental parameters.

[0023] It should be noted that the construction of the interference feature library is based on the environmental background characteristics and high-frequency interference patterns of the target area. For example, the construction process may be as follows: During periods of low traffic or closure in the target area (such as a department) (e.g., 2:00-4:00 AM daily), the response signal sequence and environmental parameter sequence (temperature, humidity, airflow velocity) of a multi-channel gas sensor array are simultaneously collected; the signal is segmented using a 90-second sliding time window (30-second step), and the steady-state feature vector and variance vector of each channel are calculated to form the basic entries of the interference feature library; for several common environmental interference sources (including 75% ethanol disinfectant, chlorine-containing cleaning agents, isopropanol solvent, etc.), Repeated sampling was performed under simulated laboratory environmental parameter combinations (temperature: 10 / 20 / 30 / 40℃, humidity: 30 / 50 / 70%) to extract standardized feature vectors. Both the standardized and steady-state feature vectors were [normalized steady-state mean, normalized peak amplitude, normalized signal integral area, normalized peak-to-peak ratio of adjacent channels]. Finally, a three-level index library was constructed through hierarchical management of environmental parameters: Level 1 is categorized by temperature (<15℃, 15-25℃, >25℃), Level 2 by humidity (<40%, 40-60%, >60℃), and Level 3 is coded by interfering object type, achieving coverage of common clinical interference scenarios.

[0024] S22: Calculate the cosine similarity between the response signal segment and each candidate interference item, and select the candidate interference items that exceed the preset similarity threshold as interference sources; for example, the preset similarity threshold can be 0.82.

[0025] S23: Input the temperature, humidity, and airflow velocity from the environmental parameter sequence into the piecewise linear regression model, output dynamic compensation coefficients, and apply them to the response signal segment for correction.

[0026] For example, a piecewise linear regression model can be used to establish the mapping relationship between environmental parameters (temperature T, humidity H, airflow velocity V) and baseline drift, specifically as follows: When the environmental parameters are in the range of T < 20°C and H < 50%, the compensation equation is K = 0.12T - 0.08H + 0.05V; When the environmental parameters are in the range of 20°C≤T≤30°C and H≥50%, the compensation equation is K=0.09T+0.15H-0.03V; When the environmental parameter range is T>30°C, the compensation equation is K=0.21T+0.04H+0.12V; The corrected response signal segment = the original response signal segment × (1 + K); where K is the dynamic compensation coefficient calculated based on the compensation equation.

[0027] S24: When the interference source is not empty, project the eigenvector component of the interference source from the response signal segment and then output it.

[0028] For example, the process of matching candidate interference items and performing projection subtraction is as follows: Input the mean environmental parameters of the current time window (e.g., temperature 23.5℃, humidity 58%), search the interference feature library with a tolerance range (temperature ±2℃, humidity ±5%), hit the first-level index (25℃ bin) and the second-level index (40-60% humidity bin), and output the subset of candidate items for the temperature range and humidity range; calculate the cosine similarity between the 4-dimensional steady-state feature vector of the response signal segment and each candidate item; select the candidate item with a preset similarity threshold ≥0.82 and the largest cosine similarity value as the interference source, input it together with the response signal segment vector (4-dimensional steady-state feature vector) into the projection difference formula, and output the response signal segment vector after subtracting the feature vector components of the candidate item.

[0029] The projection difference formula is as follows: ; In the formula, This is the vector of the response signal segment after projection subtraction. For the response signal segment vector, This is the standardized feature vector corresponding to the interference source.

[0030] It should be noted that this invention achieves baseline drift correction driven by environmental parameters, eliminating the influence of residual environmental interference signals on odor feature extraction.

[0031] S3: Perform the following operations for each response signal segment: Step S31: Extract the gas response dynamics feature set containing steady-state and transient information from the response signal segment, and perform inter-channel normalization and cross-channel ratio construction on the feature set to generate a dimensionless feature vector.

[0032] Specifically, step S31 of the present invention includes the following steps: S311: Calculate the steady-state mean, peak amplitude, and signal integral area of ​​each channel in the response signal segment to form a steady-state feature subset; the purpose of setting the steady-state mean, peak amplitude, and signal integral area is to avoid the limitations of a single indicator and to comprehensively reflect the concentration level and total exposure.

[0033] S312: Extract the extreme points of the rise time, recovery time and first derivative sequence of each channel in the response signal segment to form a transient feature subset; the purpose of setting the extreme points of the rise time, recovery time and first derivative sequence is to capture the inflection point of the gas response speed and dynamic change.

[0034] S313: Concatenate the steady-state feature subset and the transient feature subset according to the channel number to form the original feature vector.

[0035] S314: Obtain the median baseline library of the multi-channel gas sensor array, perform inter-channel median normalization on the original feature vector, and generate gain-invariant feature vectors.

[0036] It should be noted that the purpose of splicing by channel number is to preserve the spatial correspondence between channels; the median baseline library refers to a reference database that stores the reference response of a multi-channel sensor array, which can be built based on factory calibration data or historical stable state samples, with the purpose of providing a reference benchmark for the inherent characteristics of the sensor; the purpose of median normalization between channels is to eliminate gain inconsistencies caused by manufacturing deviations between multiple channels; the purpose of peak ratio and mean ratio of adjacent channels is to offset common drift caused by environmental factors such as temperature and humidity.

[0037] S315: Calculate the peak ratio and mean ratio of adjacent channels and combine them into a ratio vector.

[0038] S316: Concatenate the ratio vector and the gain-invariant eigenvector into a dimensionless eigenvector.

[0039] Furthermore, if the current system's computing resources are insufficient, the dimensionless eigenvectors can be reduced using the PCA (principal component analysis) method to reduce the subsequent computational load and improve the system's robustness. To facilitate understanding of the above embodiments, a specific application scenario of the above embodiments will be used as an example for illustration below: The application scenario is for the differential diagnosis of early-stage lung cancer (characteristic gas: nonanal) and chronic obstructive pulmonary disease (characteristic gas: nitric oxide). The practical problem is that the responses of the two diseases overlap on a single sensor channel (e.g., nonanal and nitric oxide in...). All sensors produced positive responses.

[0040] and The sensor synergistically captures the carbonyl polarity and carbon chain hydrophobicity of nonanal. and The sensors amplify the oxidation properties of nitric oxide in a coordinated manner, hence these four gas sensors were selected.

[0041] A1. Feature Extraction (Taking a 4-channel sensor array as an example): Steady-state feature subset: Ch1( Steady-state mean = 120, peak amplitude = 185, signal integral area = 1520; Ch2( Steady-state mean = 85, peak amplitude = 130, signal integral area = 980; Ch3( Steady-state mean = 210, peak amplitude = 320, signal integral area = 2640; Ch4( Steady-state mean = 95, peak amplitude = 150, signal integral area = 1180; Transient feature subset: Ch1: Rise time = 8.2s, recovery time = 12.1s, first derivative extreme point = 22.5; Ch2: Rise time = 12.5s, recovery time = 18.3s, first derivative extreme point = 15.8; Ch3: Rise time = 9.8s, recovery time = 14.7s, first derivative extreme point = 28.3; Ch4: Rise time = 11.3s, Recovery time = 16.9s, First derivative extreme point = 17.2; A2. Median normalization (used to compensate for sensor aging drift and environmental gain fluctuations, such as suppressing response amplitude variations caused by ±30% humidity to ±5%, improving the comparability of cross-channel response signals): For each channel (Ch1-Ch4), the mean baseline, peak baseline, area baseline, rise time baseline, recovery time baseline, and first derivative extremum are obtained from the median baseline library. These baselines are used to normalize each component in the steady-state feature subset and the transient feature subset, respectively. A 24-dimensional (4 channels × 6 features) gain-invariant feature vector is generated (detailed data calculation steps omitted).

[0042] In particular, median normalization was also performed on the peak and mean values ​​of adjacent channels: Channel Ch1: Median baseline is 100, normalized mean is 1.20, and normalized peak value is 1.85; Channel Ch2: median baseline is 80, normalized mean is 1.06, and normalized peak value is 1.63; Channel Ch3: median baseline is 200, normalized mean is 1.05, and normalized peak value is 1.60; Channel Ch4: median baseline is 90, normalized mean is 1.06, and normalized peak is 1.67.

[0043] A3. Calculate the peak ratio and mean ratio of adjacent channels (used to convert the cross-response differences of the sensor array into disease-specific features, and expand the feature space separation of similar diseases): Peak ratios: Ch1 / Ch2 = 1.85 / 1.63 ≈ 1.13; Ch2 / Ch3 = 1.63 / 1.60 ≈ 1.02; Ch3 / Ch4 = 1.60 / 1.67 ≈ 0.96; Mean ratios: Ch1 / Ch2 = 1.20 / 1.06 ≈ 1.13; Ch2 / Ch3 = 1.06 / 1.05 ≈ 1.01; Ch3 / Ch4 = 1.05 / 1.06 ≈ 0.99; The peak ratio and mean ratio of each channel are concatenated to form a ratio vector (6 dimensions): [1.13, 1.02, 0.96, 1.13, 1.01, 0.99].

[0044] A4. Finally, the ratio vector (6-dimensional) and the gain-invariant eigenvector (24-dimensional) are concatenated into a dimensionless eigenvector (30-dimensional): used to fuse absolute response intensity and relative distribution patterns, such as retaining nonanal in... and The sensor employs a "high mean ratio + low peak ratio" discrimination mode.

[0045] This invention enhances the specificity of disease odor patterns by constructing cross-channel ratios and fusing multi-dimensional features, thereby improving the identification resolution and robustness of disease odor component matching.

[0046] Step S32: Perform non-negative sparse mapping between the dimensionless feature vector and the pre-constructed disease odor component dictionary. By minimizing the reconstruction residual and sparsity constraints, output the non-negative intensity vector and its residual scoring factor.

[0047] Specifically, step S2 of the present invention performs a non-negative sparse mapping between the dimensionless feature vector and the pre-constructed dictionary of disease odor components, including the following steps: S321: Obtain the historical background residual library of the multi-channel gas sensor array; where the historical background residual library refers to the collection of historical noise data of the multi-channel gas sensor array under normal operating conditions. It can be implemented using non-volatile memory or cloud database, with the purpose of providing an objective noise distribution reference benchmark for reconstructing residuals. S322: Using the dictionary of disease odor components as the basis matrix, minimize the reconstruction residual of the dimensionless eigenvector, and apply L1 norm regularization constraints to solve for the non-negative intensity vector.

[0048] Furthermore, the expression for solving the non-negative intensity vector is: ; In the formula, This is a dictionary of disease odor components (based matrix, typically p×r in dimension). It is a dimensionless eigenvector (vector dimension is p×1). It is a non-negative intensity vector. It is an L1 norm. It is the L2 norm. Non-negativity constraint It is the L1 regularization coefficient and is set to 0.1.

[0049] It should be noted that non-negative constraints are used to ensure that the gas concentration intensity conforms to physical reality, while the L1 norm regularization sparse activation mechanism is used to screen significant disease odor components. The two work together to reduce the confusion rate of cross-diseases while suppressing overfitting of environmental noise.

[0050] S323: Extract the reconstructed residual vector corresponding to the non-negative intensity vector, calculate its KL divergence value with the historical background residual library, and generate residual scoring factors.

[0051] Furthermore, by solving the non-negative intensity vectors of each component of the disease odor through L1 regularization constraints, the sparsity screening of disease types is achieved while ensuring reconstruction accuracy (i.e., key odor components are screened out). The KL divergence value is calculated to quantify the similarity between the reconstructed residual vector and the historical background residual distribution, thereby obtaining the residual scoring factor that characterizes the reliability of the mapping. That is, the residual scoring factor can accurately reflect the degree to which the response signal deviates from the normal state.

[0052] The formula for calculating the KL divergence value is as follows: ; In the formula, To reconstruct the probability distribution of the residuals, The probability distribution of residuals in the historical context. For the feature dimension index in the probability distribution, The value is the KL divergence value. It should be noted that the probability distribution of the reconstructed residuals is obtained by normalizing the reconstructed residual vector, and the probability distribution of the historical background residuals is obtained simultaneously. All probability distributions are represented by vectors, with each term of the vector representing a feature dimension.

[0053] This invention can objectively assess the degree of signal anomaly based on historical noise distribution, distinguish between real odor signals and environmental interference noise, and achieve sparsity screening of disease types.

[0054] Furthermore, in step S32 of the present invention, the output of the non-negative intensity vector and its residual scoring factor includes exponential smoothing and a consistency verification mechanism, wherein the exponential smoothing and consistency verification mechanism includes: Before step S4 is executed, the continuous time windows for which the combined matching degree has been calculated are counted. When the count reaches a preset number, the non-negative intensity vectors of the continuous preset number of time windows are exponentially weighted smoothed and then output. The formula for exponentially weighted smoothing is as follows: ; In the formula, This is the non-negative intensity vector smoothly output during the current time window. This represents the original non-negative intensity vector for the current time window. The non-negative intensity vector of the previous time window. The non-negative intensity vectors for the preceding preset number of time windows. For the preset quantity, These are the weighting coefficients.

[0055] When the residual scoring factor is lower than the preset consistency threshold, steps S2 to S3 are executed again to recalibrate the response signal sequence and generate a new response signal segment. Based on the new response signal segment, the non-negative intensity vector is resolved and the residual scoring factor is generated again until the preset consistency threshold is met or the maximum number of iterations is reached.

[0056] For example, the data covers the entire human respiratory cycle (approximately 45 seconds) and includes at least 3 effective gas exchange events. Therefore, the preset number can be 3, corresponding to data within a single time window of 90 seconds. The preset consistency threshold is set based on the KL divergence normal distribution characteristics of the historical background residual library. The mean and standard deviation in this normal distribution are extracted, and the sum of the mean and twice the standard deviation is used as the preset consistency threshold. The maximum number of iterations can be 5.

[0057] Furthermore, the recalibration includes a parameter adjustment mechanism, which includes: increasing the preset similarity threshold; and expanding the order of piecewise linear regression.

[0058] Among them, the increase in the preset similarity threshold and the increment of the regression order are both directly proportional to the number of iterations.

[0059] For example, the increment of the similarity threshold is set to 0.02 × the number of iterations; the increment of the regression order is set to 0.5 × the number of iterations; and the upper limit of the similarity threshold increment is 0.1 to avoid excessively strict signal loss, and the upper limit of the regression order increment is 3 to prevent overfitting.

[0060] Specifically, exponential weighted smoothing is applied to the non-negative intensity vectors within a consecutive preset number of time windows to suppress jitter in time-series response signal segments. Response signal segments are generated by recalibrating the response signal sequence, and the non-negative intensity vectors and residual scoring factors are recalculated, forming a closed-loop feedback control. During calibration, a parameter adjustment mechanism is activated, which increases the preset similarity threshold to rigorously screen interference sources and expands the piecewise linear regression order to enhance compensation accuracy. Both the increase in the preset similarity threshold and the increment of the regression order increase linearly with the number of iterations, ensuring that the calibration strength is gradually optimized within a finite number of iterations. In summary, exponential smoothing addresses the stability issue of response signal segments within historical time windows, the consistency verification mechanism handles the quality monitoring of the real-time obtained non-negative intensity vectors, and the parameter adjustment mechanism provides a path to improve calibration accuracy. These three elements form an adaptive calibration closed loop.

[0061] This invention suppresses the temporal jitter of non-negative intensity vectors, improves the continuity and stability of odor component intensity characterization, and ensures the reliability of non-negative intensity vector reconstruction quality through a dynamic correction mechanism, thereby enhancing the credibility and consistency of disease warning signals and providing a high-precision data foundation for odor pattern recognition in complex environments.

[0062] Step S33: Match the intensity values ​​of each component in the non-negative intensity vector with the weight distribution of the corresponding disease in the preset odor component-disease association weight matrix using a probability model, and output the independent matching degree of each disease.

[0063] Specifically, the expression for calculating the independent matching degree can be: ; In the formula, For independent matching degree, Let be the dimension of the non-negative intensity vector. Let i be the association weight of odor in disease d. Let i be the intensity value of the i-th dimension of the non-negative intensity vector. , Let be the expected value and variance of the i-th dimension of the intensity of disease d. Let i be the Gaussian kernel function, and i be the dimension number and represent the corresponding odor.

[0064] Step S34: Based on the independent matching degree, different candidate disease sets are screened, and the expected intensity value and variance of the pre-set intensity value of each disease-related odor component are extracted. A combined probability distribution model is constructed by weighted cumulative intensity value expectation and combined variance, and the combined matching degree of the non-negative intensity vector under the combined probability distribution model is calculated.

[0065] Specifically, the calculation process for the combined matching degree can be as follows: V1, Weighted cumulative strength value expectation: ; In the formula, This is the combined expectation vector of the combined probability distribution model. For candidate disease set, Number the disease. Let d be the expected vector of the intensity values ​​of disease d. is the normalized value of the independent matching degree of disease d; where the intensity value expectation vector is composed of the intensity values ​​corresponding to all associated odors of disease d.

[0066] V2, Combined Variance: ; In the formula, The covariance matrix of the combined probability distribution model. Let be the variance vector of disease d. This is an operator that converts a vector into a diagonal matrix; where the variance vector is composed of the variances of all associated odors of disease d.

[0067] V3. Calculate the combination matching degree: ; In the formula, For combined matching degree, For dimension index variables, Let i be the value of the non-negative intensity vector in the i-th dimension. Let i be the value of the combined expected vector in the i-th dimension. Let be the value of the covariance matrix in the i-th dimension, and n be the dimension of the non-negative intensity vector.

[0068] Furthermore, step S34 of the present invention, which involves filtering different candidate disease sets based on independent matching degree, includes the following steps: S341: Based on the comparison between the independent matching degree and the preset matching degree interval, the diseases are divided into a main set and a sub-set, and the preset intensity value expectation of all disease-related odor components is extracted; for example, the preset matching degree interval includes the sub-set matching degree interval (0.5, 0.85) and the main set matching degree interval [0.85, 1.0].

[0069] S342: Construct different initial subsets from one or more diseases in the main set, and combine them with the intensity value expectation weighted by the odor component-disease association weight matrix to generate an intensity value expectation vector.

[0070] S343: With a minimum number as a constraint, independently select diseases or combinations thereof from the subset and include them in the initial subset. Iterate and update the expected intensity vector until its Euclidean distance with the non-negative intensity vector meets the preset deviation range or reaches the maximum number of inclusions. The minimum number constraint refers to the minimum number of diseases to be included in the subset, which is dynamically adjusted according to the number of diseases in the main set. For example, when the number of diseases in the main set is less than 3, the minimum number ranges from 1 to 3; when the number of diseases in the main set exceeds 3, the minimum number ranges from 2 to 4. Based on the normal distribution of the distance between the expected intensity vector and the measured non-negative intensity vector of historical confirmed cases, extract their mean and standard deviation, and use the sum of the two as the boundary value of the preset deviation range (typical value ±0.15). In order to meet the clinical diagnostic timeliness requirements (≤5 seconds), it was determined through complexity testing that most scenarios can be covered within 7 times, so the maximum number of inclusions can be 7 times.

[0071] S344: Denote the initial subset that meets the preset deviation range as the candidate disease set, calculate the sum of the independent matching degree of each disease in it, and output the candidate disease set in descending order accordingly.

[0072] Specifically, by dividing diseases into a main set and a subset, and extracting the expected intensity values ​​of all disease-associated odor components, an initial subset is constructed from the diseases in the main set. This subset is then weighted and fused with the expected intensity values ​​using an odor component-disease association weight matrix to generate a unified expected intensity vector. This integrates the odor components of multiple diseases into a quantitative representation of pathological association intensity. Diseases are selected from the subset and included in the initial subset with a minimum number of diseases as a constraint. Iterative updates gradually bring the expected intensity vector closer to a non-negative intensity vector while preventing excessive expansion of the combination. Finally, subsets that deviate from the specified range are designated as candidate disease sets and output in descending order based on the sum of independent matching degrees, thus quantifying the synergistic effect of the candidate disease sets.

[0073] This invention integrates the structured complementary relationships between disease odor components. The aim is to enumerate as many different single or multiple disease combinations as possible under the constraint that the intensity value expectation vector approaches a non-negative intensity vector, and to calculate the combination matching degree for each of these combinations, thereby selecting the disease combination with the highest confidence and improving the accuracy of disease early warning.

[0074] S4: The calculated combination matching degree is weighted and fused according to the time window and compared with the adaptive threshold. The weight is the residual scoring factor, and the output includes the warning signal containing the disease combination and the intensity value of the odor component.

[0075] Specifically, the pre-construction process of the disease odor component dictionary is as follows: its essence is a basis matrix (D), with each column corresponding to the odor component feature vector of a disease. The specific steps are as follows: Data Acquisition: During low-interference periods (e.g., 2:00–4:00 AM) in the target area (e.g., a respiratory clinic), the response signal sequence and environmental parameter sequence (temperature, humidity, airflow velocity) of the multi-channel gas sensor array are collected. For preset diseases (disease types are extracted from the "odor component-disease association weight matrix," taking five diseases including lung cancer, chronic obstructive pulmonary disease, asthma, tuberculosis, and pneumonia as examples), exhaled breath samples from confirmed patients are collected, and the response signal sequence is repeatedly sampled in a laboratory simulating the clinic's environmental parameter combination (temperature: 10 / 20 / 30 / 40℃, humidity: 30 / 50 / 70%).

[0076] Feature extraction: The signal is segmented using a 90-second sliding time window (30-second step) to extract dimensionless feature vectors (30-dimensional) as samples; an input data matrix V (dimension p×q) is constructed, where each column is a sample, p=30, and q is the number of samples; a basis matrix D (dimension p×r) is extracted, where r is the number of odors (odors are the main molecules contained in 5 diseases); a sparse coefficient matrix S is constructed by projecting all samples onto the basis matrix, and the non-negative intensity vector s is its single-sample form.

[0077] Dictionary generation: The Non-negative Matrix Factorization (NMF) algorithm is applied, with diseases as cluster centers, to solve for the basis matrix D. The expression for the Non-negative Matrix Factorization algorithm is as follows: ; In the formula, For the input data matrix, As a basis matrix, It is a sparse coefficient matrix. It is the Frobenius norm. Let be the L1 regularization coefficient and set to 0.1. It is an L1 norm. It indicates that the matrix elements are non-negative.

[0078] For example, a dictionary of disease odor components could be: Taking a 4-channel sensor array (Ch1-Ch4) as an example, they are as follows: , , , Gas sensor; Odor ID1: Formaldehyde (HCHO), dimensionless feature vector [0.85, 0.12, 0.33, 0.45, ..., 0.18...] (30 dimensions); Odor ID2: Nitric oxide (NO); Odor ID3: Nonanal; Odor ID4: Ethane; Odor ID5: Isoprene; Each ID contains a dimensionless feature vector, which is omitted here.

[0079] Furthermore, taking the dimensionless eigenvector of ID1 as an example, the numerical values ​​of the components represent the gas response dynamics characteristics of the response signals acquired by each sensor. For example, 0.85 corresponds to... Steady-state mean, 0.12 is Peak amplitude, 0.33 The signal integration area is 0.18, which is the peak ratio of Ch1 to Ch2.

[0080] For example, a pre-defined odor component-disease association weight matrix (behavioral disease types, listed as odor weights), taking two rows of the matrix as an example: Line 1 (Lung Cancer): [0.10, 0.15, 0.90, 0.05, 0.25]; Line 2 (Chronic Obstructive Pulmonary Disease): [0.05, 0.85, 0.90, 0.05, 0.25]; The disease corresponding to row 1 is lung cancer. The values ​​in the columns correspond to the association weights of formaldehyde 0.10, nitric oxide 0.15, nonanal 0.90, ethane and isoprene 0.05, respectively.

[0081] The expected intensity and variance of a disease can be denoted as follows (taking lung cancer as an example): Lung cancer: [0.12 / 0.02, 0.18 / 0.03, 0.92 / 0.01, 0.08 / 0.04, 0.30 / 0.02]; where the value 0.12 / 0.02 represents formaldehyde (expected intensity / variance).

[0082] Specifically, the response signal segments generated by the sensor array are used to generate dimensionless feature vectors through feature extraction; the dimensionless feature vectors are used to solve for non-negative intensity vectors through non-negative sparse mapping (based on a disease odor component dictionary) to represent the intensity distribution of the current odor components; the construction of the disease odor component dictionary depends on the odor component features of common diseases in the target area (common diseases are derived from preset diseases in the odor component-disease association weight matrix); the non-negative intensity vectors are probabilistically matched with the odor component-disease association weight matrix to output the independent matching degree; at the same time, the expected intensity and variance of the disease-related odor components in the weight matrix are preset to construct a combined probability distribution model.

[0083] Furthermore, the generation of the adaptive threshold in step S4 of the present invention includes the following steps: S41: Set a default value for the adaptive threshold; based on the normal distribution of the combined matching degree of historical data of healthy people, take the sum of its mean and 3 times the standard deviation as the default value, and the typical value can be 0.6.

[0084] S42: When there are more than m valid warning signals, a fixed number of time windows before the warning signal is triggered are used as the warning period, and the combination matching degree values ​​of all time windows within it are extracted to form the total matching degree sequence; where, a valid warning signal is that after the warning command is triggered, the corresponding disease is diagnosed through subsequent treatment, that is, within this period of time in the target area, there has been a combination of the corresponding disease and the concentration of the odor component has reached the warning standard.

[0085] S43: The total matching degree sequence corresponding to the most recent n warning periods is concatenated into a global matching degree sequence according to the time window order and input into the moving average model to output the baseline threshold curve; where n is less than m; at least 5 effective warning signals need to be accumulated to build statistical significance, so m is 5; in order to balance the real-time performance of the model and historical representativeness, the data of the most recent 3 warning periods are selected, which can capture short-term seasonal patterns and meet the computational efficiency, so n is 3.

[0086] It should be noted that, based on the typical metabolic response cycle of the respiratory system, clinical studies have confirmed that the rise in the concentration of pathological markers needs to last for 15 minutes. Taking a 90-second time window as an example, the fixed number is 30.

[0087] S44: Collect all residual score factor values ​​from the most recent n warning periods to form a global residual sequence, calculate its standard deviation and mean, and use the ratio of the standard deviation to the mean as the adjustment coefficient.

[0088] S45: Generate a smoothly adjusted adaptive threshold by scaling the baseline threshold curve by adjusting the coefficients.

[0089] Specifically, by setting a default value as the initial operating benchmark for the adaptive threshold, the system is ensured to still have basic early warning capabilities even when historical data is insufficient. When more than m effective early warning signals accumulate, a fixed number of time windows before the early warning signal is triggered are defined as the early warning period. The combined matching degree values ​​within this period are extracted to form a total matching degree sequence, accurately capturing typical matching degree patterns before the early warning occurs. The total matching degree sequences of the most recent n early warning periods are concatenated in chronological order to form a global matching degree sequence and input into a moving average model. By smoothing short-term random fluctuations, long-term early warning patterns are highlighted, and the output benchmark threshold curve reflects the stable trend of combined matching degree changes. The residual scoring factor values ​​of the most recent n early warning periods are combined to form a global residual sequence, and the ratio of its standard deviation to the mean is calculated as an adjustment coefficient. This coefficient quantifies the uncertainty of historical early warnings based on the residual scoring factor. Finally, the benchmark threshold curve is scaled by the adjustment coefficient to generate a smoothed adaptive threshold. When historical data fluctuates greatly, the coefficient expands the threshold range to reduce false alarms, and when fluctuations are small, the coefficient tightens the threshold to reduce false alarms, achieving dynamic matching between the adaptive threshold and the actual scenario.

[0090] The adaptive threshold of this invention can be dynamically calibrated based on the statistical characteristics of historical warning signals, enabling it to intelligently adjust according to the fluctuations of the actual warning scenario. This effectively distinguishes real disease signals from interference noise when environmental parameters fluctuate or sensor time-varying drift occurs, reducing false alarm and false negative rates, especially in scenarios where multiple diseases coexist or common molecules overlap.

[0091] Another embodiment of the present invention provides a disease early warning system based on odor pattern recognition, wherein the disease early warning system is applied to the disease early warning method in the above embodiment.

[0092] Figure 2 This is a schematic diagram of a disease early warning system based on odor pattern recognition provided in an embodiment of the present invention. Please refer to... Figure 2 The disease early warning system based on odor pattern recognition in this embodiment includes a signal synchronous acquisition module, a signal preprocessing module, a response signal segment processing module, and a disease early warning signal output module.

[0093] The signal synchronization acquisition module is used to synchronously acquire the response signal sequence of the multi-channel gas sensor array within the target area, as well as the environmental parameter sequence with a unified timestamp.

[0094] The signal preprocessing module is used to segment the response signal sequence through a sliding time window, combine it with a pre-built interference feature library, and use the environmental parameter sequence to correct the baseline drift of the response signal driven by the environmental parameters, and output a response signal segment aligned with the time window.

[0095] The response signal segment processing module includes a feature extraction and normalization construction unit, an odor component sparse mapping unit, a disease independent matching degree unit, and a combined matching degree calculation unit.

[0096] The feature extraction and normalization construction unit is used to extract a gas response dynamics feature set containing steady-state and transient information for each response signal segment, and to perform inter-channel normalization and cross-channel ratio construction on the feature set to generate a dimension-compressed dimensionless feature vector.

[0097] The odor component sparse mapping unit is used to perform non-negative sparse mapping between the dimensionless feature vector and the pre-constructed disease odor component dictionary. By minimizing the reconstruction residual and sparsity constraints, it outputs a non-negative intensity vector and its residual scoring factor.

[0098] The disease independent matching degree unit is used to perform probabilistic model matching between the intensity value of each component in the non-negative intensity vector and the weight distribution of the corresponding disease in the preset odor component-disease association weight matrix, and output the independent matching degree of each disease.

[0099] The combined matching degree calculation unit is used to filter different candidate disease sets based on independent matching degree, extract the preset intensity value expectation and variance of each disease-related odor component, construct a combined probability distribution model by summing the intensity value expectation and the combined variance, and calculate the combined matching degree of the non-negative intensity vector under the combined probability distribution model.

[0100] The disease early warning signal output module is used to perform time-window weighted fusion of the calculated combination matching degree and compare it with an adaptive threshold to output an early warning signal containing the disease combination and the intensity value of the odor component. The weight of the weighted fusion is the residual scoring factor.

[0101] Another embodiment of the present invention provides a computer storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the disease early warning method based on odor pattern recognition described above.

[0102] Another embodiment of the present invention provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the disease early warning method based on odor pattern recognition described above.

[0103] As can be seen from the above description, the advantages of this invention are: 1. The disease early warning method of the present invention solves the defects of insufficient identification resolution in scenarios with multiple coexisting pathologies or overlapping common molecules by constructing dimensionless multi-component gas response dynamic characteristics and performing cross-channel structured modeling. 2. The disease early warning method of the present invention generates a non-negative intensity vector and residual scoring factor through non-negative sparse mapping, and constructs a combined probability distribution model to match the disease weight matrix, which solves the credible quantification defect of missing correlation between pathological features and odor feature intensity, and realizes the output of early warning credibility quantified by probability distribution confidence interval. 3. The disease early warning method of the present invention overcomes the shortcomings of existing methods that rely solely on a single mapping of the main constituent molecules for suspected cases and are difficult to reasonably explain the overlap of common components by establishing a combined probability discrimination model for the coexistence of multiple diseases. It achieves joint inference and false alarm suppression for the coexistence of multiple diseases, and ultimately improves reliability and clinical usability.

[0104] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0105] To ensure that the technical solution of this application complies with the provisions of Article 5, Paragraph 1 of the Patent Law and does not constitute a violation of the law, social morality, or harm to the public interest, the applicant hereby makes the following compliance statement.

[0106] The data collection, storage, and processing involved in this application strictly comply with current laws and regulations. Specifically, all user-related data to be processed in this solution originates from legal channels, and explicit authorization and consent from the data subjects have been obtained through user agreements or privacy policies during the data collection phase. For cases involving sensitive personal information, separate consent from the users has been further obtained, strictly adhering to the principle of informed consent. In the actual data processing process, if identity-identifying information is involved, this solution employs anonymization or de-identification techniques to ensure that the processed data cannot identify any specific natural person and cannot be restored. The data processing activities do not infringe upon personal privacy rights and comply with the provisions of the "Personal Information Protection Law of the People's Republic of China," the "Cybersecurity Law of the People's Republic of China," and other relevant laws. If this solution involves collecting information in public places, the purpose of collection is limited to maintaining public safety or specific scenarios with explicit user authorization, and does not exceed the necessary limits.

[0107] Regarding the algorithmic decision-making mechanism, the technical solution of this application fully considers technological ethics and social morality. The training data, label settings, and decision-making logic used in the model construction process do not introduce discriminatory or unfair parameters based on factors such as gender or age. The algorithm output is optimized based on objective technical parameters, rather than making value judgments based on inherent individual attributes. For scenarios involving decision-making models or recommendation algorithms, the decision-making basis of this solution is strictly limited to objective physical parameters, technical indicators, or a pre-defined priority of legal obligations, and does not contain any content that violates the right to equality of life or social morality. The business logic and algorithmic recommendation content of this solution do not contain content that harms public interests, disrupts social order, or induces undesirable behavior.

[0108] In summary, the technical solution claimed in this application ensures its legality, compliance, and ethics through technical means throughout the entire chain of data acquisition, data processing, and result output, and there is no situation that violates the law, social morality, or harms the public interest.

[0109] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

Claims

1. A disease early warning method based on odor pattern recognition, characterized in that, Includes the following steps: S1: Synchronously acquire the response signal sequence of the multi-channel gas sensor array within the target area, as well as the environmental parameter sequence with a unified timestamp; S2: The response signal sequence is segmented by a sliding time window, combined with a pre-built interference feature library, and the baseline drift of the response signal driven by environmental parameters is corrected by the environmental parameter sequence, and the response signal segment aligned with the time window is output. S3: Perform the following operations for each response signal segment: Step S31: Extract the gas response dynamics feature set containing steady-state and transient information from the response signal segment, and perform inter-channel normalization and cross-channel ratio construction on the feature set to generate a dimensionless feature vector; Step S32: Perform non-negative sparse mapping between the dimensionless feature vector and the pre-constructed disease odor component dictionary. By minimizing the reconstruction residual and sparsity constraints, output the non-negative intensity vector and its residual scoring factor. Step S33: Match the intensity values ​​of each component in the non-negative intensity vector with the weight distribution of the corresponding disease in the preset odor component-disease association weight matrix using a probability model, and output the independent matching degree of each disease. Step S34: Based on the independent matching degree, different candidate disease sets are screened, and the expected intensity value and variance of the pre-set intensity value of each disease-related odor component are extracted. A combined probability distribution model is constructed by weighted cumulative intensity value expectation and combined variance, and the combined matching degree of the non-negative intensity vector under the combined probability distribution model is calculated. S4: The calculated combination matching degree is weighted and fused according to the time window and compared with the adaptive threshold. The weight is the residual scoring factor, and the output includes the warning signal containing the disease combination and the intensity value of the odor component.

2. The disease early warning method based on odor pattern recognition as described in claim 1, characterized in that, Step S2, which combines a pre-built interference feature library and uses environmental parameter sequences to correct the baseline drift of the response signal driven by environmental parameters, includes the following steps: S21: Extract the environmental parameters within the current time window from the environmental parameter sequence, retrieve the interference feature library based on the time window index, and obtain candidate interference entries that match the current environmental parameters; S22: Calculate the cosine similarity between the response signal segment and each candidate interference item, and select the candidate interference items that exceed the preset similarity threshold as interference sources; S23: Input the temperature, humidity and airflow velocity in the environmental parameter sequence into the piecewise linear regression model, output the dynamic compensation coefficient and apply it to the response signal segment for correction; S24: When the interference source is not empty, project the eigenvector component of the interference source from the response signal segment and then output it.

3. The disease early warning method based on odor pattern recognition as described in claim 1, characterized in that, Step S31 includes the following steps: S311: Calculate the steady-state mean, peak amplitude, and signal integral area of ​​each channel in the response signal segment to form a steady-state feature subset; S312: Extract the extreme points of the rise time, recovery time and first derivative sequence of each channel in the response signal segment to form a transient feature subset; S313: Concatenate the steady-state feature subset and the transient feature subset according to the channel number to form the original feature vector; S314: Obtain the median baseline library of the multi-channel gas sensor array, perform inter-channel median normalization on the original feature vector, and generate gain-invariant feature vectors; S315: Calculate the peak value ratio and mean value ratio of adjacent channels and combine them into a ratio vector; S316: Concatenate the ratio vector and the gain-invariant eigenvector into a dimensionless eigenvector.

4. The disease early warning method based on odor pattern recognition as described in claim 1, characterized in that, Step S32 involves performing a non-negative sparse mapping between the dimensionless feature vector and the pre-constructed dictionary of disease odor components, including the following steps: S321: Obtain the historical background residual library of the multi-channel gas sensor array; S322: Using the dictionary of disease odor components as the basis matrix, minimize the reconstruction residual of the dimensionless eigenvector, and apply L1 norm regularization constraints to solve for the non-negative intensity vector. S323: Extract the reconstructed residual vector corresponding to the non-negative intensity vector, calculate its KL divergence value with the historical background residual library, and generate residual scoring factors.

5. The disease early warning method based on odor pattern recognition as described in claim 1, characterized in that, The output of the non-negative intensity vector and its residual scoring factor in step S32 includes exponential smoothing and a consistency verification mechanism. The exponential smoothing and consistency verification mechanism includes: Before step S4 is executed, the number of consecutive time windows for which the combined matching degree has been calculated is counted. When the count reaches a preset number, the non-negative intensity vectors of the consecutive preset number of time windows are exponentially weighted and smoothed before being output. When the residual scoring factor is lower than the preset consistency threshold, steps S2 to S3 are executed again to recalibrate the response signal sequence to generate a new response signal segment. Based on the new response signal segment, the non-negative intensity vector is resolved and the residual scoring factor is generated again until the preset consistency threshold is met or the maximum number of iterations is reached.

6. The disease early warning method based on odor pattern recognition as described in claim 5, characterized in that, The recalibration includes a parameter adjustment mechanism, which includes: increasing the preset similarity threshold and expanding the order of piecewise linear regression.

7. The disease early warning method based on odor pattern recognition according to claim 1, characterized in that, Step S34, which involves filtering different candidate disease sets based on independent matching degree, includes the following steps: S341: Based on the comparison between independent matching degree and preset matching degree interval, the diseases are divided into a main set and a sub-set, and the preset intensity value expectation of all disease-related odor components is extracted. S342: Construct different initial subsets from one or more diseases in the main set, and combine them with the intensity value expectation weighted by the odor component-disease association weight matrix to generate an intensity value expectation vector; S343: With the minimum number as a constraint, independently select diseases or combinations thereof from the subset and include them in the initial subset. Iterate and update the expected intensity vector until its Euclidean distance with the non-negative intensity vector meets the preset deviation range or reaches the maximum number of inclusions. S344: Denote the initial subset that meets the preset deviation range as the candidate disease set, calculate the sum of the independent matching degree of each disease in it, and output the candidate disease set in descending order accordingly.

8. The disease early warning method based on odor pattern recognition as described in claim 1, characterized in that, The generation of the adaptive threshold in step S4 includes the following steps: S41: Set a default value for the adaptive threshold; S42: When there are more than m valid warning signals, a fixed number of time windows before the warning signal is triggered are taken as the warning period, and the combined matching degree values ​​of all time windows within it are extracted to form the total matching degree sequence. S43: Concatenate the total matching degree sequence corresponding to the most recent n warning periods into a global matching degree sequence according to the time window order, input the global matching degree sequence into the moving average model, and output the baseline threshold curve; S44: Collect all residual score factor values ​​from the most recent n warning periods to form a global residual sequence, calculate its standard deviation and mean, and use the ratio of the standard deviation to the mean as the adjustment coefficient; S45: Generate a smoothly adjusted adaptive threshold by scaling the baseline threshold curve by adjusting the coefficients.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.