Lung cancer auxiliary diagnosis method based on exhaled gas base large model

By using a method based on the exhaled gas base large model in the auxiliary diagnosis of lung cancer, VOCs characteristics and multimodal data are extracted and fused, the problem of poor generalization ability of lung cancer-related VOCs patterns in the prior art is solved, and more stable identification and personalized prediction are achieved.

CN120072276APending Publication Date: 2025-05-30CHENGDU ALIEBN SCI & TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510540884.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has poor generalization ability when identifying volatile organic compounds (VOCs) patterns associated with lung cancer, and cannot accurately identify new or unseen data.

Method used

The lung cancer-assisted diagnosis method based on the exhaled gas base model is adopted. By extracting features from high-dimensional mass spectrometry data, non-standard training is carried out to learn the cross-sample common mode of VOCs. In the training stage of the lung cancer-assisted diagnosis model, the generalized features extracted from the basic big model are input together with multimodal data, and the fusion model of pathological specific signals and clinical context information is achieved through supervised training.

Benefits of technology

It enhances the stable recognition ability of lung cancer marker combinations, improves cross-scenario generalization performance, and provides personalized lung cancer risk prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072276A_ABST
    Figure CN120072276A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of model training, and particularly relates to a lung cancer auxiliary diagnosis method based on an exhaled gas base large model. In a basic large model training stage, n features are extracted from high-dimensional mass spectrum data, and through non-standard training (learning a cross-sample generality mode of VOCs, noise interference is eliminated and deep characterization of complex metabolism association is established; then, in the lung cancer auxiliary diagnosis model training stage, generalization features extracted by the basic large model and multi-modal data are jointly input, fusion modeling of pathological specific signals and clinical context information is achieved through supervised training, and therefore on the basis that the adaptive capacity of the large model to VOCs dynamic time-varying characteristics and individual heterogeneity is reserved, the large model can be used as a model for auxiliary diagnosis of lung cancer. And the stable identification of the lung cancer marker combination is enhanced, and finally the cross-scene generalization performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of model training. More specifically, it relates to a lung cancer auxiliary diagnosis method based on an exhaled gas base large model. Background Art

[0002] Exhaled breath (also known as respiratory gas) analysis is an emerging method for disease diagnosis in recent years. Exhaled breath contains a variety of volatile organic compounds (VOCs), and the concentration changes of these compounds are closely related to various diseases, especially respiratory diseases (such as lung cancer, asthma, COPD, etc.). Compared with traditional blood or imaging examinations, exhaled breath analysis has the advantages of being non-invasive, fast, and convenient, so it shows great potential in early disease screening and health monitoring.

[0003] Since the patterns of VOCs (volatile organic compounds) in exhaled breath are very complex, due to metabolic differences among different patients, environmental factors, high-dimensional noise in mass spectrometry detection, etc., the data has high heterogeneity and non-linearity. Existing traditional methods may perform well on specific datasets, but they have poor generalization ability when encountering new and unseen data and cannot accurately identify the VOC patterns related to lung cancer. Summary of the Invention

[0004] The present invention provides a lung cancer auxiliary diagnosis method based on an exhaled gas base large model, aiming to solve the technical problem of the generalization ability of complex VOC patterns.

[0005] The lung cancer auxiliary diagnosis method based on an exhaled gas base large model includes the following steps: Basic large model training: Obtain m feature data from human exhaled gas mass spectrometry data, perform feature extraction based on the m feature data, extract n features, use the n features as the input of the basic large model, and perform non-standard training on the basic large model to obtain a trained basic large model; Lung cancer auxiliary diagnosis model training: Extract n features from human exhaled gas mass spectrometry data as the input of the trained basic large model, obtain output data based on the basic large model, fuse the output data and multi-modal data as the input of the lung cancer auxiliary diagnosis model, and use the supervised training method to train the lung cancer auxiliary diagnosis model to obtain a trained lung cancer auxiliary diagnosis model. Use the trained lung cancer auxiliary diagnosis model to output the actual prediction result to provide decision support for medical staff.

[0006] In the training stage of the basic large model of the present invention, n features are extracted from high-dimensional mass spectrometry data, and cross-sample common patterns of VOCs are learned through non-standard training (such as self-supervised pre-training) to eliminate noise interference and establish deep representations of complex metabolic associations; subsequently, in the training stage of the lung cancer auxiliary diagnosis model, the generalized features extracted by the basic large model and multi-modal data are jointly input, and supervised training is used to achieve the fusion modeling of pathological specific signals (such as lung cancer-related VOC combinations) and clinical context information (such as imaging features), so as to enhance the stable recognition of lung cancer biomarker combinations on the basis of retaining the adaptability of the large model to the dynamic time-variability and individual heterogeneity of VOCs, and finally improve the cross-scene generalization performance.

[0007] Preferably, the basic large model includes a multi-head attention mechanism and a Transformer architecture; Wherein the Transformer architecture includes an encoder and a decoder. The encoder maps the input data to a latent space representation, and the decoder is used to restore the latent space representation to the output data; and the multi-head attention mechanism is introduced into the encoder and decoder structures.

[0008] Preferably, the multi-modal data includes human exhaled gas mass spectrometry data, spectral text descriptions, patient text descriptions, calibration gas spectra, instrument status parameters, and environmental spectra.

[0009] Preferably, before training the lung cancer auxiliary diagnosis model by the supervised training method, it further includes extracting features from the human exhaled gas mass spectrometry data, extracting peak height and peak area features, and labeling the extracted features, and using the labels as the target variables for supervised learning.

[0010] Preferably, before extracting features from the human exhaled gas mass spectrometry data, it further includes generating calibration coefficients based on the calibration gas spectra, and correcting the human exhaled gas mass spectrometry data and environmental spectra based on the calibration coefficients.

[0011] Preferably, before training the lung cancer auxiliary diagnosis model by the supervised training method, the text data is vectorized, and the vectorized text data is used to describe the peak height or peak area features.

[0012] The beneficial effects of the present invention include: in the basic large model training stage, n features are extracted from high-dimensional mass spectrometry data, and cross-sample common patterns of VOCs are learned through non-standard training (such as self-supervised pre-training), noise interference is eliminated, and a deep representation of complex metabolic associations is established; subsequently, in the lung cancer auxiliary diagnosis model training stage, the generalized features extracted by the basic large model and multi-modal data are jointly input, and supervised training is used to realize the fusion modeling of pathological specific signals (such as VOCs combinations related to lung cancer) and clinical context information (such as imaging features), so as to enhance the stable recognition of lung cancer biomarker combinations while retaining the adaptability of the large model to the dynamic time-variability and individual heterogeneity of VOCs, and finally improve the cross-scene generalization performance; secondly, through the deep learning of a large amount of exhaled gas data, the model can accurately identify features related to lung cancer, and combined with individualized feature tags, it can provide personalized lung cancer risk prediction. Brief Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 It is a schematic diagram of the specific implementation logic provided for the embodiment of the present invention. Detailed Embodiments

[0015] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the following further details the present application with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0016] A lung cancer auxiliary diagnosis method based on an exhaled gas base large model includes the following steps: Basic large model training: Obtain m feature data from human exhaled gas mass spectrometry data, perform feature extraction based on the m feature data, extract n features, use the n features as the input of the basic large model, and perform non-standard training on the basic large model to obtain a trained basic large model; the basic large model includes a multi-head attention mechanism and a Transformer architecture; For example, 25,088 features are obtained from human exhaled gas mass spectrometry data (representing different data points, each data point representing a different m / z value and corresponding intensity). 6,000 of the most representative features are selected from the 25,088 features and used as the input to the base large model. Without tagging, they are directly input into the base large model for training to obtain a trained base large model. The 6,000 most representative features are selected based on task relevance. For example, statistical screening, feature importance ranking, or dimensionality reduction methods can all be used to achieve this.

[0017] The Transformer architecture includes an encoder and a decoder. The encoder maps the input data to a latent space representation, and the decoder is used to restore the latent space representation to the output data. The multi-head attention mechanism is introduced into the encoder and decoder structures.

[0018] The structure of an exemplary base large model is as follows: Encoder part (data encoding): Input embedding: The input words are converted into a fixed-length vector representation through an embedding layer. It contains some position encoding information to represent the position of the words in the sequence.

[0019] Multi-head self-attention mechanism: Each element (word vector) of the input sequence is processed through the multi-head self-attention mechanism. The multi-head attention mechanism calculates the relationships (attention scores) between each word in the input sequence. Through multiple attention heads, the model can parallelly capture different levels of features or relationships.

[0020] Each attention head calculates a weighted sum, where the weights are determined by the influence of other words on the current word. The outputs of these weighted sums are concatenated and integrated through a linear transformation.

[0021] Feed-forward network: The output of each self-attention module passes through a feed-forward neural network (such as two fully connected layers) and is processed through a non-linear activation function to enhance the model's expressive power.

[0022] Residual connection and layer normalization: The output of each layer is added to the input through a residual connection and then undergoes layer normalization processing to alleviate the vanishing gradient problem in deep neural networks and accelerate training.

[0023] Output of the Encoder: The final output of the Encoder is a high-dimensional representation of the context information of the input sequence. This output is a latent representation containing the relationships between various elements within the sequence for further decoding by the Decoder.

[0024] Decoder section: Decoder input: The Decoder receives the output of the Encoder and the output of the previous step of the Decoder (or during training, the prefix of the target sequence given as input); the goal of the Decoder is to transform the representation in the latent space (the output of the Encoder) into the target sequence.

[0025] Masked multi-head self-attention mechanism: The first step in the Decoder is to calculate the attention scores through Masked Attention, which ensures that at each time step, the Decoder can only utilize the information of previous words, ensuring that the model does not see future words in advance when generating the sequence.

[0026] Encoder-Decoder attention mechanism: In the next step of the Decoder, the attention scores between the Encoder output and the Decoder input are calculated through the Encoder-Decoder Attention layer; allowing the Decoder to obtain the global context information of the input sequence, thereby helping to generate accurate outputs.

[0027] Feed-forward network and residual connection: Similar to the Encoder section, the output of each layer in the Decoder also passes through a feed-forward neural network, and residual connection and layer normalization are performed.

[0028] Generating the output: The output of the Decoder passes through a linear layer and uses the Softmax function to generate a probability distribution.

[0029] It should be noted that the specific structure of the above model provided in this embodiment is only exemplary. The main purpose of the present invention is to provide a method for training a model on how to train a model with strong generalization ability. Therefore, on the basis of changing the model structure, using the training method described in the present invention to obtain the final model also belongs to the protection scope of this application.

[0030] Training of the lung cancer auxiliary diagnosis model: Extract n features from the human exhaled gas mass spectrometry data as the input of the trained basic large model, obtain the output data based on the basic large model, fuse the output data and the multi-modal data as the input of the lung cancer auxiliary diagnosis model, and train the lung cancer auxiliary diagnosis model using the supervised training method to obtain the trained lung cancer auxiliary diagnosis model, and use the trained lung cancer auxiliary diagnosis model to output the actual prediction result to provide decision support for medical staff.

[0031] The multi-modal data includes human exhaled gas mass spectrometry data, spectral text description, patient text description, calibration gas spectrum, instrument status parameters, and environmental spectrum.

[0032] Before training the lung cancer auxiliary diagnosis model using the supervised training method, it also includes extracting features from the human exhaled gas mass spectrometry data, extracting peak height and peak area features, and labeling the extracted features, using the labels as the target variables for supervised learning; Exemplarily, peak height and peak area features are extracted to reflect the concentration and distribution of specific substances in the exhaled gas. If 212 features are extracted, then the 212 extracted features are labeled as the target variables for supervised learning.

[0033] Before extracting features from the human exhaled gas mass spectrometry data, it also includes generating a correction coefficient based on the calibration gas spectrum, and correcting the human exhaled gas mass spectrometry data and the environmental spectrum based on the correction coefficient; Exemplary technical solutions are as follows: Solution 1: Using the relationship between the data of the calibration gas spectrum and the concentration of the target gas, establish a correction coefficient model through linear regression; specifically as follows: Collect the spectral data of the calibration gas through a gas analyzer and record the concentration of each gas; Use the calibration gas data for linear regression analysis to obtain the regression coefficients of each gas component. Then collect the spectral data of the human exhaled gas sample, and use the regression coefficients to correct the spectral data of the exhaled gas to calculate the concentration of each gas component.

[0034] Solution 2: Collect the spectra of standard gases with different concentrations, use the polynomial regression method to fit the relationship between the calibration gas spectrum data and the corresponding concentrations to obtain accurate correction coefficients, and input the exhaled gas spectrum data into the polynomial regression model to obtain accurate concentration estimates.

[0035] Solution 3: Reduce the high-dimensional data in the calibration gas spectrum to a low-dimensional space through PCA, and extract the principal components that can best represent the data characteristics; Perform regression modeling on the principal component data and the calibration gas concentration to obtain a correction coefficient; Reduce the dimension of the exhaled gas spectrum through PCA and correct it with the regression model to obtain the corresponding gas concentration.

[0036] The purpose of providing 3 correction solutions in this embodiment is to illustrate that there are various methods for correcting the human exhaled gas mass spectrometry data through the calibration gas spectrum, which is not limited to the 3 technical solutions provided above. Therefore, it should not be understood as a limitation to the present invention. Within the idea of the present invention, equivalent transformations of the existing technology are also within the protection scope of this application.

[0037] Secondly, the correction method for the environmental spectrum is the same as that for the human exhaled gas mass spectrometry data.

[0038] Before training the lung cancer auxiliary diagnosis model using the supervised training method, the text data is vectorized, and the vectorized text data is used to describe the peak height or peak area feature.

[0039] It should be noted that in this embodiment, the lung cancer auxiliary diagnosis model can adopt the trained basic large model or the Titans model improved based on Transformer (the core of the Titans architecture lies in a brand-new neural long-term memory module, aiming to dynamically learn and store historical information); that is, the lung cancer auxiliary diagnosis model described in this embodiment can be trained according to the actual situation, and the two model structures given above are not limitations to the present invention.

[0040] An exemplary structure in which the output data and multi-modal data are fused and used as the input of the lung cancer auxiliary diagnosis model is as follows: Adopt the Star Aggregate-Redistribute Module processing architecture. Divided by channels, each channel has a data sequence on the time axis. This architecture contains multiple Star Aggregate-Redistribute Modules, and each Star Aggregate-Redistribute Module consists of the following parts: Series data block: The input time series data enters the series data block for preliminary processing; Pool: The input series data block is pooled (Pooling), aiming to reduce the data size while retaining important features; the pooled data will be sent to a multi-layer perceptron (MLP) for further processing.

[0041] MLP (multi-layer perceptron): High-level feature extraction is performed on the pooled data; the extracted features will be used for the Aggregate operation; Aggregate: Through a predetermined mathematical operation or mechanism (such as summation, averaging, etc.), the extracted features are summarized; the aggregated features will be used as a new set of comprehensive features; Repeat&Concat (repeat and concatenate): The data after the repeat and concatenate operation enters the core module, and the core module redistributes this data to different channels.

[0042] Residual: The redistributed data will be connected with the original input data by residual connection, which is beneficial to information transmission and preventing gradient disappearance.

[0043] Through this multi-module and multi-step processing mechanism, it is possible to effectively extract, aggregate, enhance, and reallocate features from time series data, thereby improving the data representation ability and the learning effect of the model.

[0044] In the basic large model training stage of the present invention, n features are extracted from high-dimensional mass spectrometry data, and the cross-sample common patterns of VOCs (volatile organic compounds) are learned through non-standard training (such as self-supervised pre-training) to eliminate noise interference and establish a deep representation of complex metabolic associations; subsequently, in the lung cancer auxiliary diagnosis model training stage, the generalized features extracted by the basic large model are jointly input with multi-modal data, and the fusion modeling of pathological specific signals (such as lung cancer-related VOC combinations) and clinical context information (such as imaging features) is achieved through supervised training, so as to enhance the stable recognition of lung cancer biomarker combinations on the basis of retaining the adaptability of the large model to the dynamic time-variability and individual heterogeneity of VOCs, and finally improve the cross-scene generalization performance.

[0045] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A lung cancer auxiliary diagnosis method based on a large exhaled gas base model, characterized in that: The following steps are involved: Basic large model training: m feature data are obtained from human exhaled gas spectrum data, and feature extraction is performed based on the m feature data to extract n features, and the n features are used as inputs of the basic large model, and non-standard training is performed on the basic large model to obtain a trained basic large model; Training of lung cancer auxiliary diagnosis model: extract n features based on human exhaled gas spectrum data as the input of the trained basic large model, obtain output data based on the basic large model, fuse the output data and multimodal data as the input of the lung cancer auxiliary diagnosis model, train the lung cancer auxiliary diagnosis model using supervised training, obtain the trained lung cancer auxiliary diagnosis model, use the trained lung cancer auxiliary diagnosis model to output the actual prediction results, and provide decision support for medical staff.

2. The lung cancer auxiliary diagnosis method based on the exhaled gas base large model according to claim 1 is characterized in that: The basic large model includes a multi-head attention mechanism and a Transformer architecture; The Transformer architecture includes an encoder and a decoder, wherein the encoder maps input data to a latent space representation, and the decoder is used to restore the latent space representation to output data; And introduce the multi-head attention mechanism into the encoder and decoder structure.

3. The lung cancer auxiliary diagnosis method based on the exhaled gas base large model according to claim 1, characterized in that: The multimodal data includes human exhaled gas spectrum data, spectrum text description, patient text description, standard gas spectrum, instrument status parameters and environmental spectrum.

4. The lung cancer auxiliary diagnosis method based on the exhaled gas base large model according to claim 3 is characterized in that: Before training the lung cancer auxiliary diagnosis model in a supervised training manner, the method also includes extracting features from the human exhaled gas spectrum data, extracting peak height and peak area features, and labeling the extracted features, using the labels as target variables for supervised learning.

5. The lung cancer auxiliary diagnosis method based on the exhaled gas base large model according to claim 4 is characterized in that: Before extracting features from the human exhaled gas spectrum data, the method further includes generating a correction coefficient based on the standard gas spectrum, and correcting the human exhaled gas spectrum data and the environmental spectrum based on the correction coefficient.

6. The lung cancer auxiliary diagnosis method based on the exhaled gas base large model according to claim 4, characterized in that: Before training the lung cancer auxiliary diagnosis model in a supervised training manner, the text data is vectorized, and the vectorized text data is used to describe the peak height or peak area features.

Citation Information

Patent Citations

  • Pulmonary nodule intelligent grading method and system based on multi-modal feature fusion

    CN116883768A

  • Lung cancer electronic nose data classification method based on adversarial training and multi-scale attention

    CN118568565A

  • Method for constructing lung cancer screening model based on PTR-TOF-MS

    CN118983078A

  • Establishment method of pulmonary nodule malignant risk prediction model integrating iconography indexes and expired gas VOCs (Volatile Organic Compounds) data

    CN119108096A

  • Portable detection device and method for oral expired gas

    CN120044109A