Multi-mode-based vocal cord cancer lesion feature extraction method and analysis system

Through multimodal data fusion and deep learning technology, sound, image and structured data are integrated and high-level features are extracted, which solves the problems of insufficient utilization of multimodal data and low accuracy in feature extraction in the existing technology, and achieves more comprehensive and accurate extraction of vocal cord cancer lesions and more efficient diagnosis.

CN120048491APending Publication Date: 2025-05-27JILIN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510201384.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing vocal cord cancer lesion feature extraction methods are insufficient to utilize multimodal data, have low accuracy in feature extraction and poor interpretability of the model.

Method used

Multimodal data fusion technology is adopted to integrate sound, image and structured data through multi-scale feature stitching and modal embedding. Use the Transformer encoder to extract high-level features and build a network model through deep learning architecture, including the embedding layer, the encoder layer and the output layer, to complete the feature extraction and classification tasks.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of feature extraction, provides non-invasive and efficient analytical methods, enhances the interpretability of the model, and improves the efficiency of early lesion feature extraction and diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048491A_ABST
    Figure CN120048491A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal-based vocal cord cancerous lesion feature extraction method and analysis system. The method comprises the following steps: step 1, acquiring biochemical indexes, audio and laryngoscope pictures of patients with polyps, precancerous lesions and cancerous lesions in otolaryngological departments of cooperative hospitals in several years as multi-modal data; 2, image processing; 3, constructing a network model based on a deep learning architecture; and 4, performing model preheating by adopting a cross entropy loss function and a self-adaptive optimizer. The lesion feature extraction and analysis system comprises a data acquisition and processing module, a model training and evaluation module and a classification result output module; wherein the data acquisition and processing module executes the steps of data acquisition and data processing, the model training and evaluation module constructs a model based on a deep learning architecture, and the method has the beneficial effects that the omission and error rate of feature extraction is effectively reduced, the success rate of early feature discovery is improved, and the stability and reliability of a feature extraction result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for extracting vocal cord cancer lesion features and an analysis system, and particularly to a method for extracting vocal cord cancer lesion features and an analysis system based on multi-modal data. Background Art

[0002] Currently, in the field of vocal cord cancer lesion analysis, traditional methods mainly rely on doctors' experience and single-modal data. For example, only the morphological changes of laryngeal lesions are observed through laryngoscope pictures, or only blood routine and biochemical indexes are analyzed simply, and the utilization of voice features is also relatively limited. These methods have obvious deficiencies in data dimension and information integration, and it is difficult to comprehensively and accurately extract the features of vocal cord cancer lesions, especially in the analysis of early lesions, the effect is relatively limited.

[0003] In addition, although the traditional pathological analysis method is regarded as the gold standard, its invasiveness limits the acceptance of some patients. Some patients are difficult to tolerate such examinations due to psychological factors or health conditions, which to a certain extent affects the efficiency of early extraction and analysis of vocal cord cancer lesion features, and also highlights the necessity of exploring non-invasive or minimally invasive analysis methods.

[0004] Currently, there have been some studies on using data analysis to assist in extracting vocal cord cancer lesion features, but most of them only process single-modal data and lack effective integration and in-depth analysis of multi-modal data. For example, some studies only focus on the processing of image data and do not fully combine information such as clinical biochemical indexes and voice features, and cannot fully explore the potential associations between different modal data, resulting in insufficient comprehensiveness and reliability of feature extraction.

[0005] In terms of data processing, the existing technologies have not effectively solved key problems such as quality control, feature alignment, and fusion of multi-modal data, which affects the performance of the model. In addition, the existing methods also have deficiencies in model interpretability, and it is difficult to clearly explain the basis of model decision-making and the importance of each modal data in feature extraction, which limits their credibility and practicality in actual applications.

[0006] In contrast, the method proposed by the present invention has the following advantages:

[0007] Multi-modal data fusion ability: Through multi-scale feature concatenation and modality embedding technologies, it effectively integrates voice, image, and structured data, fully explores the associations between different modal data, and significantly improves the comprehensiveness and accuracy of feature extraction.

[0008] High-level feature extraction ability: Using a Transformer encoder to extract high-level features, which can capture complex patterns and potential laws in the data, and is superior to the shallow analysis of single-modal data by traditional methods.

[0009] Non-invasiveness and high efficiency: The analysis method based on multi-modal data does not rely on invasive examinations, can provide a more friendly analysis experience for patients, and improve the efficiency of early lesion feature extraction at the same time.

[0010] Model interpretability: Through a clear modality embedding and feature fusion mechanism, it can better explain the model decision-making process and enhance the credibility and practicality of the results. Summary of the Invention

[0011] The purpose of the present invention is to provide a multi-modal based vocal cord cancer lesion feature extraction method and analysis system to solve many problems such as insufficient utilization of multi-modal data, low accuracy of feature extraction, and poor model interpretability in existing vocal cord cancer lesion feature extraction methods.

[0012] The multi-modal based vocal cord cancer lesion feature extraction method provided by the present invention includes the following steps:

[0013] First step: Obtain biochemical indicators, audio, and laryngoscope images of patients with polyps, precancerous lesions, and cancer in the otolaryngology department of cooperative hospitals over the years as multi-modal data, pass through the ethical review of the cooperative hospitals, obtain clinical samples by simple sampling and make clear annotations, refer to the knowledge ontology in the field of vocal cord cancer, construct data resources, and use the Cartesian product method, and fill in with "0" if the number of digits is insufficient;

[0014] Second step: Image processing: Force all data with inconsistent bit depths into RGB format, scale and center crop the data resolution, normalize the image pixel values, and divide the image into 16*16 patches, and project linearly into the embedding space; Audio processing: Adopt a centering strategy for audio, extract Mel-frequency cepstral coefficients, and perform masking operations, random noise processing, rolling processing, and normalization processing, and divide the Mel-frequency cepstral coefficients into 16*16 patches, and project linearly into the embedding space, Biochemical index processing: Perform multiple imputations of chained equations on biochemical data, and perform normalization and feature embedding processing on all data;

[0015] Third step: Build a network model based on a deep learning architecture. The network model consists of an embedding layer, an encoder layer, and an output layer. The embedding layer fuses multi-modal features and creates position encoding. The multi-head self-attention mechanism of the encoder layer discovers semantic associations between different modal data, and the feed-forward neural network enhances the expression ability of the model. The output layer completes the classification task;

[0016] Step 4: Use the cross-entropy loss function and the adaptive optimizer to warm up the model. Evaluate the model performance through accuracy, average precision, and AUC. Use gradient-weighted class activation mapping, attention weights, and Shap values to achieve model interpretability analysis. Make full use of multi-modal data to deeply mine the internal connections, which are used as clues and bases for early screening and vocal cord lesion classification.

[0017] The specific steps of the first step are as follows:

[0018] Step 1: The data comes from the real data of clinical polyps, precancerous lesions, and cancer patients in the otolaryngology department of a cooperative hospital in recent years. The data must pass the ethical review of the cooperative hospital, and the data source is real and reliable. Taking the clinical patients as the population, simple sampling is used to obtain clinical sample data, which are divided into three categories: polyps, precancerous lesions, and cancer, and are clearly labeled.

[0019] Step 2: Follow the clinical interpretation and refer to the knowledge ontology in the field of vocal cord cancer to construct data resources. Each data resource consists of 1 biochemical index, 3 images, and 2 audio segments. Using the Cartesian product method, the above data resources are integrated into a total of 6 data of 1*2*3, that is, each data only contains one biochemical index, one image, and one audio segment. If the number of digits is insufficient, it is filled with "0" to ensure the integrity and standardization of the data.

[0020] The specific steps of the second step are as follows:

[0021] Step 1: Image processing, specifically as follows:

[0022] (1) Bit-depth inconsistency processing: In the image data, some data have a bit-depth of 32, that is, 4 channels; some data have a bit-depth of 24, that is, 3 channels. To ensure the consistency of the bit-depth, the transparency channel of the 4-channel data is discarded, and all data are forcibly converted to the RGB format.

[0023] (2) Scaling and central cropping processing: To fix the data length, the image is scaled to 256*256. It is observed that the data border is black useless information, and the key information is concentrated in the center of the image. Combining with the input of the Transformer, the data is centrally cropped, and the cropped size is 224*224.

[0024] (3) Standardization processing: The pixel value range of the image is fixed at 0-255. Each pixel point is divided by 255 for normalization operation, and the range of the normalized pixel value is a floating point number in [0.0, 1.0].

[0025] (4), Patch Embedding: The image is divided into 16*16 patches, which are projected into the embedding space linearly. The specific method is as follows: Use a 16*16 convolutional kernel with a stride of 16, and the embedding feature is 768 dimensions. Use 2D convolution + linear projection, and the image changes from 1*3*224*224 to a feature of 1*196*768;

[0026] Step 2, Audio Processing, is as follows:

[0027] (1), Centering Processing: The read audio is a one-dimensional array. Adopt the centering strategy to eliminate the mean shift. The formula is as follows:

[0028] X_centered = X - μ

[0029] X is the original data matrix, each row is a sample, each column is a feature, and μ is the mean vector of each feature;

[0030] (2), Extract Mel-Frequency Cepstral Coefficients MFCC: Based on the non-linear perception of the human auditory system, through frame-by-frame processing, applying a Hanning window, and discrete Fourier transform of the audio signal, the time-domain signal is converted into a frequency-domain signal, and 128 frequency bands are divided. Based on a 10-second length benchmark, the extracted time-domain dimension is approximately 1024. Truncate the part exceeding 1024, pad 0 for the part less than 1024, and fix the time-domain length at 1024 to obtain a feature tensor of 1*1024*128;

[0031] (3), Masking Operation: Usually, the mask is 10%-20% in the frequency domain and time domain. Calculate that the frequency-domain mask is 24 and the time-domain mask is 256. Randomly occlude the Mel-Frequency Cepstral Coefficients MFCC in the time-domain and frequency-domain dimensions to improve the generalization ability of the model;

[0032] (4), Random Noise Processing: Add random noise in both the time domain and frequency domain, with a range of 0-1. The purpose is to enhance the robustness of the model and improve the generalization ability;

[0033] (5), Rolling Processing: Randomly roll forward or backward in the time domain with a unit of 10 to improve data diversity, extract the features of the data from different angles, and achieve data augmentation;

[0034] (6), Normalization Processing: Normalize the mean and standard deviation of the Mel-Frequency Cepstral Coefficients MFCC. The formula is as follows:

[0035] fbank = (fbank - fbank.mean()) / fbank.std();

[0036] Fbank represents the MFCC matrix, mean calculates the mean, and std calculates the standard deviation;

[0037] (7) Patch Embedding: The Mel Frequency Cepstral Coefficients (MFCC) are divided into 16*16 patches, which are projected into the embedding space through linear projection. Specifically, a 16*16 convolutional kernel is used, with a stride of 10 in both the time domain and the frequency domain, overlapping sampling with an overlap coefficient of 6, and an embedding feature of 768. 2D convolution + linear projection is adopted, and the Mel Frequency Cepstral Coefficients (MFCC) are transformed from a feature of 1*1024*128 to a feature of 1*1212*768.

[0038] Step 3: Biochemical Index Processing, which is specifically as follows:

[0039] (1) Multiple Imputation by Chained Equations: Multiple Imputation by Chained Equations is adopted to construct a MICE model based on all data for data imputation.

[0040] (2) Normalization Processing: The maximum and minimum values are calculated respectively for all data in the column direction, and normalization is performed using the maximum and minimum values in the row direction. The formula is (x - min(x)) / (max(x) - min(x)).

[0041] (3) Feature Embedding Processing: A fully connected linear layer is used to map 51 biochemical indexes to a 768-dimensional space, which is the same dimension as images and audio, and the shape changes from 1*51 to 1*1*768.

[0042] The specific steps of the third step are as follows:

[0043] The network model consists of an embedding layer, an encoder layer, and an output layer, which are specifically as follows:

[0044] Step 1: Embedding Layer: After feature embedding of all modal data, feature concatenation is performed in the first dimension. The images of 1*196*768, the audio of 1*1212*768, and the biochemical indexes of 1*1*768 are concatenated to 1*1409*768; additionally, a classification with a shape of 1*1*768 is added, and the final shape is 1*1410*768. In order to enable the model to understand the data order, position encoding is created and embedded into the data.

[0045] Step 2: Encoder Layer, which is specifically as follows:

[0046] (1) Multi-Head Self-Attention Mechanism: An Encoder layer consisting of 12 Block blocks is designed. Each Block block uses the multi-head self-attention mechanism to calculate the correlation between each position in the sequence and other positions, allowing the model to focus on different parts of the sequence. In a sequence containing text and image information, the model can determine the degree of association between a certain word in the text and a certain region in the image through the self-attention mechanism.

[0047] (2) Feedforward neural network: Perform a non-linear transformation on the features after self-attention processing to further enhance the model's expressive ability. Process the features at each position independently, including two fully connected layers, and use GELU + LayerNorm for non-linear activation and layer normalization in each layer;

[0048] Step 3, Output layer: For classification tasks, use a fully connected layer to convert the 0th dimension of the features output by the Encoder layer into class probabilities.

[0049] The specific steps of the fourth step are as follows:

[0050] Step 1, Loss function: Use binary cross-entropy for binary classification and cross-entropy for multi-classification;

[0051] Step 2, Optimizer: Use an adaptive optimizer, with an initial learning rate of 1e-5, an L2-norm regularization weight decay value of 5e-7, an exponential decay rate of 0.95 for the first moment estimate, and an exponential decay rate of 0.999 for the second moment estimate;

[0052] Step 3, Model warm-up: Design a 200-batch model warm-up. Within the first 200 batches, dynamically adjust the learning rate to achieve the purpose of stable training;

[0053] Step 4, Evaluation metrics: Calculate evaluation metrics for each round, including accuracy, average precision, and AUC;

[0054] Step 5, Model interpretability, specifically as follows:

[0055] (1) Calculate the average gradient of each modality using gradient-weighted class activation mapping, and interpret the importance of each modality according to the average gradient score;

[0056] (2) Calculate the attention weights and obtain the weight scores through the average weights of each modality;

[0057] (3) Analyze the contribution rate of each modality to the classification result through the average shap value of each modality.

[0058] The multi-modal vocal cord cancer lesion feature extraction and analysis system provided by the present invention includes a data acquisition and processing module, a model training and evaluation module, and a classification result output module; among them, the data acquisition and processing module performs the steps of data acquisition and data processing to obtain and process multi-modal data; the model training and evaluation module constructs a model based on a deep learning architecture and performs model training, evaluation, and interpretability analysis, including operations such as setting the loss function, optimizer, model warm-up, calculating evaluation metrics, and obtaining modality importance; the classification result output module receives the output result of the model training and evaluation module and outputs and displays the final vocal cord lesion classification result.

[0059] Advantages of the present invention:

[0060] (1) Improve the comprehensiveness and accuracy of feature extraction: Traditional single-modal methods only focus on a certain type of data, resulting in limited obtained lesion features. In contrast, the present invention can fully explore the correlation information between different modal data by integrating multi-modal data, and extract the lesion features of vocal cord cancer from multiple dimensions. Compared with traditional methods, the present invention can capture the features of vocal cord cancer more comprehensively and accurately, effectively reducing the omission and error rates of feature extraction, and significantly improving the success rate of early feature discovery.

[0061] (2) Enhance the adaptability of the model: Traditional methods in the past were relatively single in data processing and lacked the ability to adapt to changes in data scenarios. In the data processing process of the present invention, various enhancement operations are performed on image and audio data, such as masking, adding noise, rolling processing, etc., and biochemical indicators are reasonably processed at the same time. This enables the model to learn richer feature patterns, greatly enhancing the model's adaptability in different data scenarios and improving the stability and reliability of feature extraction results.

[0062] (3) Endow the model with partial interpretability: Most traditional methods are "black box" models and it is difficult to explain their decision-making processes. The present invention uses methods such as Grad-Cam to calculate the modal average gradient, obtain attention-weights weight information, and the average shap value of each modality, which can reveal the decision-making logic of the model to a considerable extent and illustrate the importance of each modality data in feature extraction. This interpretability is a new ability that traditional methods do not possess, which helps professionals understand and trust the feature extraction results and provides a strong basis for further clinical research. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic structural diagram of the basic technical route described in the present invention.

[0064] Figure 2 It is a schematic structural diagram of the patch embedding processing described in the present invention.

[0065] Figure 3 It is a schematic structural diagram of the voice patch embedding processing described in the present invention.

[0066] Figure 4 It is a schematic structural diagram of the model encoder layer described in the present invention.

[0067] Figure 5 It is a schematic structural diagram of the overall model described in the present invention.

[0068] Figure 6 It is a schematic flow diagram of the method for extracting vocal cord cancer lesion features based on multi-modal described in the present invention. Detailed implementation mode

[0069] Please refer to Figures 1 to 6 as shown in

[0070] The method for extracting vocal cord cancer lesion characteristics based on multi-modal provided by the present invention includes the following steps:

[0071] The first step: Obtain the biochemical indexes, audio and laryngoscope pictures of patients with polyps, precancerous lesions and canceration in the otolaryngology department of cooperative hospitals in recent years as multi-modal data, and pass the ethical review of the cooperative hospitals. Obtain clinical samples by simple sampling and make clear annotations. Refer to the knowledge ontology in the field of vocal cord cancer, construct data resources, and use the Cartesian product method, and fill in the insufficient digits with "0";

[0072] The second step: Image processing: Force all data with inconsistent bit depths into RGB format, scale and center crop the data resolution, normalize the image pixel values, and divide the image into 16*16 patches, and project linearly into the embedding space; Audio processing: Adopt a centering strategy for audio, extract Mel frequency cepstral coefficients, and perform masking operation, random noise processing, rolling processing, normalization processing, and divide the Mel frequency cepstral coefficients into 16*16 patches, and project linearly into the embedding space, Biochemical index processing: Perform multiple imputation of chained equations on biochemical data, and perform normalization and feature embedding processing on all data;

[0073] The third step: Construct a network model based on the deep learning architecture. The network model consists of an embedding layer, an encoder layer and an output layer. The embedding layer fuses multi-modal features and creates position encoding. The multi-head self-attention mechanism of the encoder layer discovers the semantic associations between different modal data, and the feed-forward neural network enhances the expression ability of the model. The output layer completes the classification task;

[0074] The fourth step: Use the cross-entropy loss function and the adaptive optimizer to preheat the model, evaluate the model performance through accuracy, average precision and AUC, and use gradient-weighted class activation mapping, attention weights and Shap values to realize the interpretability analysis of the model, make full use of multi-modal data, deeply explore the internal connections, and use them as clues and bases for early screening and vocal cord lesion classification.

[0075] The specific steps of the first step are as follows:

[0076] Step 1: The data comes from the real data of clinical polyps, precancerous lesions and canceration patients in the otolaryngology department of cooperative hospitals in recent years. The data must pass the ethical review of the cooperative hospitals. The data source is true and reliable. Taking the clinical patients as the population, clinical sample data is obtained by simple sampling, divided into three categories: polyps, precancerous lesions and canceration, and clear annotations are made;

[0077] Step 2: Follow the clinical interpretation, refer to the knowledge ontology in the field of vocal cord cancer, and construct data resources. Each data resource consists of 1 biochemical index, 3 images, and 2 audio segments. Using the Cartesian product method, integrate the above data resources into a total of 6 data with 1*2*3, that is, each data only contains one biochemical index, one image, and one audio segment. Pad with "0" if the number of digits is insufficient to ensure the integrity and standardization of the data.

[0078] The specific steps of the second step are as follows:

[0079] Step 1: Image processing, specifically as follows:

[0080] (1) Bit-depth inconsistency processing: In the image data, some data have a bit-depth of 32, that is, 4 channels; some data have a bit-depth of 24, that is, 3 channels. To ensure the consistency of the bit-depth, discard the transparency channel of the 4-channel data and force all data to the RGB format.

[0081] (2) Scaling and central cropping processing: To fix the data length, scale the image to 256*256. Observe that the data border is black useless information, and the key information is concentrated in the center of the image. Combine the input of the Transformer and perform central cropping on the data. The cropped size is 224*224.

[0082] (3) Normalization processing: The pixel value range of the image is fixed at 0-255. Divide each pixel by 255 for normalization operation. After normalization, the pixel value range is a floating point number in [0.0, 1.0].

[0083] (4) Patch embedding: Divide the image into 16*16 patches and project them into the embedding space linearly. The specific method is as follows: Use a 16*16 convolutional kernel with a stride of 16, and the embedding feature is 768 dimensions. Use 2D convolution + linear projection, and the image changes from 1*3*224*224 to 1*196*768 features.

[0084] Step 2: Audio processing, specifically as follows:

[0085] (1) Centering processing: The read audio is a one-dimensional array. Adopt the centering strategy to eliminate the mean shift. The formula is as follows:

[0086] X_centered = X - μ

[0087] X is the original data matrix, each row is a sample, each column is a feature, and μ is the mean vector of each feature.

[0088] (2) Extract Mel Frequency Cepstral Coefficients (MFCC): Based on the non - linear perception of the human auditory system, through frame - splitting, Hanning windowing, and discrete Fourier transform of the audio signal, the time - domain signal is converted into a frequency - domain signal. 128 frequency bands are divided. With a 10 - second length as the benchmark, the extracted time - domain dimension is approximately 1024. For the part exceeding 1024, it is truncated, and for the part less than 1024, it is padded with 0s to fix the time - domain length at 1024, obtaining a feature tensor of 1*1024*128;

[0089] (3) Masking operation: Usually, the mask is 10% - 20% in both the frequency domain and the time domain. It is calculated that the frequency - domain mask is 24 and the time - domain mask is 256. Random occlusion is performed on the Mel Frequency Cepstral Coefficients (MFCC) in the time - domain and frequency - domain dimensions to improve the generalization ability of the model;

[0090] (4) Random noise processing: Random noise is added in both the time - domain and frequency - domain dimensions, with a range of 0 - 1. The purpose is to enhance the robustness of the model and improve the generalization ability;

[0091] (5) Rolling processing: Random forward or backward rolling with a unit of 10 is performed in the time - domain dimension to increase data diversity, extract data features from different angles, and achieve data augmentation;

[0092] (6) Standardization processing: Standardization is performed on the mean and standard deviation of the Mel Frequency Cepstral Coefficients (MFCC). The formula is as follows:

[0093] fbank=(fbank - fbank.mean()) / fbank.std();

[0094] Fbank represents the MFCC matrix, mean calculates the mean, and std calculates the standard deviation;

[0095] (7) Patch embedding: The Mel Frequency Cepstral Coefficients (MFCC) are divided into 16*16 patches and projected into the embedding space through linear projection. Specifically: A 16*16 convolutional kernel is used, with a stride of 10 in both the time - domain and frequency - domain, overlapping sampling, an overlapping coefficient of 6, and an embedding feature of 768. 2D convolution + linear projection is adopted, and the Mel Frequency Cepstral Coefficients (MFCC) change from a feature of 1*1024*128 to a feature of 1*1212*768;

[0096] Step 3: Biochemical index processing, specifically as follows:

[0097] (1) Multiple imputation by chained equations: Multiple imputation by chained equations is adopted. An MICE model is constructed based on all data to impute the data;

[0098] (2) Normalization: Calculate the maximum and minimum values of all data in the column direction respectively, and use the maximum and minimum values for normalization in the row direction. The formula is (x - min(x)) / (max(x) - min(x));

[0099] (3) Feature embedding processing: Use a fully connected linear layer to map 51 biochemical indicators to a 768-dimensional space, keeping the same dimension as images and audio. The shape changes from 1 * 51 to 1 * 1 * 768.

[0100] The specific steps of the third step are as follows:

[0101] The network model consists of an embedding layer, an encoder layer, and an output layer, specifically as follows:

[0102] Step 1, Embedding layer: After performing feature embedding on all modal data, perform feature concatenation in the first dimension. The images 1 * 196 * 768, audio 1 * 1212 * 768, and biochemical indicators 1 * 1 * 768 are concatenated to 1 * 1409 * 768; additionally, add a classification with a shape of 1 * 1 * 768, and the final shape is 1 * 1410 * 768. To enable the model to understand the data order, create position encoding and embed it into the data;

[0103] Step 2, Encoder layer, specifically as follows:

[0104] (1) Multi-head self-attention mechanism: Design an Encoder layer consisting of 12 Block blocks. Each Block block uses the multi-head self-attention mechanism to calculate the correlation between each position in the sequence and other positions, allowing the model to focus on different parts of the sequence. In a sequence containing text and image information, the model can determine the degree of association between a certain word in the text and a certain region in the image through the self-attention mechanism;

[0105] (2) Feed-forward neural network: Perform a non-linear transformation on the features after self-attention processing to further enhance the model's expressive ability. Process the features of each position independently, including two fully connected layers, and use GELU + LayerNorm for non-linear activation and layer normalization in each layer;

[0106] Step 3, Output layer: For the classification task, use a fully connected layer to convert the 0th dimension of the features output by the Encoder layer into class probabilities.

[0107] The specific steps of the fourth step are as follows:

[0108] Step 1, Loss function: Use binary cross-entropy for binary classification and cross-entropy for multi-class classification;

[0109] Step 2, Optimizer: The optimizer uses an adaptive optimizer with an initial learning rate of 1e-5, an L2 regularization weight decay value of 5e-7, an exponential decay rate of 0.95 for the first moment estimate, and an exponential decay rate of 0.999 for the second moment estimate;

[0110] Step 3, Model Warm-up: Design 200 batches of model warm-up. Within the first 200 batches, the learning rate is dynamically adjusted to achieve the purpose of stable training;

[0111] Step 4, Evaluation Metrics: Calculate evaluation metrics for each round, including accuracy, average precision, and AUC;

[0112] Step 5, Model Interpretability, specifically as follows:

[0113] (1). Calculate the average gradient of each modality using gradient-weighted class activation mapping, and interpret the importance of each modality based on the average gradient score;

[0114] (2). Calculate the attention weights and obtain the weight scores through the average weights of each modality;

[0115] (3). Analyze the contribution rate of each modality to the classification result through the average shap value of each modality.

[0116] The multi-modal vocal cord cancer lesion feature extraction and analysis system provided by the present invention includes a data acquisition and processing module, a model training and evaluation module, and a classification result output module; among them, the data acquisition and processing module executes the steps of data acquisition and data processing to obtain and process multi-modal data; the model training and evaluation module constructs a model based on a deep learning architecture and performs model training, evaluation, and interpretability analysis, including operations such as setting the loss function, optimizer, model warm-up, calculating evaluation metrics, and obtaining the importance of modalities; the classification result output module receives the output results of the model training and evaluation module and outputs and displays the final vocal cord lesion classification results.

[0117] The specific verification process is as follows:

[0118] In the data acquisition stage, the established process is strictly followed. The data source is the real data of clinical polyps, precancerous lesions, and cancer patients in the otolaryngology department of the cooperative hospital from 2020 to 2024. These data have successfully passed the ethical review of the cooperative hospital, ensuring the reliability and legality of the source. Taking the clinical patients as the population, clinical sample data is carefully selected by simple sampling method and clearly divided into three categories: polyps, precancerous lesions, and cancer, and clearly labeled, laying a solid foundation for subsequent research.

[0119] In the data integration process, we closely refer to the knowledge ontology in the field of vocal cord cancer and follow the clinical interpretation. We carefully construct data resources, stipulating that each data resource consists of 1 biochemical indicator, 3 images and 2 audio clips, and integrate them using the Cartesian product method. In the case of insufficient digits, "0" is strictly used to fill in the gaps to ensure the integrity and standardization of the data.

[0120] Next is the data processing stage, where special processing measures are taken for different types of data.

[0121] Image processing: Due to the inconsistency of image data bit depth, some are 32 bits (4 channels) and some are 24 bits (3 channels). In order to achieve uniform bit depth, the transparency channel of the 4-channel data is decisively discarded, and all data is forced to be converted to RGB format. Given that the image resolution provided is generally around 400*400, in order to fix the data length, the image is first scaled to 256*256. In addition, it is observed that the data border is black and useless information, and the key information is concentrated in the center of the image. Combined with the Transformer input requirements, the data is further center-cropped, and the size after cropping is accurately determined to be 224*224. Then standardization is performed to fix the image pixel value range to 0-255, and normalization is achieved by dividing each pixel by 255, so that the pixel value range becomes a floating point number [0.0,1.0]. Finally, the patch embedding operation is performed to divide the image into 16*16 patches. A 16*16 convolution kernel, a step size of 16 and a linear projection to a 768-dimensional embedding space are used to successfully convert the image from 1*3*224*224 to 1*196*768 features.

[0122] In terms of audio processing: The read audio is a one-dimensional array. First, a centering strategy is adopted to eliminate the mean shift, and the formula is X_centered = X - μ (where X is the original data matrix and μ is the mean vector of each feature). Based on the non-linear perception principle of the human auditory system, by framing the audio signal, applying a Hanning Window, and performing a discrete Fourier transform (DFT), the time-domain signal is converted into a frequency-domain signal, and 128 frequency bands are divided. Taking 10 seconds as the length benchmark, the audio with a time-domain dimension of approximately 1024 is processed. The part exceeding 1024 is truncated, and the part less than 1024 is padded with 0s to fix the time-domain length at 1024, obtaining a feature tensor of 1*1024*128. Then, a masking operation is performed. Based on the fact that the usual mask is 10%-20% in the frequency domain and time domain, it is calculated that the frequency-domain mask is 24 and the time-domain mask is 256. The Mel-Frequency Cepstral Coefficients (MFCC) are randomly occluded in the time-domain and frequency-domain dimensions to improve the generalization ability of the model. At the same time, random noise in the range of 0-1 is added in both the time-domain and frequency-domain dimensions to enhance the robustness of the model and improve the generalization ability. A rolling operation with a unit of 10 in the forward or reverse direction is performed in the time-domain dimension to increase data diversity. Finally, the mean and standard deviation of the Mel-Frequency Cepstral Coefficients (MFCC) are normalized, and the formula is fbank = (fbank - fbank.mean()) / fbank.std(). The MFCC is segmented into 16*16 patches, and a 16*16 convolutional kernel, a time-domain and frequency-domain stride of 10, an overlap coefficient of 6, and a linear projection into a 768-dimensional embedding space are used to change the MFCC from 1*1024*128 to a feature of 1*1212*768.

[0123] In terms of biochemical index processing: By carefully observing the data, it is found that there are missing values in some columns. At this time, the Multiple Imputation by Chained Equations (MICE) multiple imputation method is adopted, and a MICE model is constructed based on all the data to impute the data. Then, the maximum and minimum values are calculated separately in the column direction, and the maximum and minimum values are used for normalization in the row direction, and the formula is (x - min(x)) / (max(x) - min(x)). Finally, a fully connected linear layer is used to map 51 biochemical indexes to a 768-dimensional space to make it have the same dimension as the image and audio, and the shape changes from 1*51 to 1*1*768.

[0124] In the model construction stage, the network model consists of an embedding layer, an encoder layer, and an output layer in sequence. In the embedding layer, after feature embedding of all modal data, feature concatenation is performed in the first dimension. For example, after processing, the image data is 1*196*768, the audio data is 1*1212*768, and the biochemical index is 1*1*768. After concatenation, it becomes 1*1409*768. Then, a classification token with a shape of 1*1*768 is added, and the final shape is 1*1410*768. And positional encoding (PositionalEncoding) is created to embed the data, enabling the model to understand the data order. In the encoder layer, an Encoder layer composed of 12 Block blocks is designed. Each Block block uses the multi-head self-attention mechanism to calculate the correlation between each position in the sequence and other positions, which is of great significance for cross-modal data and helps to discover the semantic associations between different modal data. For example, in a sequence containing speech and video information, the model can use this mechanism to determine the degree of association between a certain feature in the speech and a certain area in the video. Then, the features processed by self-attention are non-linearly transformed through a feed-forward neural network to further enhance the model's expressive ability. It contains two fully connected layers, and each layer uses GELU+LayerNorm for non-linear activation and layer normalization. The output layer performs the classification task, and a fully connected layer is used to convert the 0th dimension of the features output by the Encoder layer into class probabilities.

[0125] During the experimental setup, binary cross-entropy (BCEWithLogitsLos) is used for binary classification, and cross-entropy (CrossEntropyLoss) is used for multi-classification. The optimizer uses the Adaptive Moment Estimation (Adam) optimizer, with an initial learning rate of 1e-5, an L2-norm regularization weight decay value of 5e-7, an exponential decay rate of 0.95 for the first moment estimate, and an exponential decay rate of 0.999 for the second moment estimate. At the same time, 200 batches of model warm-up are designed, and the learning rate is dynamically adjusted within the first 200 batches to ensure that the model reaches a stable training state. Evaluation metrics are calculated in each round of training, including accuracy, average precision, and AUC, to comprehensively evaluate the model performance. In terms of model interpretability, Gradient-weighted Class Activation Mapping (Grad-Cam) is used to calculate the average gradient of each modality, and the importance of each modality is explained according to the average gradient score; attention weights (Attention-Weights) are calculated, and the weight score is obtained through the average weight of each modality; through the average SHapley Additive exPlanations (Shap value) of each modality, the contribution rate of each modality to the classification result is analyzed.

[0126] After completing the model training, with the trained model as the core, we will fully develop the "Method for Extracting Lesion Features of Vocal Cord Cancer" and put it into the clinical testing process. We will comprehensively evaluate and optimize the model using the results obtained from the clinical testing, continuously improve the software system, and enhance the accuracy and reliability of the system to provide strong support for the method for extracting lesion features of vocal cord cancer.

[0127] Through the above implementation methods, the present invention effectively integrates multi-modal data, uses a deep learning architecture to achieve a method for extracting lesion features of vocal cord cancer, and has also actively explored in terms of model interpretability, which is expected to play an important role in the early screening and diagnosis of vocal cord cancer and improve the diagnostic efficiency and accuracy.

Claims

1. A method for extracting vocal cord cancer lesion features based on multimodality, characterized in that: The method comprises the following steps: The first step is to obtain biochemical indicators, audio and laryngoscope images of patients with polyps, precancerous lesions and cancers in the otolaryngology department of the cooperative hospital over the past few years as multimodal data, and pass the ethical review of the cooperative hospital. Simple sampling is used to obtain clinical samples and clearly mark them. With reference to the knowledge ontology in the field of vocal cord cancer, data resources are constructed, and the Cartesian product method is used to fill the insufficient digits with "0"; Step 2: Image processing: All data with inconsistent bit depth are converted to RGB format, the data resolution is scaled and center-cropped, the image pixel values ​​are normalized, and the image is divided into 16*16 patches and linearly projected into the embedding space; Audio processing: A centralized strategy is used for audio to extract Mel frequency cepstral coefficients, and masking, random noise processing, rolling processing, and standardization are performed. The Mel frequency cepstral coefficients are divided into 16*16 patches and linearly projected to the embedding space. Biochemical index processing: Chain equation multiple filling is performed on biochemical data, and all data are normalized and feature embedded. The third step is to build a network model based on the deep learning architecture. The network model consists of an embedding layer, an encoder layer, and an output layer. The embedding layer fuses multimodal features and creates position encoding. The multi-head self-attention mechanism of the encoder layer discovers the semantic association between different modal data, and the feedforward neural network enhances the expression ability of the model. The output layer completes the classification task. In the fourth step, the cross entropy loss function and adaptive optimizer were used to warm up the model. The model performance was evaluated by accuracy, average precision and AUC. Gradient weighted class activation mapping, attention weight and Shap value were used to implement model interpretability analysis, making full use of multimodal data and deeply exploring the internal connections as clues and basis for early screening and classification of vocal cord lesions.

2. The method for extracting vocal cord cancer lesion features based on multimodality according to claim 1, characterized in that: The specific steps of the first step are as follows: Step 1. The data comes from the real data of patients with clinical polyps, precancerous lesions and cancers in the ENT department of the cooperative hospital over the past few years. The data must pass the ethical review of the cooperative hospital. The data source is authentic and reliable. The clinical patients are taken as the overall population and the clinical sample data is obtained by simple sampling. The data are divided into three categories: polyps, precancerous lesions and cancers, and are clearly marked. Step 2: Following clinical interpretation and referring to the knowledge ontology in the field of vocal cord cancer, data resources were constructed. Each data resource consisted of one biochemical indicator, three images, and two audio clips. Cartesian product was used to integrate the above data resources into 1*2*3, a total of 6 data. That is, each data contained only one biochemical indicator, one image, and one audio clip. Insufficient digits were filled with "0" to ensure data integrity and standardization.

3. The method for extracting vocal cord cancer lesion features based on multimodality according to claim 1, characterized in that: The specific steps of the second step are as follows: Step 1: Image processing, as follows: (1) Bit depth inconsistency processing: In the image data, some data has a bit depth of 32, that is, 4 channels; some data has a bit depth of 24, that is, 3 channels. To ensure the consistency of bit depth, the transparency channel of the 4-channel data is discarded and all data is forcibly converted to RGB format; (2) Scaling and center cropping: In order to fix the data length, the image is scaled to 256*256. The data border is observed to be black and useless information, and the key information is concentrated in the center of the image. Combined with the input of the Transformer, the data is center cropped, and the size after cropping is 224*224. (3) Standardization: The pixel value range of the image is fixed to 0-255. Each pixel is divided by 255 for normalization. The normalized pixel value range is a floating point number of [0.0, 1.0]. (4) Patch embedding: The image is divided into 16*16 patches and linearly projected into the embedding space. The specific method is as follows: a 16*16 convolution kernel with a step size of 16 is used, the embedding feature is 768 dimensions, and 2D convolution + linear projection is used. The image is transformed from 1*3*224*224 to 1*196*768 features; Step 2: Audio processing, as follows: (1) Centralization processing: The read audio is a one-dimensional array. A centralization strategy is used to eliminate mean shift. The formula is as follows: X_centered=X-μ X is the original data matrix, each row is a sample, each column is a feature, and μ is the mean vector of each feature; (2) Extracting Mel-frequency cepstral coefficients (MFCC): Based on the nonlinear perception of the human auditory system, the time domain signal is converted into a frequency domain signal by framing the audio signal, adding a Hanning window, and performing a discrete Fourier transform. The signal is divided into 128 frequency bands, and the extracted time domain dimension is approximately 1024, with a length of 10 seconds as the benchmark. The part exceeding 1024 is truncated, and the part less than 1024 is padded with 0. The time domain length is fixed to 1024, and a feature tensor of 1*1024*128 is obtained. (3) Masking: Usually the mask is 10%-20% of the frequency domain and time domain. The calculated frequency domain mask is 24 and the time domain mask is 256. The Mel frequency cepstral coefficients MFCC are randomly masked in the time domain and frequency domain dimensions to improve the generalization ability of the model. (4) Random noise processing: Add random noise in the time domain and frequency domain, ranging from 0 to 1, in order to enhance the robustness of the model and improve the generalization ability; (5) Rolling processing: Randomly roll forward or backward in units of 10 in the time domain dimension to improve data diversity, extract data features from different angles, and achieve data enhancement; (6) Standardization: The mean and standard deviation of the Mel frequency cepstral coefficient MFCC are standardized. The formula is as follows: fbank=(fbank-fbank.mean()) / fbank.std(); Fbank represents the MFCC matrix, mean is the mean, and std is the standard deviation; (7) Patch embedding: The Mel frequency cepstral coefficients (MFCC) are divided into 16*16 patches and linearly projected into the embedding space. Specifically, a 16*16 convolution kernel is used, the step size of the time domain and the frequency domain is 10, overlapping sampling, the overlap coefficient is 6, the embedding feature is 768, and 2D convolution + linear projection is used. The Mel frequency cepstral coefficients (MFCC) are transformed from 1*1024*128 to 1*1212*768 features; Step 3: Biochemical index processing, as follows: (1) Chain equation multiple imputation: Chain equation multiple imputation is used to construct the MICE model based on all data and impute the data; (2) Normalization: Calculate the maximum and minimum values ​​of all data in the column direction, and use the maximum and minimum values ​​to normalize in the row direction. The formula is (x-min(x)) / (max(x)-min(x)); (3) Feature embedding processing: A fully connected linear layer is used to map the 51 biochemical indicators into a 768-dimensional space, keeping the same dimension as the image and audio, and the shape changes from 1*51 to 1*1*768.

4. The method for extracting vocal cord cancer lesion features based on multimodality according to claim 1, characterized in that: The specific steps of the third step are as follows: The network model consists of an embedding layer, an encoder layer, and an output layer, as follows: Step 1, Embedding layer: After embedding all modal data, perform feature concatenation in the first dimension. The image is 1*196*768, the audio is 1*1212*768, and the biochemical index is 1*1*768, which is concatenated to 1*1409*768. In addition, a classification with a shape of 1*1*768 is added, and the final shape is 1*1410*768. In order to make the model understand the data order, create a positional encoding and embed it into the data. Step 2, encoder layer, as follows: (1) Multi-head self-attention mechanism: An encoder layer consisting of 12 blocks is designed. Each block uses a multi-head self-attention mechanism to calculate the correlation between each position in the sequence and other positions, allowing the model to focus on different parts of the sequence. In a sequence containing text and image information, the model can determine the degree of correlation between a word in the text and a region in the image through the self-attention mechanism. (2) Feedforward neural network: Performs nonlinear transformation on the features after self-attention processing to further enhance the model's expressiveness. It processes the features at each position independently and includes two fully connected layers. Each layer uses GELU+LayerNorm for nonlinear activation and layer normalization. Step 3, output layer: classification task, use the fully connected layer to convert the 0th dimension of the feature output of the Encoder layer into category probability.

5. The method for extracting vocal cord cancer lesion features based on multimodality according to claim 1, characterized in that: The specific steps of the fourth step are as follows: Step 1, loss function: binary cross entropy is used for binary classification, and cross entropy is used for multi-classification; Step 2, optimizer: The optimizer uses an adaptive optimizer, with an initial learning rate of 1e-5, an L2 norm regularization weight decay value of 5e-7, an exponential decay rate of 0.95 for the first-order moment estimate, and an exponential decay rate of 0.999 for the second-order moment estimate; Step 3: Model preheating: Design 200 batches of model preheating. In the first 200 batches, dynamically adjust the learning rate to achieve stable training. Step 4: Evaluation indicators: For each round, calculate the evaluation indicators, including accuracy, average precision and AUC; Step 5: Model interpretability, as follows: (1) Gradient-weighted class activation mapping is used to calculate the average gradient of each modality, and the importance of each modality is explained based on the average gradient score; (2) Calculate the attention weight and obtain the weight score by the average weight of each modality; (3) The contribution rate of each modality to the classification result is analyzed through the average shap value of each modality.

6. A multimodal vocal cord cancer lesion feature extraction and analysis system, characterized by: It includes a data acquisition and processing module, a model training and evaluation module, and a classification result output module; the data acquisition and processing module performs the steps of data acquisition and data processing, and acquires and processes multimodal data; the model training and evaluation module builds a model based on a deep learning architecture, and performs model training, evaluation, and interpretability analysis, including setting loss functions, optimizers, model preheating, calculating evaluation indicators, and obtaining modal importance operations; The classification result output module receives the output results of the model training and evaluation module, and outputs and displays the final vocal cord lesion classification results.

Citation Information

Cited By

  • Multi-modal children voice data processing method based on deep learning and federated learning

    CN120470245A